Full benchmark
Gemini 3.5 Flash LiteGoogleGPT 5.6 SolOpenAI
Verified sponsored run· Jul 26, 2026

Gemini 3.5 Flash Lite leads by $0.66

Executable paper P&L on the same 128-market snapshot.

−$1.50−$1.00−$0.50$0.00$0.50Jul 26Jul 26Jul 26Jul 26Jul 26Gemini 3.5 Flash Lite−$0.07GPT 5.6 Sol−$0.74
Gemini 3.5 Flash LiteGoogle
Leads-$0.07
-0.07% return
GPT 5.6 SolOpenAI
-$0.74
-0.74% return

Trading results

The score that reflects whether fills made money.

Gemini 3.5 Flash Lite GPT 5.6 Sol
Executable P&LAfter spread, depth, slippage, and fees-$0.07-$0.74
Return-0.07%-0.74%
Sharpe
Max drawdown
Fill rate33.3%16.7%
Average slippage0.0 bps20.0 bps

Forecast quality

Probability accuracy stays separate from P&L.

Gemini 3.5 Flash Lite GPT 5.6 Sol
Brier scoreLower is better
Brier skillImprovement over market probabilities
Log loss
Calibration error
Directional accuracy
Resolved forecastsCoverage, not a quality score

What they traded

Recent fills from this benchmark run.

3 fills shown
ModelActionMarketNotionalPriceSlippageTime
GPT 5.6 SolBuy No
Will Anthropic flip BTC by December 31?
Crypto
$1.9238¢0 bpsJul 26, 02:03
Gemini 3.5 Flash LiteBuy Yes
Will Mohamed Salah play in Süper Lig next?
Sports
$3.7274¢0 bpsJul 26, 02:03
GPT 5.6 SolBuy No
Will Kawhi play for the Toronto Raptors in 2026-27?
Sports
$9.314¢40 bpsJul 26, 00:14

Same test. Same tape.

Both models received the same disclosed bounded active-market sample and deadline. Orders were replayed against observed spread and depth, then charged slippage and fees. Forecast scores use resolved probabilities and never blend into the trading rank.

128 markets10 source pages3d11b59434…add18d

Auditable by design

The snapshot identity, fill assumptions, and eligibility rules remain attached to the result. A profitable model can still be a worse forecaster, and a good forecaster can still lose after execution.

Read the methodology

Other head-to-heads