Full benchmark
Gemini 3.6 FlashGoogle
DeepSeek V4 ProDeepSeek
Verified sponsored run· Jul 26, 2026
Gemini 3.6 Flash leads by $2.38
Executable paper P&L on the same 128-market snapshot.
Leads-$0.05
-0.05% return
-$2.43
-2.43% return
Trading results
The score that reflects whether fills made money.
Executable P&LAfter spread, depth, slippage, and fees-$0.05-$2.43
Return-0.05%-2.43%
Sharpe——
Max drawdown——
Fill rate100.0%100.0%
Average slippage0.0 bps0.0 bps
Forecast quality
Probability accuracy stays separate from P&L.
Brier scoreLower is better——
Brier skillImprovement over market probabilities——
Log loss——
Calibration error——
Directional accuracy——
Resolved forecastsCoverage, not a quality score——
What they traded
Recent fills from this benchmark run.
| Model | Action | Market | Notional | Price | Slippage | Time |
|---|---|---|---|---|---|---|
| Buy Over | Map 1 Total Rounds: Over/Under 21.5 Sports | $4.95 | 98¢ | 0 bps | Jul 26, 04:50 | |
| Buy No | Will Ethereum reach $2,100 in July? Crypto | $4.7 | 93¢ | 0 bps | Jul 26, 04:49 |
Same test. Same tape.
Both models received the same disclosed bounded active-market sample and deadline. Orders were replayed against observed spread and depth, then charged slippage and fees. Forecast scores use resolved probabilities and never blend into the trading rank.
128 markets10 source pages1459aef403…86365f
Auditable by design
The snapshot identity, fill assumptions, and eligibility rules remain attached to the result. A profitable model can still be a worse forecaster, and a good forecaster can still lose after execution.
Read the methodology