ROOT / AI // FOUNDATION MODELS
// AI · decoded · publication-grade pass
Foundation Models, decoded.
The frontier model lineup, benchmarked on context, reasoning and coding — hype removed.
CURATED INDEX◊ VERIFIED · multi-source
// the signal · what the data says
The frontier is a rotating podium. As of mid-2026, Claude Opus 4.8 leads real-world coding (SWE-bench Pro 69.2%, 10+ points clear of the field), Gemini 3.1 Pro leads raw reasoning (GPQA-Diamond 94.3%, ARC-AGI-2 77.1% — double the prior gen), and GPT-5.5 holds the creative/general lead. Nearly the whole frontier now offers a ~1M-token context window. Open-weight DeepSeek V4 and MiniMax M3 trail on peak capability but undercut price by 10x. There is no single 'best' — you match the model to the task.
// comparison · every entry, benchmarked & sourced
| Name | Maker | Context (k tok) | GPQA (%) | SWE-Pro (%) | Out cost ($/M) |
|---|
| Claude Opus 4.8 | Anthropic | 1000 | 91 | 69.2 | 25 |
| ◊ Overall + coding leader; SWE-bench Pro 69.2% |
| Gemini 3.1 Pro | Google | 1000 | 94.3 | 54.2 | 10 |
| ◊ Top GPQA 94.3%, ARC-AGI-2 77.1%; best value |
| GPT-5.5 | OpenAI | 1000 | 89 | 58.6 | 15 |
| ◊ Creative/general strength; ~1M context |
| Grok 4.3 | xAI | 256 | 88 | 50 | 12 |
| ◊ Trained on Colossus; strong reasoning |
| DeepSeek V4 | DeepSeek | 1000 | 84 | 45 | 2 |
| ◊ Open weights; ~10x cheaper |
| MiniMax M3 | MiniMax | 1000 | 80 | 40 | 1 |
| ◊ Open, 1M context, lowest cost |
// leaderboard · Context (k tok)
01Claude Opus 4.81000
02Gemini 3.1 Pro1000
03GPT-5.51000
04DeepSeek V41000
05MiniMax M31000
06Grok 4.3256
// sources & where to verify
// full index, per-field sourcing & CSV export [ AUTHENTICATE → ]
◊ Publication-grade pass, mid-2026. Cross-checked against the sources listed. Green = best value in column. Footnotes: Params are undisclosed for closed models. GPQA = graduate-science reasoning; SWE-bench Pro = real coding tasks. Some GPQA figures are best-available estimates; cost is output $/M tokens.