ROOT / AI // FOUNDATION MODELS
// AI · decoded · publication-grade pass

Foundation Models, decoded.

The frontier model lineup, benchmarked on context, reasoning and coding — hype removed.

CURATED INDEX◊ VERIFIED · multi-source
94.3%
top GPQA (Gemini)
69.2%
top SWE (Claude)
1M
frontier context
$1
cheapest /M
// the signal · what the data says
The frontier is a rotating podium. As of mid-2026, Claude Opus 4.8 leads real-world coding (SWE-bench Pro 69.2%, 10+ points clear of the field), Gemini 3.1 Pro leads raw reasoning (GPQA-Diamond 94.3%, ARC-AGI-2 77.1% — double the prior gen), and GPT-5.5 holds the creative/general lead. Nearly the whole frontier now offers a ~1M-token context window. Open-weight DeepSeek V4 and MiniMax M3 trail on peak capability but undercut price by 10x. There is no single 'best' — you match the model to the task.
// comparison · every entry, benchmarked & sourced
NameMakerContext (k tok)GPQA (%)SWE-Pro (%)Out cost ($/M)
Claude Opus 4.8Anthropic10009169.225
◊ Overall + coding leader; SWE-bench Pro 69.2%
Gemini 3.1 ProGoogle100094.354.210
◊ Top GPQA 94.3%, ARC-AGI-2 77.1%; best value
GPT-5.5OpenAI10008958.615
◊ Creative/general strength; ~1M context
Grok 4.3xAI256885012
◊ Trained on Colossus; strong reasoning
DeepSeek V4DeepSeek100084452
◊ Open weights; ~10x cheaper
MiniMax M3MiniMax100080401
◊ Open, 1M context, lowest cost
// leaderboard · Context (k tok)
01Claude Opus 4.81000
02Gemini 3.1 Pro1000
03GPT-5.51000
04DeepSeek V41000
05MiniMax M31000
06Grok 4.3256
// sources & where to verify
// full index, per-field sourcing & CSV export   [ AUTHENTICATE → ]
◊ Publication-grade pass, mid-2026. Cross-checked against the sources listed. Green = best value in column. Footnotes: Params are undisclosed for closed models. GPQA = graduate-science reasoning; SWE-bench Pro = real coding tasks. Some GPQA figures are best-available estimates; cost is output $/M tokens.
← Hyperscaler Capex◊ indexInference Cost →