ROOT / AI // INFERENCE COST
// AI · decoded · deep-research pass
Inference Cost, decoded.
Price and speed per million tokens across every serving provider.
LIVE · feed-ready◊ confidence: Medium
// the signal · what the data says
Inference economics decide who can deploy AI at scale. Specialized silicon (Groq, Cerebras) undercuts latency dramatically; open-weight models on commodity clouds undercut price. Prices fall roughly an order of magnitude per year, so this table must be refreshed continuously.
// comparison · every entry, benchmarked & annotated
| Name | Maker | Input ($/M) | Output ($/M) | Latency (ms) | Context (k) |
|---|
| GPT-frontier | OpenAI | 5 | 15 | 320 | 256 |
| ◊ Premium frontier pricing |
| Gemini | Google | 3.5 | 10 | 280 | 1000 |
| ◊ Cheapest long-context frontier |
| Claude | Anthropic | 3 | 15 | 300 | 500 |
| ◊ Strong agentic value |
| Llama 4 (Groq) | Groq | 0.6 | 0.8 | 90 | 256 |
| ◊ LPU, sub-100ms |
| DeepSeek V4 | DeepSeek | 0.3 | 0.5 | 400 | 128 |
| ◊ Lowest headline price |
// leaderboard · Input ($/M)
01DeepSeek V40.3
02Llama 4 (Groq)0.6
03Claude3
04Gemini3.5
05GPT-frontier5
// sources & where to verify
// full index, per-field sourcing & CSV export [ AUTHENTICATE → ]
◊ Deep-research pass, mid-2026. Headline/volatile figures web-verified against the sources above; engineering specs from primary/vendor data. Green = best value in column · confidence: Medium · verify volatile figures (prices, counts, live feeds) before publishing.