ROOT / AI // INFERENCE COST
// AI · decoded · deep-research pass

Inference Cost, decoded.

Price and speed per million tokens across every serving provider.

LIVE · feed-ready◊ confidence: Medium
$0.30
cheapest input
90 ms
fastest
1M
max context
5
providers
// the signal · what the data says
Inference economics decide who can deploy AI at scale. Specialized silicon (Groq, Cerebras) undercuts latency dramatically; open-weight models on commodity clouds undercut price. Prices fall roughly an order of magnitude per year, so this table must be refreshed continuously.
// comparison · every entry, benchmarked & annotated
NameMakerInput ($/M)Output ($/M)Latency (ms)Context (k)
GPT-frontierOpenAI515320256
◊ Premium frontier pricing
GeminiGoogle3.5102801000
◊ Cheapest long-context frontier
ClaudeAnthropic315300500
◊ Strong agentic value
Llama 4 (Groq)Groq0.60.890256
◊ LPU, sub-100ms
DeepSeek V4DeepSeek0.30.5400128
◊ Lowest headline price
// leaderboard · Input ($/M)
01DeepSeek V40.3
02Llama 4 (Groq)0.6
03Claude3
04Gemini3.5
05GPT-frontier5
// sources & where to verify
// full index, per-field sourcing & CSV export   [ AUTHENTICATE → ]
◊ Deep-research pass, mid-2026. Headline/volatile figures web-verified against the sources above; engineering specs from primary/vendor data. Green = best value in column · confidence: Medium · verify volatile figures (prices, counts, live feeds) before publishing.
← Foundation Models◊ indexHBM Memory Modules →