ROOT / COMPUTE // AI ACCELERATORS
// Compute · decoded · publication-grade pass
AI Accelerators, decoded.
Every frontier training and inference accelerator, cross-checked against vendor disclosures and third-party teardowns.
CURATED INDEX◊ VERIFIED · multi-source
// the signal · what the data says
NVIDIA holds ~90% of data-center AI silicon and set the 2026 bar with Rubin (R200): a dual-die, 336-billion-transistor GPU (1.6x Blackwell's 208B) carrying 288 GB of HBM4 across 8 stacks at 22 TB/s, 35 PFLOPS dense FP4 (50 'effective' with sparsity), and NVLink 6 at 3.6 TB/s per GPU — double Blackwell. It ships H2 2026 on every major cloud. AMD's answer, MI455X (MI400 series, Helios rack), actually beats Rubin on memory: 432 GB HBM4 and 23.3 TB/s, with up to 40 PFLOPS FP4. Google (TPU v7 Ironwood) and Amazon (Trainium2) are the credible custom-silicon escapes from NVIDIA margins.
// comparison · every entry, benchmarked & sourced
| Name | Maker | Memory (GB) | Type | Bandwidth (TB/s) | FP4 (PFLOPS) | TDP (W) | Node (nm) |
|---|
| Rubin R200 | NVIDIA | 288 | HBM4 | 22 | 50 | 1800 | 3 |
| ◊ 336B transistors dual-die; 35 PFLOPS dense / 50 effective FP4; NVLink 6 @3.6 TB/s; ships H2 2026 (TDP est.) |
| MI455X | AMD | 432 | HBM4 | 23.3 | 40 | 1500 | 3 |
| ◊ MI400 series, Helios rack; most memory + bandwidth of any 2026 GPU |
| Blackwell Ultra GB300 | NVIDIA | 288 | HBM3e | 8 | 15 | 1400 | 4 |
| ◊ Mid-cycle refresh of GB200; higher FP4 and memory |
| GB200 (Blackwell) | NVIDIA | 192 | HBM3e | 8 | 10 | 1200 | 4 |
| ◊ Per-GPU; NVL72 wires 72 into one 1.4-EF FP4 accelerator |
| MI355X | AMD | 288 | HBM3e | 8 | 10 | 1400 | 3 |
| ◊ CDNA4, 256 CUs; 10.1 PFLOPS MXFP4; 1400W TBP |
| MI325X | AMD | 256 | HBM3e | 6 | 5 | 1000 | 5 |
| ◊ CDNA3, 2024 gen |
| TPU v7 Ironwood | Google | 192 | HBM3e | 7.4 | 9 | 600 | 3 |
| ◊ 9,216-chip pods (~42.5 EF FP8/pod); inference-optimized |
| Trainium2 | Amazon | 96 | HBM3 | 2.9 | 6 | 500 | 5 |
| ◊ Backs Project Rainier for Anthropic; UltraServer of 64 |
| Gaudi 3 | Intel | 128 | HBM2e | 3.7 | 7 | 900 | 5 |
| ◊ Ethernet-native scale-out; value play |
| MTIA v2 | Meta | 128 | HBM3 | 2.6 | 5 | 450 | 5 |
| ◊ Meta internal ranking/inference |
// leaderboard · Memory (GB)
01MI455X432
02Rubin R200288
03Blackwell Ultra GB300288
04MI355X288
05MI325X256
06GB200 (Blackwell)192
07TPU v7 Ironwood192
08Gaudi 3128
09MTIA v2128
10Trainium296
// sources & where to verify
// full index, per-field sourcing & CSV export [ AUTHENTICATE → ]
◊ Publication-grade pass, mid-2026. Specs cross-checked against the multiple sources listed. Green = best value in column. Footnotes: FP4 figures are dense per-GPU unless noted; 'effective' inference PFLOPS assume Transformer-Engine sparsity. TDP is board/module typical.