NVIDIA's next platform is a four-tile module with 288 GB of memory.
The compute frontier's next leap is NVIDIA's Rubin generation, launching in the second half of 2026. Built by TSMC on a 3-nanometre process with next-generation HBM4 memory, the flagship R200 is a true multi-chip module — two compute dies plus two dedicated I/O dies in one package — carrying 288 GB of HBM4 at roughly 22 TB/s of bandwidth.
Rubin is rated at about 50 petaflops of FP4 inference — some 2.5× its Blackwell predecessor — with a "Rubin Ultra" refresh set to double that again and move to TSMC's 2 nm (N2) node around 2027. Early supply goes to hyperscalers and frontier AI labs; broad cloud availability isn't expected until late 2026.