Supermicro has started shipping NVIDIA Vera Rubin NVL72 racks. Rubin is NVIDIA’s next AI platform after Blackwell, and it’s now leaving the factory as liquid-cooled, fully integrated systems. Customers can also order a complete Scalable Unit: 16 racks, 1,152 Rubin GPUs, delivered production-ready.
“Our customers can now order a Scalable Unit and receive production-ready systems with end-to-end integration,” says president and CEO Charles Liang.
What’s in a rack
Each NVL72 rack holds 72 Rubin GPUs and 36 Vera CPUs across 18 compute trays, with four GPUs and two CPUs per tray. Nine trays of sixth-generation NVLink switches tie them together with 216 TB/s of scale-up bandwidth, so the whole rack behaves like one very large accelerator.
Memory is where Rubin makes the biggest jump. Every rack carries 20.7 TB of HBM4, plus up to 54 TB of LPDDR5X attached to the Vera CPUs. A full Scalable Unit adds up to 331 TB of HBM4. That’s the number to watch if you follow the memory makers. HBM4 supply is one of the main limits on how fast these racks can ship.
Cooling is the product now
Supermicro’s pitch is its DCBBS data center building blocks, which cover most of the stack around the servers: cold plates, manifolds and hose kits, rack power shelves, 1.8 MW in-row coolant distribution units, in-rack CDUs, rear-door heat exchangers and the cooling towers outside. Everything can be built with N+1 redundancy.
At these power densities air cooling is off the table. The companies that can deliver the plumbing along with the compute have the advantage, and that is Supermicro’s argument.
Supermicro designs and builds in the US, Taiwan and the Netherlands. It handles site surveys, design, L11 and L12 rack- and cluster-level testing, on-site deployment and ongoing support. Networking follows NVIDIA’s reference architecture. Deployments start around 5 MW and go up to gigawatt scale.
For the AI build-out, Rubin racks leaving the factory mark the start of the next upgrade cycle. Hyperscalers and GPU clouds that are just finishing Blackwell deployments now have to decide how fast to move again.