Vera Rubin NVL72 brings CoreWeave's AI factory lifecycle from validation through operations

CoreWeave achieved up to 96% goodput on NVIDIA Hopper GPUs for long-running distributed AI training, with 20% higher MFU than other reported benchmarks.

Categorized in: AI News Operations
Published on: Aug 13, 2026
Vera Rubin NVL72 brings CoreWeave's AI factory lifecycle from validation through operations

Long-running distributed AI training jobs don't leave room to notice an interruption a day late. When a cluster moves from validated to operational, the question shifts from whether it works to whether it keeps working - and how much of that compute time turns into finished model progress instead of recovery time. CoreWeave and NVIDIA have released the second entry in their AI factory lifecycle series, detailing the operating discipline required after a system goes live.

The number that governs AI factory economics is goodput - the share of time a system spends on useful work rather than recovering from interruptions. On long-running distributed training workloads, CoreWeave has demonstrated up to 96% goodput on NVIDIA Hopper GPUs, presented as an upper bound rather than a guarantee. Everything below that line is lost productivity.

Two variables drive goodput

Throughput. Model FLOPs utilization (MFU) measures how much of a GPU's peak throughput becomes model progress. CoreWeave measured 20% higher MFU on Hopper-based systems than other publicly reported benchmarks. MFU is workload-dependent - it varies with model architecture, sequence length, and parallelism - so the figure reflects its test context rather than a universal result.

Resiliency. Mean time to failure improved roughly 10x, with an effective training time ratio of up to 98%. Restarts and lost checkpoints stop setting the timeline when interruptions drop by that much. Those gains come from fewer interruptions and faster recovery, not from eliminating failures entirely.

On the NVIDIA side, each new generation improves resiliency. Blackwell introduced a dedicated RAS engine applying AI-based predictive maintenance across thousands of hardware and software data points. Vera Rubin carries a second-generation RAS engine for proactive maintenance and real-time health checks without downtime. Modular cable-free trays speed assembly and simplify serviceability, while software-defined NVLink routing reroutes around faults.

The operational layer underneath

CoreWeave Mission Control processes over 200 million metrics samples per second across all customer environments. Its job is to catch and contain infrastructure failures before they cost a customer a training run, through straggler detection, automated node draining, and recovery workflows. SUNK, CoreWeave's Slurm on Kubernetes offering, dynamically reallocates GPUs across training and inference workloads on the same cluster as demand shifts.

CoreWeave completed the industry-first bring-up and validation of NVIDIA Vera Rubin NVL72 in June 2026. Customers inherit a validated system instead of debugging new architecture in production. Six purpose-built chips - Vera CPU, Rubin GPU, NVLink 6 Switch, ConnectX-9 SuperNIC, BlueField-4 DPU, and Spectrum-6 Ethernet Switch - are unified into a single supercomputer.

At the rack level, CoreWeave built two patent-pending innovations. Valvey is a per-rack liquid-cooling valve assembly that turns cooling into a software-defined control surface, monitoring flow rate, temperature, and pressure continuously, and isolating a rack automatically without taking down neighboring racks. Racky is the per-rack control point aggregating power, cooling, and environmental sensors into one standardized surface.

Craig Falls, Head of Quantitative Research at Jane Street, said: "Our research depends on infrastructure that's both powerful and reliable, and CoreWeave has delivered on this as we've scaled across NVIDIA Hopper and Blackwell. Their ability to deliver highly performant clusters with full cluster observability gives us the confidence to partner with them on NVIDIA Vera Rubin."

CoreWeave shared the first measured Vera Rubin performance results: up to 10x more tokens per second per megawatt than NVIDIA GB200 NVL72, tested using the same DeepSeek R1 workload at a matched interactivity target with all major optimizations enabled - expert parallelism, NVFP4 precision, multi-token prediction, and disaggregated prefill and decode.

Why this matters for operations leaders

For operations professionals managing AI infrastructure, the takeaway is that the hard part was never the silicon. It's the network, storage, scheduler, and recovery path behaving as one system while a run is live, and the discipline that keeps them there day after day. The lifecycle - not the launch - determines whether a factory delivers. CoreWeave and NVIDIA plan for the next generation using reference designs from the NVIDIA DSX Platform, which applies intelligent optimizations, power policies, and simulation tools before hardware is deployed physically. Operations teams looking to build and run these systems can explore structured learning through resources like the AI for Operations Managers Learning Path or the Vice Presidents of Operations Learning Path.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)