Pull to refresh
Logo
CoreWeave brings up first multi-rack NVIDIA Rubin cluster

CoreWeave brings up first multi-rack NVIDIA Rubin cluster

New Capabilities

Seven Vera Rubin NVL72 racks (504 GPUs) run as one production cluster across two regions

Yesterday: First multi-rack Rubin cluster goes live

Overview

Updated 56 minutes ago

CoreWeave has become the first cloud provider to run NVIDIA's newest AI chips at multi-rack scale. On September 16 it brought up seven Vera Rubin NVL72 racks — 504 Rubin GPUs — as a single production cluster spanning two regions.

The jump from a single validated rack to a multi-rack fabric matters because agentic AI workloads make many small, latency-sensitive calls. Delays compound across repeated model and tool interactions. CoreWeave says the multi-rack setup lets agents iterate faster, and new storage features cut data-access latency eightfold.

Why it matters

CoreWeave holds the first-mover slot on NVIDIA's Rubin generation — and agentic AI customers pay a premium for low-latency, rack-scale compute.

Questions about this story

Free account needed to ask — your question is kept and asked for you right after sign-up. Answers are public.

No questions yet — be the first to ask.

Key Indicators

504
Rubin GPUs in first multi-rack cluster
Seven NVL72 racks, each with 72 GPUs, running as one fabric across two regions.
72
Rubin GPUs per NVL72 rack
Each rack pairs 72 GPUs with 36 Vera CPUs, NVLink 6, ConnectX-9 SuperNICs, and BlueField-4 DPUs.
1.6 Tb/s
Scale-out bandwidth per GPU
Two ConnectX-9 SuperNICs per GPU support roughly 128,000 GPUs per rail in a non-blocking fabric.
8x
Latency reduction from local caching
CoreWeave's LOTA caching reads at local NVMe speeds, up to 7 GB/s per GPU.

Voices

Curated perspectives — historical figures and your fellow readers.

Ever wondered what historical figures would say about today's headlines?

Sign up to generate historical perspectives on this story.

People Involved

Organizations Involved

Timeline

June 2026 September 2026

3 events Latest: Yesterday
  1. First multi-rack Rubin cluster goes live

    Latest Deployment

    CoreWeave connected seven racks (504 Rubin GPUs) into one production cluster across two regions.

  2. CoreWeave adds cross-region writes and Archive tier

    Announcement

    AI Object Storage gained background cross-region replication and a low-cost Archive tier for checkpoints and datasets.

  3. Industry-first single-rack Rubin validation

    Milestone

    CoreWeave completed end-to-end bring-up and validation of the first Vera Rubin NVL72 rack.

Scenarios

1

CoreWeave scales Rubin to thousands of GPUs by mid-2027

Likely Resolves by Q2 2027

Discussed by: 650 Group's analysis of the bring-up and CoreWeave's stated expansion plans

The fabric design supports roughly 128,000 GPUs per rail without redesign, so adding racks is configuration rather than re-engineering. If demand from multi-step agentic workloads holds, CoreWeave's first seven racks grow into dozens, and customers run training and reinforcement-learning jobs across thousands of Rubin GPUs.

2

A hyperscaler matches with multi-rack Rubin by Q2 2027

Possible Resolves by Q2 2027

Discussed by: Industry analysts covering the GPU cloud race, including 650 Group

CoreWeave's head start is not permanent; rivals want Rubin clusters as proof of scale. Any delivery slip on CoreWeave's or NVIDIA's side opens a window for another cloud to announce its own multi-rack Rubin fabric and bid for the same workloads.

3

Rubin ramp slows on power or liquid-cooling constraints

Uncertain Resolves by Q2 2027

Discussed by: Datacenter infrastructure observers; CoreWeave's own rack-automation effort signals the risks

Rack-scale Rubin demands liquid cooling and substantial power per rack. CoreWeave built Racky and Valvey specifically to automate rack and cooling control, a sign these systems break when power or thermal limits are mismanaged. If datacenter sites cannot meet the power draw or cooling loops fail, expansion slips and early clusters run below capacity.

Historical Context

2 moments from history that rhyme with this story — and how they unfolded.

2023–2024

H100 Capacity Race (2023–2024)

When OpenAI's success made NVIDIA H100 GPUs the scarcest resource in tech, cloud providers scrambled to secure supply. CoreWeave, then a small GPU renter, locked in early access and landed a multibillion-dollar Microsoft-backed deal for its capacity.

Then

CoreWeave became one of the fastest-growing AI infrastructure providers almost overnight.

Now

The episode showed that whoever secures scarce chips first wins enterprise contracts, even without a legacy cloud brand.

Why this matters now

The same first-mover dynamic is playing out again with Rubin multi-rack deployments.

2025

Blackwell GB200 NVL72 rollout (2025)

NVIDIA's previous generation, Blackwell, introduced the same rack-scale NVL72 concept with liquid cooling. Early deployments hit integration problems — power delivery, cooling loops, and firmware — that slowed ramp-ups at several clouds.

Then

Some NVL72 deployments slipped by quarters; only providers with the most mature rack engineering brought clusters up quickly.

Now

The episode set the bar for what it takes to deploy new-generation NVIDIA racks; CoreWeave's Racky and Valvey are direct responses.

Why this matters now

Rubin repeats the rack-scale challenge, and the winners will be those who solved Blackwell's integration lessons.

Sources

(8)