Pull to refresh
Logo
Clockwork.io ships multi-node snapshots that cut AI training recovery to under 20 seconds

Clockwork.io ships multi-node snapshots that cut AI training recovery to under 20 seconds

New Capabilities

Company raises $31M as LinkedIn and Together AI deploy its GPU fault-tolerance software

Yesterday: Clockwork ships multi-node snapshots, raises $31M

Overview

Updated 41 minutes ago

A GPU fails mid-run on a distributed AI training job, and the run rolls back hours or days of compute. Clockwork.io's TorchPass snapshots now capture a whole multi-node job across every GPU in under 20 seconds, with no changes to the training code.

The capability, announced October 5 alongside a $31 million funding round, is the first shipped multi-node snapshot for training; NVIDIA's own version is still in prototype. LinkedIn and Together AI run Clockwork's software in production, and WhiteFiber is expanding its use. NVIDIA expects general availability by the end of 2026, so the competitive window may be short.

Why it matters

A failed GPU can wipe out days of expensive AI training; Clockwork's snapshots now cut recovery to under 20 seconds.

Questions about this story

Free account needed to ask — your question is kept and asked for you right after sign-up. Answers are public.

No questions yet — be the first to ask.

Key Indicators

$31M
New funding round
Co-led by Premji Invest, Wing Venture Capital and Seligman Ventures.
$73M
Total funding to date
Includes earlier investment from NEA and e& Capital.
Under 20 seconds
Multi-node snapshot pause time
Measured on an 8-node H200 Megatron-LM test, about 72 GiB of GPU memory per rank.
1.1%
Snapshot overhead at 30-minute intervals
Rises to about 3.3 percent of wall-clock time at 10-minute intervals.
Tens of thousands
GPU-hours saved per month at LinkedIn
Attributed to Clockwork CEO Raghu Hiremagalur, via LinkPass.

Voices

Curated perspectives — historical figures and your fellow readers.

Ever wondered what historical figures would say about today's headlines?

Sign up to generate historical perspectives on this story.

People Involved

Organizations Involved

Timeline

September 2026 October 2026

2 events Latest: Yesterday
  1. Clockwork ships multi-node snapshots, raises $31M

    Latest Funding / Product launch

    Clockwork announced multi-node platform snapshots for training, fast asynchronous application checkpoints, a $31 million round, and production deployments at LinkedIn and Together AI.

  2. NVIDIA reveals prototype multi-node checkpointing

    Announcement

    NVIDIA announced NCCL prototype support for cuda-checkpoint allowing multi-node checkpoints, with general availability expected by the end of 2026.

Scenarios

1

NVIDIA ships native multi-node checkpointing, narrowing Clockwork's edge

Likely Resolves by Feb 28, 2027

Discussed by: NVIDIA's September 15 announcement; Supercomputing.news analysis of the competitive timeline

NVIDIA said September 15 that its NCCL library has prototype support for multi-node cuda-checkpointing, with general availability expected by the end of 2026. If the native feature ships and works well, platform teams may prefer driver-level checkpointing built into the stack they already run. Clockwork counters that its snapshots need no code changes and pause the whole job at a coordinated safe point.

2

Clockwork expands beyond LinkedIn and Together AI to new production fleets

Possible Resolves by Q2 2027

Discussed by: Clockwork's October 5 release, which earmarks the capital for enterprise adoption and cloud-partner delivery

Clockwork says the capital will accelerate deployment across training, inference and reinforcement learning, and scale delivery through cloud partners. Additional named customers would show the platform is gaining traction beyond early adopters LinkedIn, Together AI and WhiteFiber.

3

Together AI launches managed GPU fault tolerance for tenants

Possible Resolves by End of 2027

Discussed by: Clockwork and Together AI's joint plan, reported by Supercomputing.news

Together AI is integrating LinkPass and TorchPass into its GPU clusters, planning to make them opt-in for tenants, though the details are still being worked out. A launch would put Clockwork's tools in front of any model builder renting GPU time, expanding the market beyond enterprises that run their own fleets. The two companies plan to demonstrate the technology publicly at the PyTorch Conference on October 20-21.

Historical Context

3 moments from history that rhyme with this story — and how they unfolded.

1974–1976

Tandem Computers builds fault-tolerant systems (1974)

Tandem Computers was founded in 1974 and shipped the NonStop in 1976, the first commercially successful computer that kept running through hardware failure. Banks and telecom carriers adopted it because a single failed component never interrupted a transaction.

Then

Tandem dominated the fault-tolerant server market for two decades.

Now

Turned 'hardware will fail' into a design assumption for critical systems, a lesson mainstream computing absorbed gradually.

Why this matters now

Clockwork applies the same assumption to AI training: hardware failures are inevitable, so the system should absorb them without losing work.

2003

VMware vMotion (2003)

VMware's ESX Server 2.0 introduced VMotion, which moved a running virtual machine between physical servers with no downtime. What once required stopping the workload became a live operation.

Then

Live migration became a standard feature of virtualization platforms and underpinned the cloud's ability to move work between hosts.

Now

Established the pattern of relocating running workloads; Clockwork's TorchPass extends the same idea to GPU state across many machines.

Why this matters now

The principle, move live work instead of restarting it, is exactly what Clockwork applies to distributed neural network training.

2022

Microsoft Singularity (2022)

Microsoft researchers published the Singularity paper, describing a scheduler that could transparently preempt a live deep learning job, migrate it to different nodes, and resume it without code changes on Microsoft's own AI fleet.

Then

Demonstrated the technique in-house but did not ship it as a commercial product.

Now

Showed transparent migration of distributed AI jobs was technically feasible, opening the path for companies like Clockwork to productize it.

Why this matters now

Clockwork turns that research idea into software platform teams can buy and deploy without touching training code.

Sources

(8)