Pull to refresh
Logo
World Labs debuts Atlas, a model that builds and simulates 3D worlds from a few photos

World Labs debuts Atlas, a model that builds and simulates 3D worlds from a few photos

New Capabilities

The 'omni' model handles text, images, video, and 3D, with camera poses and depth as native inputs.

Yesterday: Atlas debuts in early access

Overview

Updated 2 hours ago

Point a phone camera at a room and this model rebuilds the room in 3D, including the parts the camera never saw. On September 1, World Labs, the startup co-founded by Stanford computer vision pioneer Fei-Fei Li, unveiled Atlas.

Atlas is an 'omni' model: one system that handles text, images, video, and 3D, with camera position and depth as native inputs. That lets it control generated video with pixel precision and reconstruct real scenes from as few as two or three photos. World Labs says Atlas beats specialized video and 3D-reconstruction models at their own tasks, though it released no research paper to back the numbers.

Why it matters

If Atlas scales, robotics training, game design, and 3D capture shift from manual pipelines to a single AI that builds worlds from photos.

Questions about this story

Free account needed to ask — your question is kept and asked for you right after sign-up. Answers are public.

No questions yet — be the first to ask.

Key Indicators

1 min @ 1440p
Maximum video generation
Atlas outputs up to a minute of camera-controlled video at 1440p from one or more reference images.
2–3
Minimum photos for faithful reconstruction
World Labs says two or three images typically produce a faithful 3D reconstruction of a real scene.
100+
Maximum input photos for scene reconstruction
Atlas can use more than a hundred images in its spatial context to recreate an environment.
3
Cameras needed for bullet-time reframing
Footage from as few as three ordinary cameras lets Atlas freeze time and shift the viewing angle.
Gemini Omni Flash, FLUX
Rivals beaten in blind camera-path test
Human raters preferred Atlas over Google's Gemini Omni Flash and Black Forest Labs' FLUX for following camera paths.

Voices

Curated perspectives — historical figures and your fellow readers.

Ever wondered what historical figures would say about today's headlines?

Sign up to generate historical perspectives on this story.

Play

Exploring all sides of a story is often best achieved with Play.

Most of these play right now — no account needed. Sign up to save scores, keep a streak, and unlock Debate and Predict. Log in Sign Up
Predict 3 ways this could play out. Back the one you believe — contrarian picks score more when a scenario has a resolution date. Log in to play

People Involved

Organizations Involved

Timeline

February 2024 September 2026

2 events Latest: Yesterday
  1. Atlas debuts in early access

    Latest Product Launch

    World Labs unveils Atlas, an omni world model, and opens early access to select partners.

  2. World Labs founded

    Founding

    Fei-Fei Li and colleagues launch World Labs to pursue spatial intelligence in AI.

Historical Context

3 moments from history that rhyme with this story — and how they unfolded.

March 2020

Neural Radiance Fields (2020)

UC Berkeley and Google researchers showed that a small neural network could encode a 3D scene from dozens of photos and render it from new camera angles. The technique, called NeRF, produced photorealistic novel views that earlier methods could not match.

Then

NeRF became the default method for 3D view synthesis within months, spawning hundreds of follow-up papers.

Now

Each new scene required training a fresh network, so the method stayed slow and demanded dense photo coverage. It never became a general-purpose model.

Why this matters now

Atlas attacks the same problem, novel view synthesis from sparse images, but replaces per-scene training with a pretrained foundation model that reconstructs from two or three photos.

February 2024

Sora, OpenAI's video model (2024)

OpenAI demonstrated Sora, a text-to-video model that generated up to a minute of footage with unusually consistent objects and motion. OpenAI's technical report described such models as 'world simulators.'

Then

Sora triggered a wave of video-generation models from Google, Meta, and startups, though physical consistency and controllability remained unreliable.

Now

Video models proved they could learn physical regularities from pixels, but they had no explicit 3D structure and only coarse camera control.

Why this matters now

Atlas is built to close those gaps: it takes camera poses as precise geometric inputs and outputs explicit 3D geometry alongside pixels.

February 2024

Genie, DeepMind's generative world model (2024)

DeepMind unveiled Genie, a model trained on unlabeled internet videos that could turn a single image into a playable 2D platformer environment. It established 'generative interactive environments' as a research category.

Then

Genie set off a race to world models that simulate rather than just render, drawing in OpenAI, World Labs, and others.

Now

By 2026 that race includes Google DeepMind's Genie 3, which World Labs now faces as a direct competitor for interactive 3D worlds.

Why this matters now

Atlas extends Genie's core idea, generating a world from a single image, into 3D with explicit geometry, reconstruction, and robot sensor simulation.

Sources

(8)