World Labs debuts Atlas, a model that builds and simulates 3D worlds from a few photos
New CapabilitiesThe 'omni' model handles text, images, video, and 3D, with camera poses and depth as native inputs.
Yesterday: Atlas debuts in early accessNew here? Follow stories to track developments over time. Create a free account to get updates when stories you care about change.
Overview
Updated 2 hours agoPoint a phone camera at a room and this model rebuilds the room in 3D, including the parts the camera never saw. On September 1, World Labs, the startup co-founded by Stanford computer vision pioneer Fei-Fei Li, unveiled Atlas.
Atlas is an 'omni' model: one system that handles text, images, video, and 3D, with camera position and depth as native inputs. That lets it control generated video with pixel precision and reconstruct real scenes from as few as two or three photos. World Labs says Atlas beats specialized video and 3D-reconstruction models at their own tasks, though it released no research paper to back the numbers.
Why it matters
If Atlas scales, robotics training, game design, and 3D capture shift from manual pipelines to a single AI that builds worlds from photos.
Questions about this story
Free account needed to ask — your question is kept and asked for you right after sign-up. Answers are public.
No questions yet — be the first to ask.
Key Indicators
Voices
Curated perspectives — historical figures and your fellow readers.
Play
Exploring all sides of a story is often best achieved with Play.
Higher or Lower
A number from this story, against one from elsewhere in the news — guess which is bigger, then keep the chain going. 5 rounds, 3 strikes; a miss costs a strike and resets your streak.
Keyboard: ↓/L lower · ↑/H higher
0 points — sign up to put that on the leaderboard.
Connections
Sixteen names from the news. Find the four hidden groups of four. Four mistakes max.
Sign up to keep a daily streak — a new puzzle lands every day.
Exit debate?
Your progress in this debate will be lost.
- 1 Two AI personas square off on this story.
- 2 You predict who'll win each round — correct picks earn XP.
- 3 One crossfire question is yours to fire. Pick it carefully.
Couldn't generate a topic
Select Your Champions
Choose one persona for each side of the debate
DEBATE TOPIC
Choose personas with different perspectives for a more dynamic debate.
Select debater for this side:
No debate personas available right now.
Select debater for this side:
No debate personas available right now.
Who's Got This Round?
Make your prediction before the referee scores
The referee scores both sides on
Round Results
Set the Crossfire
Pick the question both personas must answer in the final round
Debate Oracle! You called every round!
Sharp Instincts! You know your debaters!
The Coin Flip Strategist! Perfectly balanced!
The Contrarian! Bold predictions!
Inverse Genius! Try betting the opposite next time!
XP Breakdown
Prediction History
People Involved
Organizations Involved
Timeline
February 2024 September 2026
-
Atlas debuts in early access
Latest Product LaunchWorld Labs unveils Atlas, an omni world model, and opens early access to select partners.
-
World Labs founded
FoundingFei-Fei Li and colleagues launch World Labs to pursue spatial intelligence in AI.
Historical Context
3 moments from history that rhyme with this story — and how they unfolded.
Neural Radiance Fields (2020)
UC Berkeley and Google researchers showed that a small neural network could encode a 3D scene from dozens of photos and render it from new camera angles. The technique, called NeRF, produced photorealistic novel views that earlier methods could not match.
NeRF became the default method for 3D view synthesis within months, spawning hundreds of follow-up papers.
Each new scene required training a fresh network, so the method stayed slow and demanded dense photo coverage. It never became a general-purpose model.
Atlas attacks the same problem, novel view synthesis from sparse images, but replaces per-scene training with a pretrained foundation model that reconstructs from two or three photos.
Sora, OpenAI's video model (2024)
OpenAI demonstrated Sora, a text-to-video model that generated up to a minute of footage with unusually consistent objects and motion. OpenAI's technical report described such models as 'world simulators.'
Sora triggered a wave of video-generation models from Google, Meta, and startups, though physical consistency and controllability remained unreliable.
Video models proved they could learn physical regularities from pixels, but they had no explicit 3D structure and only coarse camera control.
Atlas is built to close those gaps: it takes camera poses as precise geometric inputs and outputs explicit 3D geometry alongside pixels.
Genie, DeepMind's generative world model (2024)
DeepMind unveiled Genie, a model trained on unlabeled internet videos that could turn a single image into a playable 2D platformer environment. It established 'generative interactive environments' as a research category.
Genie set off a race to world models that simulate rather than just render, drawing in OpenAI, World Labs, and others.
By 2026 that race includes Google DeepMind's Genie 3, which World Labs now faces as a direct competitor for interactive 3D worlds.
Atlas extends Genie's core idea, generating a world from a single image, into 3D with explicit geometry, reconstruction, and robot sensor simulation.
