Pull to refresh
Logo
Light Origins open-sources model that turns internet video into robot skills

Light Origins open-sources model that turns internet video into robot skills

New Capabilities

Light-O1-Preview, a 6B-parameter model, learns whole-body actions from human video and adapts them to different robot bodies

Today: Light-O1 launched, preview open-sourced

Overview

Updated 53 minutes ago

Robot training data is scarce and expensive. Light Origins thinks internet video can replace most of it. On September 21, the Shenzhen startup released Light-O1, a model that learns whole-body human actions from online video and adapts them to different humanoid robots.

The 6-billion-parameter model, open-sourced under Apache 2.0, is the clearest test yet of whether the scaling playbook that worked for language models can work for physical behavior. Light Origins reports power-law improvements in action prediction as pretraining data grows, up to 100,000 hours of recovered human action.

Why it matters

If internet video can pretrain robot skills, humanoid development stops being bottlenecked by expensive, hand-collected robot data.

Questions about this story

Free account needed to ask — your question is kept and asked for you right after sign-up. Answers are public.

No questions yet — be the first to ask.

Key Indicators

6B
Model parameters
Light-O1-Preview has 6 billion parameters, traced to Qwen3.5-4B-Base.
100,000
Hours of recovered human action
Largest pretraining budget tested, corresponding to 120 billion multimodal tokens.
79.3%
Macro success rate on RoboCasa GR-1 tasks
Measured across 24 simulated kitchen tasks with 50 evaluation episodes each.
138
Values per frame in action representation
Covers root movement, pelvis height, yaw, 22 joints, and hand-open states at 20 fps.

Voices

Curated perspectives — historical figures and your fellow readers.

Ever wondered what historical figures would say about today's headlines?

Sign up to generate historical perspectives on this story.

People Involved

Organizations Involved

Timeline

December 2024 September 2026

3 events Latest: Today
  1. Light-O1 launched, preview open-sourced

    Today Product Launch

    Light Origins releases Light-O1 and open-sources Light-O1-Preview on Hugging Face under Apache 2.0.

  2. Pre-A round announced

    Funding

    Light Origins announces a Pre-A round of several hundred million yuan, led by CAS Investment.

  3. Light Origins founded

    Founding

    Roger Jiang founds Light Origins in Shenzhen after leaving OpenAI.

Scenarios

1

Internet-video pretraining becomes the standard for robot learning

Possible Resolves by Sep 21, 2027

Discussed by: Light Origins technical report; robotics researchers presenting at CoRL and ICRA

Light Origins' scaling study shows power-law improvement in action prediction as pretraining data grows. If independent labs replicate this, internet-video pretraining becomes a standard first step in robot learning, before embodiment-specific adaptation.

2

Light-O1-Preview becomes a community standard for motion generation

Possible Resolves by Sep 21, 2027

Discussed by: Open-source robotics community on Hugging Face

The Apache 2.0 model gains traction as a text-to-motion base. Developers build on it for animation, simulation, and robot control research.

3

Full robot-deployment system stays closed, limiting open-source impact

Likely Resolves by Sep 21, 2027

Discussed by: Runtimewire analysis noting the preview requires separate controllers

Light-O1-Preview generates motion but can't control a robot without proprietary adaptation data and controllers. If Light Origins never releases these, the open-source model is a demo, not a deployable system.

Historical Context

3 moments from history that rhyme with this story — and how they unfolded.

May 2020

GPT-3 (2020)

OpenAI showed that language model performance scales predictably with data and parameters. GPT-3, with 175 billion parameters trained on internet text, could perform tasks it was never explicitly trained for.

Then

Pretraining on internet-scale data became the default approach in natural language processing.

Now

Established the scaling law framework that now guides most large AI model development.

Why this matters now

Light-O1 applies the same scaling logic to physical behavior, using internet video instead of text as the pretraining data.

January 2021

CLIP (2021)

OpenAI's CLIP learned visual concepts from 400 million internet text-image pairs, showing that internet-scale data can teach models generalizable concepts without manual labels.

Then

CLIP became a standard building block for vision and multimodal models.

Now

Proved that internet-scale data can substitute for curated, hand-labeled datasets.

Why this matters now

Light-O1's bet is that human video can play the same role for robot motion that text-image pairs played for vision.

July 2023

RT-2 (2023)

Google's Robotics Transformer 2 adapted a vision-language model pretrained on internet data to robot control. It showed that web knowledge could transfer to physical actions, but was limited by the model's understanding of embodiment.

Then

Demonstrated that internet-pretrained models could be fine-tuned for robot control.

Now

Inspired a wave of work on using large pretrained models as robot policy backbones.

Why this matters now

Light-O1 goes further by pretraining specifically on human whole-body motion, not just language and vision.

Sources

(5)