GPT-3 (2020)
OpenAI showed that language model performance scales predictably with data and parameters. GPT-3, with 175 billion parameters trained on internet text, could perform tasks it was never explicitly trained for.
Pretraining on internet-scale data became the default approach in natural language processing.
Established the scaling law framework that now guides most large AI model development.
Light-O1 applies the same scaling logic to physical behavior, using internet video instead of text as the pretraining data.
