AI agents quadruple their score on a real freelance-work benchmark
New CapabilitiesThe Remote Labor Index jumps from 2.5% to 16.1% of paid projects completed at professional quality in under eight months
July 6th, 2026: Frontier score more than quadruplesNew here? Follow stories to track developments over time. Create a free account to get updates when stories you care about change.
Overview
Updated Jul 9Eight months ago, the best AI agent could finish 2.5% of real freelance jobs well enough for a paying client to accept. On July 6, the Center for AI Safety reported that a new model now clears 16.1%.
The number is small, but the slope is steep. The test uses 240 real paid projects, and human professionals earned about $144,000 doing them. AI agents are still failing most of the work. They are failing less of it, fast.
Why it matters
If AI agents keep closing this gap at the current pace, freelance work in design, video, and data analysis is the first labor market to feel it.
Questions about this story
Free account needed to ask — your question is kept and asked for you right after sign-up. Answers are public.
No questions yet — be the first to ask.
Key Indicators
Voices
Curated perspectives — historical figures and your fellow readers.
Play
Exploring all sides of a story is often best achieved with Play.
Higher or Lower
A number from this story, against one from elsewhere in the news — guess which is bigger, then keep the chain going. 5 rounds, 3 strikes; a miss costs a strike and resets your streak.
Keyboard: ↓/L lower · ↑/H higher
0 points — sign up to put that on the leaderboard.
Connections
Sixteen names from the news. Find the four hidden groups of four. Four mistakes max.
Sign up to keep a daily streak — a new puzzle lands every day.
Exit debate?
Your progress in this debate will be lost.
- 1 Two AI personas square off on this story.
- 2 You predict who'll win each round — correct picks earn XP.
- 3 One crossfire question is yours to fire. Pick it carefully.
Couldn't generate a topic
Select Your Champions
Choose one persona for each side of the debate
DEBATE TOPIC
Choose personas with different perspectives for a more dynamic debate.
Select debater for this side:
No debate personas available right now.
Select debater for this side:
No debate personas available right now.
Who's Got This Round?
Make your prediction before the referee scores
The referee scores both sides on
Round Results
Set the Crossfire
Pick the question both personas must answer in the final round
Debate Oracle! You called every round!
Sharp Instincts! You know your debaters!
The Coin Flip Strategist! Perfectly balanced!
The Contrarian! Bold predictions!
Inverse Genius! Try betting the opposite next time!
XP Breakdown
Prediction History
People Involved
Organizations Involved
A San Francisco nonprofit that studies risks from advanced AI and builds tests to measure what AI systems can do.
A company that supplies data and evaluation services to AI developers, providing the human graders behind the benchmark.
Timeline
October 2025 July 2026
-
Frontier score more than quadruples
Latest ResearchNew results show Fable 5 at 16.1%, Opus 4.8 at 8.3%, and GPT-5.5 at 6.3%. The top score has more than quadrupled from the previous leader's 4.2% in under eight months.
-
Benchmark launches with a 2.5% ceiling
ResearchThe Center for AI Safety and Scale AI release the Remote Labor Index. The best agent, Manus, completes 2.5% of the 240 paid projects at professional quality.
Historical Context
3 moments from history that rhyme with this story — and how they unfolded.
ImageNet and the deep-learning breakout (2012)
A neural network called AlexNet cut the error rate on the ImageNet image-recognition contest by roughly 10 percentage points in one year. The benchmark had moved slowly before. Then it moved fast.
Research funding and talent poured into deep learning within months.
The approach became the base for modern computer vision, from phone cameras to self-driving car perception.
A single benchmark can flip from slow crawl to steep climb. The Remote Labor Index's jump from 2.5% to 16.1% has the same early shape, though the outcome is not yet settled.
Self-driving 90% problem (2015 onward)
Autonomous-driving demos quickly handled most road situations, and companies predicted robotaxis within a few years. The last fraction of hard cases proved far slower to solve.
Billions in investment chased a near-term rollout that kept slipping.
Deployment came, but narrowly, city by city, more than a decade after the early hype.
Fast early gains do not guarantee a fast finish. The messy final projects on the Remote Labor Index could resist automation the way rare road cases did.
DeepMind's protein-folding leap on CASP (2020)
AlphaFold scored high enough on the CASP protein-structure contest that organizers called the 50-year problem largely solved. Prior systems had plateaued for years.
Structural biology labs began using predicted structures almost immediately.
DeepMind released structures for most known proteins, speeding drug and disease research worldwide.
A trusted, hard-to-game benchmark turned a research milestone into real-world use. The Remote Labor Index aims for the same credibility by grading against paid human work.
