Can AI agents handle real freelance work online? A new study says not yet

October 31, 2025  13:25

Are your remote freelancers asking for raises as inflation soars? Some companies might be tempted to replace them with AI agents — but, according to new research, that’s likely to end in failure.

As reported by Wired and Futurism, a recent study reveals how ineffective current AI models are at automating tasks — let alone entire professions — compared to the human workers they are meant to replace. The tests were conducted by researchers from the nonprofit Center for AI Safety (CAIS) and Scale AI, a major data annotation company whose freelancers perform much of the routine work underpinning the AI industry. The teams assigned six leading AI agents a set of simulated freelance jobs.

The results were discouraging. None of the AI agents successfully completed more than 3% of the assigned tasks, collectively earning only $1,810 out of a possible $143,991. Dan Hendrycks, director of CAIS, told Wired that the study was intended to provide a clearer picture of the real capabilities of AI systems.

A New Benchmark: Remote Labor Index

To measure AI performance, the researchers created a new benchmark called the Remote Labor Index (RLI), designed to assess how well AI agents can complete economically valuable work in fields ranging from game development to data analysis.

The benchmark includes 240 real-world remote projects drawn from freelance platforms, covering 23 domains. Each project represents a complete, self-contained task that typically requires about 11.5 hours of human labor and is valued at $200.

Unlike tests that evaluate partial assistance, RLI assesses whether an AI can complete the entire project from start to finish at a quality acceptable for real-world delivery. Expert human evaluators manually reviewed each AI submission against a “gold standard” human output, with annotator agreement exceeding 90%.

The best performer was Manus, an AI agent from a Chinese startup, which managed to automate just 2.5% of the projects. In second place, with 2.1%, were Grok 4 (developed by Elon Musk’s team) and Claude Sonnet 4.5 from Anthropic — the latter marketed as “the world’s best coding model” and “the most advanced for building complex agents.”

Next came OpenAI’s GPT-5, which achieved 1.7% completion despite being described as having “PhD-level intelligence.” OpenAI’s CEO Sam Altman had claimed that GPT-5 marked “a significant step toward AGI” (artificial general intelligence) capable of outperforming humans in most economically valuable work — yet the RLI results suggest otherwise.

Ironically, OpenAI’s own specialized ChatGPT Agent ranked second to last with 1.3%, while Google’s Gemini 2.5 Pro came in last with a dismal 0.8%.

Why AI Still Can’t Replace Human Workers

The push to sell AI agents as a replacement for employees has become an industry obsession, with companies like OpenAI seeking to monetize the popularity of chatbots — many of which are free to use. However, despite executive enthusiasm and layoffs driven by automation, it remains unclear whether AI actually improves productivity.

Bing Liu, director of research at Scale AI, told Wired that discussions about AI and jobs have largely been hypothetical until now.

Anecdotally, many managers who replaced staff with AI ended up rehiring them after realizing the tools were inadequate. Studies confirm this pattern: research from MIT found that 95% of companies testing AI initiatives saw no revenue growth, while another study showed that AI integration often results in an influx of low-quality “workslop” — content that requires extensive editing, slowing workflows and increasing tension among colleagues.

Hendrycks noted several fundamental flaws in AI agents: they lack long-term memory, cannot continuously learn from experience, and do not acquire on-the-job skills as humans do.

Despite these shortcomings, AI-driven layoffs continue across industries.

In Short

The joint CAIS and Scale AI study using the Remote Labor Index found that leading AI agents — including Grok 4, Claude 4.5, and GPT-5 — successfully completed fewer than 3% of real freelance projects, earning only a fraction of what human workers did. Manus led with 2.5%, while Google’s Gemini trailed with 0.8%.

For now, AI remains unfit for true work automation: it lacks memory, learning, and consistent quality. The hype around replacing employees with AI may be premature, but the measurable progress shows that competition — and the stakes — continue to rise.


 
 
 
 
  • Archive