…Study Finds Most AI Agents Successfully Complete Only About One-Quarter of Work Assignments

BERKELEY, Calif. — With credit unions racing to embrace AI as fast as everyone else, a new analysis has found artificial intelligence agents remain far from replacing workers across most knowledge-based professions, successfully completing only about one-quarter of real-world work assignments in a new benchmark developed by researchers at the University of California, Berkeley’s Center for Responsible, Decentralized Intelligence.

The findings, based on nearly 1,500 work assignments collected from professionals across dozens of industries, challenge predictions that AI agents will surpass humans in most knowledge jobs within the next one to two years, according to the research center.

“There are predictions everywhere that AI agents will surpass humans in almost all jobs between 2026 and 2027,” Yiyou Sun, a core author of the benchmark, told FrontierNews.ai. “So, we created this exam to verify this claim.”

Researchers assembled 1,490 real-world assignments from more than 250 professionals representing 55 industries. The tasks reflected projects that typically require hours to weeks to complete, including:

  • Preparing legal filings.
  • Building financial models.
  • Designing manufacturing parts.

How Evaluations Worked

Each AI system was evaluated on whether it produced a complete, accurate final product. Researchers awarded no partial credit, meaning assignments with missing or incorrect deliverables were considered failures, according to the University of California, Berkeley’s Center for Responsible, Decentralized Intelligence

According to the benchmark, the highest-performing system was OpenAI’s Codex running on the GPT-5.5 model, which correctly completed 26.2% of all assignments.

Performance dropped sharply on more complex projects requiring multiple steps and sustained reasoning.

Among the findings:

  • AI systems averaged a 2.6% success rate on the most difficult assignments.
  • OpenAI’s Codex achieved an 8.6% success rate on those complex tasks.
  • Anthropic’s Claude Code, running on the Opus 4.7 model, failed every assignment in the most difficult category.

How One Experiment Worked

Researchers also tested whether substantially increasing computing resources improved results. In one experiment, an AI system consumed approximately $630 in computing resources and processed 763 million tokens—the units AI models use to process information—but still completed only 2.9% of the most difficult assignments successfully.

The researchers concluded that simply adding more computing power does not reliably improve an AI agent’s ability to perform complex, multi-step work.

The Better the Definition, Better the Results

The study found AI agents performed significantly better on simpler, well-defined tasks.

Among 59 assignments categorized as more manageable for current AI systems:

  • Top-performing models successfully completed about 30% of the tasks.
  • OpenAI’s Codex running GPT-5.5 achieved a 42.4% success rate.

According to the researchers, current AI systems often perform well on isolated questions but struggle when required to execute an entire workflow, including gathering information, verifying results, correcting errors and producing a finished product suitable for professional use.

Facebook
Twitter
LinkedIn

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.