Berkeley Test Shows AI Agents Fail Three in Four Real-World Tasks
AI agents are emerging as a central bet in corporate efforts to automate knowledge work, but answering questions is not the same as delivering finished work. UC Berkeley’s Center for Responsible, Decentralized Intelligence created Agents’ Last Exam, or ALE, to assess complete professional workflows rather than narrow academic exercises. The benchmark matters because it tests claims that autonomous systems could replace workers across many knowledge-intensive jobs as soon as 2026 or 2027.
Berkeley RDI announced ALE on June 14, 2026, and a July 24 report detailed results from 1,490 assignments supplied by more than 250 professionals across 55 industries. OpenAI’s Codex running GPT-5.5 led the systems tested but completed only 26.2% of assignments, implying a failure rate of nearly 74%. Average success on the hardest tier was 2.6%, while one compute-heavy run consumed $630 and 763 million tokens yet passed just 2.9% of tasks.
All Coverage
1 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.