Mark RadarMARK RADAR
EN
Event File AI Agentic AI

Berkeley Test Shows AI Agents Fail Three in Four Real-World Tasks

1 reports · First detected 2026-07-25 · Last active 2026-07-25

AI agents are emerging as a central bet in corporate efforts to automate knowledge work, but answering questions is not the same as delivering finished work. UC Berkeley’s Center for Responsible, Decentralized Intelligence created Agents’ Last Exam, or ALE, to assess complete professional workflows rather than narrow academic exercises. The benchmark matters because it tests claims that autonomous systems could replace workers across many knowledge-intensive jobs as soon as 2026 or 2027.

Berkeley RDI announced ALE on June 14, 2026, and a July 24 report detailed results from 1,490 assignments supplied by more than 250 professionals across 55 industries. OpenAI’s Codex running GPT-5.5 led the systems tested but completed only 26.2% of assignments, implying a failure rate of nearly 74%. Average success on the hardest tier was 2.6%, while one compute-heavy run consumed $630 and 763 million tokens yet passed just 2.9% of tasks.

All Coverage

1 original reports

The Backstory

The history behind this event

No historical echoes for this signal

Mark Radar|MARK RADAR
All times are in Taipei time (GMT+8)