Mark RadarMARK RADAR
About
EN
Sign in
Event File AI Anthropic

Anthropic’s TASTE Benchmark Exposes Weak AI Research Judgment

1 reports · First detected 2026-08-31 · Last active 2026-08-31

Anthropic has introduced TASTE, a benchmark designed to test whether artificial intelligence models can evaluate research proposals and make judgments similar to those of experienced researchers. The assessment focuses on higher-order decisions such as identifying important problems and setting research priorities. Those skills are crucial for allocating resources in AI safety research and differ from the language generation and idea production tasks at which leading models often excel.

Results released by Anthropic show that the strongest model agreed with the expert benchmark only about 60% of the time, while most models performed near the level of random guessing. The findings suggest that current systems may generate plausible research ideas without reliably determining which questions deserve attention first. Human oversight therefore remains essential if AI is used to review proposals, direct research programs or shape safety priorities.

All Coverage

1 original reports

The Backstory

The history behind this event

No historical echoes for this signal

Mark Radar|MARK RADAR

If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →

All times are in Taipei time (GMT+8)