Anthropic’s TASTE Benchmark Exposes Weak AI Research Judgment
Anthropic has introduced TASTE, a benchmark designed to test whether artificial intelligence models can evaluate research proposals and make judgments similar to those of experienced researchers. The assessment focuses on higher-order decisions such as identifying important problems and setting research priorities. Those skills are crucial for allocating resources in AI safety research and differ from the language generation and idea production tasks at which leading models often excel.
Results released by Anthropic show that the strongest model agreed with the expert benchmark only about 60% of the time, while most models performed near the level of random guessing. The findings suggest that current systems may generate plausible research ideas without reliably determining which questions deserve attention first. Human oversight therefore remains essential if AI is used to review proposals, direct research programs or shape safety priorities.
All Coverage
1 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.
If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →