Mark RadarMARK RADAR
About
EN
Sign in
Event File AI Agentic AI

Keenable AI Open-Sources Live Search Benchmark NEEDLE

1 reports · First detected 2026-09-01 · Last active 2026-09-01

Static search benchmarks can be compromised when agents fetch public answer keys or when models already retain the answers in their parameters, making strong scores a poor proxy for live retrieval. Keenable AI designed NEEDLE — News, Everyday, Expert, Deep-tail and Legal Evaluation — as a continuously refreshed, open-source test of agentic search. Its common protocol is intended to reduce leakage and overfitting while separating ranking failures from cases where no provider retrieved strong evidence.

Keenable AI released NEEDLE under an MIT license on Aug. 27. News queries are rebuilt every hour from about 124 curated RSS feeds and Google Trends, while finance, scholar, deep-tail and legal suites refresh daily. The framework tests 15 search APIs with identical queries and a 2,000-character evidence cap. In seven-day averages through Aug. 28, Exa scored 0.910 on finance, Keenable 0.872 and Perplexity 0.871, against the pooled “ultimate” ceiling of 0.965.

All Coverage

1 original reports

The Backstory

The history behind this event

No historical echoes for this signal

Mark Radar|MARK RADAR

If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →

All times are in Taipei time (GMT+8)