Keenable AI Open-Sources Live Search Benchmark NEEDLE
Static search benchmarks can be compromised when agents fetch public answer keys or when models already retain the answers in their parameters, making strong scores a poor proxy for live retrieval. Keenable AI designed NEEDLE — News, Everyday, Expert, Deep-tail and Legal Evaluation — as a continuously refreshed, open-source test of agentic search. Its common protocol is intended to reduce leakage and overfitting while separating ranking failures from cases where no provider retrieved strong evidence.
Keenable AI released NEEDLE under an MIT license on Aug. 27. News queries are rebuilt every hour from about 124 curated RSS feeds and Google Trends, while finance, scholar, deep-tail and legal suites refresh daily. The framework tests 15 search APIs with identical queries and a 2,000-character evidence cap. In seven-day averages through Aug. 28, Exa scored 0.910 on finance, Keenable 0.872 and Perplexity 0.871, against the pooled “ultimate” ceiling of 0.965.
All Coverage
1 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.
If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →