General-Purpose LLMs Beat Specialized Clinical AI in Nature Study
Clinical AI vendors have long argued that domain training and retrieval-augmented generation can make specialized tools more reliable than general-purpose models. The Nature Medicine study directly tested that proposition, comparing OpenEvidence and UpToDate Expert AI with frontier systems from OpenAI, Google and Anthropic. The result matters for hospitals weighing costly clinical products and for vendors whose premium pricing rests on specialized performance. Still, benchmark superiority is not the same as improved patient outcomes, and the findings should not be read as clinical deployment clearance.
Published on June 12, 2026, the study tested GPT-5.2, Gemini 3.1 Pro Preview and Claude Opus 4.6 against OpenEvidence and UpToDate Expert AI. On 500 MedQA questions, Gemini led with 97.4% accuracy, followed by GPT at 94.2% and Claude at 90.2%; OpenEvidence and UpToDate scored 89.6% and 88.4%. The researchers also evaluated 500 HealthBench items and 100 de-identified, real-world physician queries. Twelve U.S. clinicians conducted randomized, blinded reviews, generating 1,800 model-question annotations, and placed the three frontier models in the top performance tier across the real-query test.
All Coverage
1 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.
If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →