Musk Freezes Grok Updates Over Baldur’s Gate 3 Errors; Accuracy Later Hits 92%
xAI’s Grok competes with OpenAI’s ChatGPT and Anthropic’s Claude in general-purpose AI capabilities. Baldur’s Gate 3, developed by Larian Studios, has a complex story and rules despite the wealth of publicly available guides. Grok still got details wrong, making the episode a test of model knowledge reliability and xAI’s approach to product decisions.
In 2025, Elon Musk delayed xAI model updates for several days and reassigned several senior engineers after Grok answered a Baldur’s Gate 3 question incorrectly. In February 2026, BaldurBench compared Grok, ChatGPT, Claude and Gemini using five questions. The optimized Grok achieved 92% accuracy, with answers tending to use tables and gamer terminology.
All Coverage
1 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.
If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →