Grok Leads Agent Tests, Gemini Wins on Speed as DeepSeek Raises Prices
Competition among large language models is shifting beyond raw benchmark scores toward a three-way test of intelligence, execution speed and cost. Elon Musk’s xAI, Alphabet’s Google and Hangzhou-based DeepSeek are pitching Grok 4.6, Gemini 3.7 Flash and DeepSeek V4 Pro, respectively, as workhorses for coding and autonomous agents. The comparison matters for businesses because a model that is smarter on paper may still lose if it takes longer, consumes more tokens or costs more to finish a task.
Evaluations published Aug. 12-13 put Grok 4.6 at 70.8% on CursorBench 3.2, while Gemini 3.7 Flash delivered roughly 340 output tokens a second. Google priced Gemini at an introductory $0.75 per million input tokens and $3.75 per million output tokens through Dec. 31, before rates double on Jan. 1, 2027. DeepSeek’s new time-of-day schedule took effect Aug. 17: V4 Pro output costs $1.98 per million tokens off-peak and $3.96 at peak, up from $0.87. The results favor Grok for long-horizon agent work and Gemini for speed and value, while DeepSeek faces criticism over the increase.
All Coverage
1 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.
If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →