OpenAI, Google Turn AI Speed Into a Selling Point
Generative AI competition has largely centered on model capability and price, but latency is becoming a commercial differentiator as companies deploy customer-service bots, coding agents and financial-analysis tools. Faster inference lets software complete more steps while an event is still unfolding, improving productivity and user engagement. OpenAI underscored the shift on Jan. 14 with a multiyear Cerebras agreement for 750 megawatts of low-latency computing capacity, a deal reported at more than $10 billion and scheduled to roll out in stages through 2028.
On Aug. 13, OpenAI previewed Ultrafast, an API service tier that runs GPT-5.6 Sol on Cerebras chips at as many as 750 output tokens a second, up to 14 times faster than Standard processing. Access is initially limited to selected customers and pricing has not been disclosed. Google released Gemini 3.7 Flash the same day, charging an introductory $0.75 per million input tokens and $3.75 per million output tokens through Dec. 31; those rates double on Jan. 1, 2027.
All Coverage
1 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.
If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →