Mark RadarMARK RADAR
About
EN
Sign in

OpenAI, Google Turn AI Speed Into a Selling Point

1 reports · First detected 2026-08-15 · Last active 2026-08-15

Generative AI competition has largely centered on model capability and price, but latency is becoming a commercial differentiator as companies deploy customer-service bots, coding agents and financial-analysis tools. Faster inference lets software complete more steps while an event is still unfolding, improving productivity and user engagement. OpenAI underscored the shift on Jan. 14 with a multiyear Cerebras agreement for 750 megawatts of low-latency computing capacity, a deal reported at more than $10 billion and scheduled to roll out in stages through 2028.

On Aug. 13, OpenAI previewed Ultrafast, an API service tier that runs GPT-5.6 Sol on Cerebras chips at as many as 750 output tokens a second, up to 14 times faster than Standard processing. Access is initially limited to selected customers and pricing has not been disclosed. Google released Gemini 3.7 Flash the same day, charging an introductory $0.75 per million input tokens and $3.75 per million output tokens through Dec. 31; those rates double on Jan. 1, 2027.

All Coverage

1 original reports

The Backstory

The history behind this event

No historical echoes for this signal

Mark Radar|MARK RADAR

If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →

All times are in Taipei time (GMT+8)