Mark RadarMARK RADAR
About
EN
Sign in

AI Labs Shift Model Race to Token Pricing

1 reports · First detected 2026-08-18 · Last active 2026-08-18

The artificial-intelligence model race is shifting from raw benchmark performance toward the cost of running models at scale. Stanford University’s AI Index found that inference prices for systems delivering roughly GPT-3.5-level capability have collapsed, while Andreessen Horowitz, or a16z, describes the trend as “LLMflation.” The comparison with broadband’s falling costs matters because cheaper access can unlock new products, higher usage and broader adoption beyond companies able to absorb large computing bills.

According to the AI Index, the cost of querying a model at GPT-3.5-equivalent performance fell from about $20 per million tokens in November 2022 to roughly $0.07 in October 2024, a decline of more than 280-fold. OpenAI, Google and Anthropic are increasingly competing through lower API prices alongside speed and model quality. The reductions are making unit economics and deployment efficiency more important in purchasing decisions as advanced capabilities become less differentiated.

All Coverage

1 original reports

The Backstory

The history behind this event

No historical echoes for this signal

Mark Radar|MARK RADAR

If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →

All times are in Taipei time (GMT+8)