AI Labs Shift Model Race to Token Pricing
The artificial-intelligence model race is shifting from raw benchmark performance toward the cost of running models at scale. Stanford University’s AI Index found that inference prices for systems delivering roughly GPT-3.5-level capability have collapsed, while Andreessen Horowitz, or a16z, describes the trend as “LLMflation.” The comparison with broadband’s falling costs matters because cheaper access can unlock new products, higher usage and broader adoption beyond companies able to absorb large computing bills.
According to the AI Index, the cost of querying a model at GPT-3.5-equivalent performance fell from about $20 per million tokens in November 2022 to roughly $0.07 in October 2024, a decline of more than 280-fold. OpenAI, Google and Anthropic are increasingly competing through lower API prices alongside speed and model quality. The reductions are making unit economics and deployment efficiency more important in purchasing decisions as advanced capabilities become less differentiated.
All Coverage
1 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.
If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →