Mark RadarMARK RADAR
About
EN
Sign in

DeepSeek Slashes API Cache-Hit Pricing to One-Tenth, Extends V4 Pro Discount

5 reports · First detected 2026-04-27 · Last active 2026-05-25

Chinese AI company DeepSeek has long challenged OpenAI and Anthropic with low-cost APIs. Input caching allows existing content to be reused, directly affecting the operating costs of services with high volumes of repetitive requests, including RAG and agents. Developers are therefore watching the price changes closely.

DeepSeek most recently announced that input cache-hit API rates across its entire model lineup would fall to one-tenth of their previous levels. It also extended the flagship V4 Pro model’s limited-time 75% discount through May 5, 2026, bringing the discounted API output price to less than NT$30 per million tokens.

All Coverage

5 original reports

The Backstory

The history behind this event
DeepSeek Quadruples V4 API Prices During Peak Hours2026-08-14 · 2 reports · similarity 0.83

DeepSeek has built its position in artificial intelligence by offering large-language-model APIs at prices below those of major rivals. V4 is the company’s flagship model, with usage billed by the million tokens. The introduction of peak-hour pricing signals a shift toward managing computing demand through variable rates, an approach that could influence when developers run workloads. Even after the increase, V4 remains relatively inexpensive compared with services from competitors including Anthropic.

DeepSeek said the new V4 API pricing will take effect on Aug. 17, 2026, with peak-hour charges rising to four times their current level. The reports did not disclose the specific per-million-token rates before or after the adjustment. The change represents a significant increase for developers running inference workloads during periods of heavy demand, although DeepSeek’s absolute pricing is expected to remain below comparable offerings from Anthropic and other leading AI providers.

DeepSeek Upgrades V4 Pro, Nears Claude Performance at Fraction of Cost2026-08-13 · 3 reports · similarity 0.82

China’s DeepSeek is intensifying competition in generative AI with DeepSeek-V4-Pro, a flagship model aimed at reasoning, coding and general-purpose workloads. The model matters because performance approaching Anthropic’s premium Claude offering at a sharply lower cost could accelerate enterprise adoption and pressure global developers to reduce inference prices. It also extends DeepSeek’s strategy of challenging better-funded U.S. rivals through efficiency rather than spending alone.

DeepSeek quietly upgraded DeepSeek-V4-Pro to its production-ready “0813” release on Aug. 13, 2026. Reported benchmark results showed the model trailing Anthropic’s Claude Fable 5 by about 5.3%, while costing roughly 46 times less; expressed another way, the rival model was about 4,500% more expensive. The reports did not disclose exact U.S.-dollar rates per million tokens for either model, leaving the pricing comparison dependent on the cited test configuration.

Mark Radar|MARK RADAR

If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →

All times are in Taipei time (GMT+8)