DeepSeek Quadruples V4 API Prices During Peak Hours
DeepSeek has built its position in artificial intelligence by offering large-language-model APIs at prices below those of major rivals. V4 is the company’s flagship model, with usage billed by the million tokens. The introduction of peak-hour pricing signals a shift toward managing computing demand through variable rates, an approach that could influence when developers run workloads. Even after the increase, V4 remains relatively inexpensive compared with services from competitors including Anthropic.
DeepSeek said the new V4 API pricing will take effect on Aug. 17, 2026, with peak-hour charges rising to four times their current level. The reports did not disclose the specific per-million-token rates before or after the adjustment. The change represents a significant increase for developers running inference workloads during periods of heavy demand, although DeepSeek’s absolute pricing is expected to remain below comparable offerings from Anthropic and other leading AI providers.
All Coverage
2 original reportsThe Backstory
The history behind this eventDeepSeek Plans Significant API Price Increase
DeepSeek, a Chinese artificial intelligence startup, provides model inference services to developers and businesses through its application programming interface. Pricing currently varies across models including Flash and Pro, as well as the volume of input and output tokens and whether requests receive cache discounts. A broad, substantial increase could raise operating costs for customers building AI products and highlight the trade-off between aggressive pricing and the expense of computing capacity.
DeepSeek said in its latest announcement that it plans to raise API prices across its services in the near term and warned users that the increase is expected to be significant. The company did not disclose an effective date, revised per-million-token charges or a percentage increase. It said the final pricing structure would be detailed in a formal notice, leaving customers to reassess budgets and usage across Flash and Pro models once the new rate card is released.
DeepSeek Slashes API Cache-Hit Pricing to One-Tenth, Extends V4 Pro Discount
Chinese AI company DeepSeek has long challenged OpenAI and Anthropic with low-cost APIs. Input caching allows existing content to be reused, directly affecting the operating costs of services with high volumes of repetitive requests, including RAG and agents. Developers are therefore watching the price changes closely.
DeepSeek most recently announced that input cache-hit API rates across its entire model lineup would fall to one-tenth of their previous levels. It also extended the flagship V4 Pro model’s limited-time 75% discount through May 5, 2026, bringing the discounted API output price to less than NT$30 per million tokens.
DeepSeek Formally Launches DeepSeek-V4 Preview Models
DeepSeek shook the global AI market in early 2025 with its low-cost R1 model, prompting investors to reassess heavy spending on computing power. V4 is its next-generation flagship, with a greater focus on long-context processing, reasoning and autonomous agent capabilities. Its ability to close the gap with OpenAI, Anthropic and Google through an open-source, low-price strategy could shape the competitive position of Chinese models.
DeepSeek formally launched and open-sourced the V4 preview on April 24, 2026, offering the 1.6-trillion-parameter Pro and 284-billion-parameter Flash models. Both support a 1-million-token context window and thinking and non-thinking modes. For every 1 million uncached tokens, API input/output prices are 1/2 yuan for Flash and 3/6 yuan for Pro.
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.