DeepSeek Upgrades V4 Pro, Narrows Claude Gap at Fraction of Cost
China’s DeepSeek has built its challenge to leading US artificial-intelligence developers around models that combine competitive performance with sharply lower inference costs. Its flagship DeepSeek-V4-Pro is aimed at users seeking capabilities close to top proprietary systems without comparable spending, a strategy that could intensify pricing pressure across the generative AI market and broaden enterprise adoption.
DeepSeek quietly upgraded DeepSeek-V4-Pro to its 0813 production release on Aug. 13. Benchmark results cited in the report showed Anthropic’s Claude Fable 5 ahead by about 5.3%, while costing roughly 4,500% more. That implies DeepSeek’s model is priced at about one forty-fifth of its rival, highlighting a large cost gap despite a relatively narrow difference in measured performance.
All Coverage
1 original reportsThe Backstory
The history behind this eventDeepSeek Upgrades V4-Flash to Boost Agentic Coding
DeepSeek, a Chinese AI startup, has built its challenge to U.S. model makers around open weights and aggressive pricing. V4-Flash uses a mixture-of-experts design with 284 billion total parameters, 13 billion activated per request and a 1 million-token context window. In May 2026, the U.S. National Institute of Standards and Technology’s Center for AI Standards and Innovation said V4 Pro cost less than GPT-5.4 mini on five of seven comparable benchmarks, underscoring why the series matters to developers and enterprise buyers.
On July 31, 2026, DeepSeek released DeepSeek-V4-Flash-0731 on Hugging Face and opened its API for public beta testing after re-running post-training while retaining the 284-billion-parameter architecture and MIT license. DeepSeek reported scores of 82.7 on Terminal Bench 2.1 and 54.4 on DeepSWE. API pricing is $0.14 per million uncached input tokens, $0.0028 for cached input and $0.28 for output, while the open weights allow companies to deploy the model on their own infrastructure.
DeepSeek Slashes API Cache-Hit Pricing to One-Tenth, Extends V4 Pro Discount
Chinese AI company DeepSeek has long challenged OpenAI and Anthropic with low-cost APIs. Input caching allows existing content to be reused, directly affecting the operating costs of services with high volumes of repetitive requests, including RAG and agents. Developers are therefore watching the price changes closely.
DeepSeek most recently announced that input cache-hit API rates across its entire model lineup would fall to one-tenth of their previous levels. It also extended the flagship V4 Pro model’s limited-time 75% discount through May 5, 2026, bringing the discounted API output price to less than NT$30 per million tokens.
DeepSeek Formally Launches DeepSeek-V4 Preview Models
DeepSeek shook the global AI market in early 2025 with its low-cost R1 model, prompting investors to reassess heavy spending on computing power. V4 is its next-generation flagship, with a greater focus on long-context processing, reasoning and autonomous agent capabilities. Its ability to close the gap with OpenAI, Anthropic and Google through an open-source, low-price strategy could shape the competitive position of Chinese models.
DeepSeek formally launched and open-sourced the V4 preview on April 24, 2026, offering the 1.6-trillion-parameter Pro and 284-billion-parameter Flash models. Both support a 1-million-token context window and thinking and non-thinking modes. For every 1 million uncached tokens, API input/output prices are 1/2 yuan for Flash and 3/6 yuan for Pro.
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.