Mark RadarMARK RADAR
About
EN
Sign in

DeepSeek Formally Launches DeepSeek-V4 Preview Models

5 reports · First detected 2026-04-24 · Last active 2026-04-28

DeepSeek shook the global AI market in early 2025 with its low-cost R1 model, prompting investors to reassess heavy spending on computing power. V4 is its next-generation flagship, with a greater focus on long-context processing, reasoning and autonomous agent capabilities. Its ability to close the gap with OpenAI, Anthropic and Google through an open-source, low-price strategy could shape the competitive position of Chinese models.

DeepSeek formally launched and open-sourced the V4 preview on April 24, 2026, offering the 1.6-trillion-parameter Pro and 284-billion-parameter Flash models. Both support a 1-million-token context window and thinking and non-thinking modes. For every 1 million uncached tokens, API input/output prices are 1/2 yuan for Flash and 3/6 yuan for Pro.

All Coverage

5 original reports
NEWS.SMOL.AI 2026-04-24
DeepSeek v4

The Backstory

The history behind this event
DeepSeek Quadruples V4 API Prices During Peak Hours2026-08-14 · 2 reports · similarity 0.82

DeepSeek has built its position in artificial intelligence by offering large-language-model APIs at prices below those of major rivals. V4 is the company’s flagship model, with usage billed by the million tokens. The introduction of peak-hour pricing signals a shift toward managing computing demand through variable rates, an approach that could influence when developers run workloads. Even after the increase, V4 remains relatively inexpensive compared with services from competitors including Anthropic.

DeepSeek said the new V4 API pricing will take effect on Aug. 17, 2026, with peak-hour charges rising to four times their current level. The reports did not disclose the specific per-million-token rates before or after the adjustment. The change represents a significant increase for developers running inference workloads during periods of heavy demand, although DeepSeek’s absolute pricing is expected to remain below comparable offerings from Anthropic and other leading AI providers.

DeepSeek Upgrades V4 Pro, Nears Claude Performance at Fraction of Cost2026-08-13 · 3 reports · similarity 0.86

China’s DeepSeek is intensifying competition in generative AI with DeepSeek-V4-Pro, a flagship model aimed at reasoning, coding and general-purpose workloads. The model matters because performance approaching Anthropic’s premium Claude offering at a sharply lower cost could accelerate enterprise adoption and pressure global developers to reduce inference prices. It also extends DeepSeek’s strategy of challenging better-funded U.S. rivals through efficiency rather than spending alone.

DeepSeek quietly upgraded DeepSeek-V4-Pro to its production-ready “0813” release on Aug. 13, 2026. Reported benchmark results showed the model trailing Anthropic’s Claude Fable 5 by about 5.3%, while costing roughly 46 times less; expressed another way, the rival model was about 4,500% more expensive. The reports did not disclose exact U.S.-dollar rates per million tokens for either model, leaving the pricing comparison dependent on the cited test configuration.

DeepSeek Upgrades V4-Flash to Boost Agentic Coding2026-08-03 · 3 reports · similarity 0.87

DeepSeek, a Chinese AI startup, has built its challenge to U.S. model makers around open weights and aggressive pricing. V4-Flash uses a mixture-of-experts design with 284 billion total parameters, 13 billion activated per request and a 1 million-token context window. In May 2026, the U.S. National Institute of Standards and Technology’s Center for AI Standards and Innovation said V4 Pro cost less than GPT-5.4 mini on five of seven comparable benchmarks, underscoring why the series matters to developers and enterprise buyers.

On July 31, 2026, DeepSeek released DeepSeek-V4-Flash-0731 on Hugging Face and opened its API for public beta testing after re-running post-training while retaining the 284-billion-parameter architecture and MIT license. DeepSeek reported scores of 82.7 on Terminal Bench 2.1 and 54.4 on DeepSWE. API pricing is $0.14 per million uncached input tokens, $0.0028 for cached input and $0.28 for output, while the open weights allow companies to deploy the model on their own infrastructure.

DeepSeek V4-Flash Launches on Ollama Cloud With Support for Major AI Development Tools2026-04-27 · 2 reports · similarity 0.81

Chinese AI startup DeepSeek released a preview of its V4 model series on April 24, 2026. V4-Flash uses a mixture-of-experts architecture with 284 billion total parameters, activating only 13 billion at a time, and supports a 1-million-token context window. The design uses fewer active parameters to process code and lengthy documents, reducing the computing and deployment barriers for ultra-long-context inference.

As of July 20, 2026, V4-Flash was available on Ollama Cloud. Developers can use ollama launch to connect it to Claude Code, OpenClaw, Codex and OpenCode with a single command. The cloud version provides a 1-million-token context window, while Ollama has yet to disclose standalone pricing for the model.

DeepSeek Reportedly Breaks with Practice, Withholds Latest V4 AI Model from Nvidia and AMD2026-04-24 · 10 reports · similarity 0.84

Chinese AI startup DeepSeek has drawn attention for its low-cost large language models. Pre-release optimization across different chips directly affects inference performance and hardware adoption. Developers have typically provided models to Nvidia and AMD before launch, but DeepSeek’s reported shift toward Huawei could advance China’s efforts to reduce its reliance on advanced U.S. chips.

People familiar with the matter said DeepSeek did not give Nvidia or AMD access to a pre-release version of its flagship V4 model, instead prioritizing access for Chinese suppliers including Huawei. V4 is reportedly scheduled for release in late April, will support a one-million-token context window and could run on Huawei’s Ascend 950PR. Separate reports claim its cost is about one-hundredth that of GPT.

Mark Radar|MARK RADAR

If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →

All times are in Taipei time (GMT+8)