Mark RadarMARK RADAR
About
EN
Sign in

DeepSeek Launches V4.1 Flash, Touting Pro-Beating Performance at Lower Cost

2 reports · First detected 2026-09-10 · Last active 2026-09-10

DeepSeek has built its challenge to larger artificial-intelligence providers around models that combine strong performance with lower operating costs. V4 Pro had served as its premium offering, but the new V4.1 Flash is positioned as a faster and cheaper successor. If DeepSeek’s internal assessment holds up under real-world workloads, the release could lower deployment costs for developers and intensify pricing pressure across the generative-AI market.

DeepSeek formally released V4.1 Flash around Sept. 10, saying the model surpasses V4 Pro across performance, cost and speed metrics. During the transition, all requests sent to Pro will be routed automatically to Flash and billed at the new rates. Off-peak output pricing will fall by two-thirds, a reduction that could materially shrink bills for developers and businesses running inference-heavy applications.

All Coverage

2 original reports

The Backstory

The history behind this event
Before this
DeepSeek Quadruples V4 API Prices During Peak Hoursfirst seen 2026-08-14 · 2 reports · similarity 0.84

DeepSeek has built its position in artificial intelligence by offering large-language-model APIs at prices below those of major rivals. V4 is the company’s flagship model, with usage billed by the million tokens. The introduction of peak-hour pricing signals a shift toward managing computing demand through variable rates, an approach that could influence when developers run workloads. Even after the increase, V4 remains relatively inexpensive compared with services from competitors including Anthropic.

DeepSeek said the new V4 API pricing will take effect on Aug. 17, 2026, with peak-hour charges rising to four times their current level. The reports did not disclose the specific per-million-token rates before or after the adjustment. The change represents a significant increase for developers running inference workloads during periods of heavy demand, although DeepSeek’s absolute pricing is expected to remain below comparable offerings from Anthropic and other leading AI providers.

DeepSeek Upgrades V4 Pro, Nears Claude Performance at Fraction of Costfirst seen 2026-08-13 · 3 reports · similarity 0.83

China’s DeepSeek is intensifying competition in generative AI with DeepSeek-V4-Pro, a flagship model aimed at reasoning, coding and general-purpose workloads. The model matters because performance approaching Anthropic’s premium Claude offering at a sharply lower cost could accelerate enterprise adoption and pressure global developers to reduce inference prices. It also extends DeepSeek’s strategy of challenging better-funded U.S. rivals through efficiency rather than spending alone.

DeepSeek quietly upgraded DeepSeek-V4-Pro to its production-ready “0813” release on Aug. 13, 2026. Reported benchmark results showed the model trailing Anthropic’s Claude Fable 5 by about 5.3%, while costing roughly 46 times less; expressed another way, the rival model was about 4,500% more expensive. The reports did not disclose exact U.S.-dollar rates per million tokens for either model, leaving the pricing comparison dependent on the cited test configuration.

DeepSeek Upgrades V4-Flash to Boost Agentic Codingfirst seen 2026-08-01 · 3 reports · similarity 0.84

DeepSeek, a Chinese AI startup, has built its challenge to U.S. model makers around open weights and aggressive pricing. V4-Flash uses a mixture-of-experts design with 284 billion total parameters, 13 billion activated per request and a 1 million-token context window. In May 2026, the U.S. National Institute of Standards and Technology’s Center for AI Standards and Innovation said V4 Pro cost less than GPT-5.4 mini on five of seven comparable benchmarks, underscoring why the series matters to developers and enterprise buyers.

On July 31, 2026, DeepSeek released DeepSeek-V4-Flash-0731 on Hugging Face and opened its API for public beta testing after re-running post-training while retaining the 284-billion-parameter architecture and MIT license. DeepSeek reported scores of 82.7 on Terminal Bench 2.1 and 54.4 on DeepSWE. API pricing is $0.14 per million uncached input tokens, $0.0028 for cached input and $0.28 for output, while the open weights allow companies to deploy the model on their own infrastructure.

DeepSeek Slashes API Cache-Hit Pricing to One-Tenth, Extends V4 Pro Discountfirst seen 2026-04-27 · 5 reports · similarity 0.81

Chinese AI company DeepSeek has long challenged OpenAI and Anthropic with low-cost APIs. Input caching allows existing content to be reused, directly affecting the operating costs of services with high volumes of repetitive requests, including RAG and agents. Developers are therefore watching the price changes closely.

DeepSeek most recently announced that input cache-hit API rates across its entire model lineup would fall to one-tenth of their previous levels. It also extended the flagship V4 Pro model’s limited-time 75% discount through May 5, 2026, bringing the discounted API output price to less than NT$30 per million tokens.

DeepSeek Formally Launches DeepSeek-V4 Preview Modelsfirst seen 2026-04-24 · 5 reports · similarity 0.85

DeepSeek shook the global AI market in early 2025 with its low-cost R1 model, prompting investors to reassess heavy spending on computing power. V4 is its next-generation flagship, with a greater focus on long-context processing, reasoning and autonomous agent capabilities. Its ability to close the gap with OpenAI, Anthropic and Google through an open-source, low-price strategy could shape the competitive position of Chinese models.

DeepSeek formally launched and open-sourced the V4 preview on April 24, 2026, offering the 1.6-trillion-parameter Pro and 284-billion-parameter Flash models. Both support a 1-million-token context window and thinking and non-thinking modes. For every 1 million uncached tokens, API input/output prices are 1/2 yuan for Flash and 3/6 yuan for Pro.

DeepSeek V4-Flash Launches on Ollama Cloud With Support for Major AI Development Toolsfirst seen 2026-04-24 · 2 reports · similarity 0.80

Chinese AI startup DeepSeek released a preview of its V4 model series on April 24, 2026. V4-Flash uses a mixture-of-experts architecture with 284 billion total parameters, activating only 13 billion at a time, and supports a 1-million-token context window. The design uses fewer active parameters to process code and lengthy documents, reducing the computing and deployment barriers for ultra-long-context inference.

As of July 20, 2026, V4-Flash was available on Ollama Cloud. Developers can use ollama launch to connect it to Claude Code, OpenClaw, Codex and OpenCode with a single command. The cloud version provides a 1-million-token context window, while Ollama has yet to disclose standalone pricing for the model.

After this
DeepSeek Prepares STAR Market IPO, Readies V4.1 Flashfirst seen 2026-09-10 · 1 reports · similarity 0.81

DeepSeek is preparing to tap China’s domestic capital markets as competition among artificial-intelligence model developers intensifies. A listing on Shanghai’s technology-focused STAR Market could provide the company with additional funding for research and computing infrastructure, while offering investors exposure to one of China’s prominent homegrown AI developers. The potential deal would also test market appetite for capital-intensive generative-AI businesses.

DeepSeek has hired CITIC Securities to prepare a STAR Market initial public offering, according to market reports, with the two sides now conducting due diligence. No fundraising target or listing timetable has been disclosed. Separately, the company plans to release its V4.1 Flash model around Sept. 10, saying the update will deliver broad improvements in performance and speed.

DeepSeek Releases V4.1-Flash Model With 1M-Token Contextfirst seen 2026-09-10 · 1 reports · similarity 0.84

DeepSeek AI has released DeepSeek-V4.1-Flash, a multimodal mixture-of-experts model aimed at handling unusually large workloads while reducing inference costs. Its 1 million-token context window is designed for tasks such as analyzing extensive code repositories, long documents and multimodal inputs. The model extends DeepSeek’s open-weight strategy, which has helped the Chinese AI developer compete with proprietary systems by giving researchers and businesses greater control over deployment.

DeepSeek-V4.1-Flash uses a causal encoder-decoder architecture to cut computation during prompt prefill, a costly stage when processing long inputs. It also introduces an FP4 KV cache and cross-layer attention reuse to substantially reduce the memory required for cached attention data. DeepSeek released the model weights under the MIT license, permitting research, modification and commercial use. The available report did not specify an exact release date, parameter count or benchmark results.

Mark Radar|MARK RADAR

If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →

All times are in Taipei time (GMT+8)