DeepSeek V4-Flash Launches on Ollama Cloud With Support for Major AI Development Tools
Chinese AI startup DeepSeek released a preview of its V4 model series on April 24, 2026. V4-Flash uses a mixture-of-experts architecture with 284 billion total parameters, activating only 13 billion at a time, and supports a 1-million-token context window. The design uses fewer active parameters to process code and lengthy documents, reducing the computing and deployment barriers for ultra-long-context inference.
As of July 20, 2026, V4-Flash was available on Ollama Cloud. Developers can use ollama launch to connect it to Claude Code, OpenClaw, Codex and OpenCode with a single command. The cloud version provides a 1-million-token context window, while Ollama has yet to disclose standalone pricing for the model.
All Coverage
2 original reportsThe Backstory
The history behind this eventDeepSeek Upgrades V4-Flash to Boost Agentic Coding
DeepSeek, a Chinese AI startup, has built its challenge to U.S. model makers around open weights and aggressive pricing. V4-Flash uses a mixture-of-experts design with 284 billion total parameters, 13 billion activated per request and a 1 million-token context window. In May 2026, the U.S. National Institute of Standards and Technology’s Center for AI Standards and Innovation said V4 Pro cost less than GPT-5.4 mini on five of seven comparable benchmarks, underscoring why the series matters to developers and enterprise buyers.
On July 31, 2026, DeepSeek released DeepSeek-V4-Flash-0731 on Hugging Face and opened its API for public beta testing after re-running post-training while retaining the 284-billion-parameter architecture and MIT license. DeepSeek reported scores of 82.7 on Terminal Bench 2.1 and 54.4 on DeepSWE. API pricing is $0.14 per million uncached input tokens, $0.0028 for cached input and $0.28 for output, while the open weights allow companies to deploy the model on their own infrastructure.
DeepSeek Formally Launches DeepSeek-V4 Preview Models
DeepSeek shook the global AI market in early 2025 with its low-cost R1 model, prompting investors to reassess heavy spending on computing power. V4 is its next-generation flagship, with a greater focus on long-context processing, reasoning and autonomous agent capabilities. Its ability to close the gap with OpenAI, Anthropic and Google through an open-source, low-price strategy could shape the competitive position of Chinese models.
DeepSeek formally launched and open-sourced the V4 preview on April 24, 2026, offering the 1.6-trillion-parameter Pro and 284-billion-parameter Flash models. Both support a 1-million-token context window and thinking and non-thinking modes. For every 1 million uncached tokens, API input/output prices are 1/2 yuan for Flash and 3/6 yuan for Pro.
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.
If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →