Mark RadarMARK RADAR
EN

DeepSeek V4-Flash Launches on Ollama Cloud With Support for Major AI Development Tools

2 reports · First detected 2026-04-24 · Last active 2026-04-27

Chinese AI startup DeepSeek released a preview of its V4 model series on April 24, 2026. V4-Flash uses a mixture-of-experts architecture with 284 billion total parameters, activating only 13 billion at a time, and supports a 1-million-token context window. The design uses fewer active parameters to process code and lengthy documents, reducing the computing and deployment barriers for ultra-long-context inference.

As of July 20, 2026, V4-Flash was available on Ollama Cloud. Developers can use ollama launch to connect it to Claude Code, OpenClaw, Codex and OpenCode with a single command. The cloud version provides a 1-million-token context window, while Ollama has yet to disclose standalone pricing for the model.

All Coverage

2 original reports

The Backstory

The history behind this event
DeepSeek Formally Launches DeepSeek-V4 Preview Models2026-04-28 · 5 reports · similarity 0.81

DeepSeek shook the global AI market in early 2025 with its low-cost R1 model, prompting investors to reassess heavy spending on computing power. V4 is its next-generation flagship, with a greater focus on long-context processing, reasoning and autonomous agent capabilities. Its ability to close the gap with OpenAI, Anthropic and Google through an open-source, low-price strategy could shape the competitive position of Chinese models.

DeepSeek formally launched and open-sourced the V4 preview on April 24, 2026, offering the 1.6-trillion-parameter Pro and 284-billion-parameter Flash models. Both support a 1-million-token context window and thinking and non-thinking modes. For every 1 million uncached tokens, API input/output prices are 1/2 yuan for Flash and 3/6 yuan for Pro.

Mark Radar|MARK RADAR
All times are in Taipei time (GMT+8)