DeepSeek Releases Open-Source V4 Models with a 1-Million Token Context Window

April 24, 2026  12:07

On Wednesday, the Chinese artificial intelligence laboratory DeepSeek unveiled preview versions of its highly anticipated V4 model, releasing two variants into the public domain. According to the company, these models rival the capabilities of the world’s most powerful closed-source artificial intelligence systems while costing significantly less. This release marks DeepSeek’s most ambitious launch since the R1 model, which shook global artificial intelligence markets in early 2025.

Two Models, One Million Tokens

DeepSeek announced the release on its official X account, stating that "DeepSeek-V4 Preview is officially launched and open to everyone," and proclaiming the arrival of the "era of the cost-effective 1-million token context window." On the same day, the company published the model weights and a technical report on Hugging Face.

The flagship DeepSeek-V4-Pro features a total of 1.6 trillion parameters, with 49 billion activated per token. The lightweight version, DeepSeek-V4-Flash, contains 284 billion parameters with 13 billion active parameters. Both models are built on the Mixture-of-Experts architecture and support a context window of one million tokens — one of the largest in the industry. A new hybrid attention mechanism, combining Compressed Sparse Attention and Heavily Compressed Attention, allows V4-Pro to use only 27% of the computational resources during inference and 10% of the cache memory compared to its predecessor — DeepSeek-V3.2 — for the same context window size.

Benchmarks and Prices

According to DeepSeek’s own evaluations, V4-Pro in maximum reasoning mode scored 93.5 points on LiveCodeBench and achieved a rating of 3206 on Codeforces, outperforming Google’s Gemini-3.1-Pro and OpenAI’s GPT-5.4 in both metrics. On the widely monitored SWE-bench Verified, V4-Pro scored 80.6% — within one percentage point of Anthropic’s Opus-4.6, which leads with a score of 80.8%. In the technical report, V4-Pro-Max is described as "the best open-source model available today," although it trails behind closed-source market leaders in several benchmarks related to knowledge and agent-based tasks.

The pricing could prove to be as revolutionary as the benchmark results. V4-Flash is offered at a price of $0.14 per million input tokens and $0.28 per million output tokens, while V4-Pro costs $1.74 and $3.48 respectively — according to the company’s published tariff plan. This makes V4-Pro approximately three times cheaper than GPT-5.5 for input tokens and nearly nine times cheaper for output tokens, according to an analysis by Apidog.


 
 
 
 
  • Archive