Just days after unveiling its long-awaited V4 large language model, Chinese AI lab DeepSeek announced on Saturday a massive price cut across its API services. The cost for cached input has dropped to one-tenth of its original price, while the flagship V4-Pro model received a temporary 75% discount—valid until May 5th—according to Technology Review.

The cumulative discounts bring the price of V4-Pro cached input down to just 0.025 yuan—approximately $0.0036—per million tokens, while standard promotional prices for input and output stand at 3 and 6 yuan per million tokens, respectively. This pricing makes DeepSeek’s offering significantly cheaper than its Western rivals. According to OpenRouter data cited by Chinese financial outlet Gelonghui, the output costs for models such as Anthropic’s Claude Opus, OpenAI’s GPT-5.4, and Google’s Gemini 3.1 Pro range from $12 to $25 per million tokens.

A New Front in the AI Price War

On April 24, DeepSeek released V4-Pro and V4-Flash in preview mode—the Haizhou-based startup’s first major release since V3.2 last December. With a total of 1.6 trillion parameters and 49 billion active parameters per inference pass, V4-Pro is the largest open-weights model available on the market, while V4-Flash offers a more compact 284-billion-parameter option.

Even before the additional discounts, V4-Pro’s standard rates—$1.74 per million input tokens and $3.48 per million output tokens—were already incomparably lower than competitor rates: approximately 98% cheaper than OpenAI’s GPT-5.5 Pro, per Yahoo Tech. The latest round of cuts has widened this gap even further. "Against the backdrop of generally rising computing costs in the AI industry this year, DeepSeek V4 has once again fully embodied the concept of 'AI price reduction,'" Gelonghui reports.

Powered by Huawei

DeepSeek reported that V4 runs on Huawei Ascend hardware rather than Nvidia chips—a development CNN called potentially more significant than the model’s performance itself. Wei Sun, principal AI analyst at Counterpoint Research, told CNN that this "opens the possibility of developing and deploying AI systems without relying exclusively on Nvidia," which could accelerate both domestic adoption and global AI progress.

The new model’s efficiency is also noteworthy. According to MIT Technology Review, with a context window of one million tokens, V4-Pro consumes only 27% of the computing resources required by its predecessor, V3.2. In its technical paper, DeepSeek acknowledged that V4 "is slightly inferior to GPT-5.4 and Gemini 3.1 Pro," trailing top-tier models by approximately three to six months.