By NUPIAO News Desk
On August 6th, DeepSeek dropped a bombshell announcement, saying they plan to hike the pricing on their API services across the board in the near future. They flagged the increase as “substantial” and urged users to plan their usage accordingly, adding that the official rates will be shared in a formal notice later.

After digging into the DeepSeek open platform, we found that their two flagship models, V4-Flash and V4-Pro, are still operating at the reduced rates from their earlier price cuts.
Here’s the current official pricing: For deepseek-v4-flash, the input cost is just 0.02 yuan per million tokens when the cache hits, 1 yuan per million tokens when it misses, and 2 yuan per million tokens for output. For the beefier deepseek-v4-pro, it’s 0.025 yuan per million tokens on cache hits, 3 yuan per million tokens on misses, and 6 yuan per million tokens for output.

DeepSeek has also rolled out a peak-valley billing system. During busy weekday hours, the price doubles, but nights, weekends, and holidays stick to the lower off-peak rates. This is clearly designed to ease the strain on their computing resources. Both models support ultra-long contexts up to a million tokens and are compatible with both OpenAI and Anthropic call formats.
Looking back at 2026, DeepSeek has been tweaking its API prices quite a bit.

Back on April 25th, they slashed the cache-hit input price and kicked off a limited-time 2.5% discount on the deepseek-v4-pro API, which ran until May 5th, 2026. After the discount, the input cost for a million tokens (cache hit) dropped to 0.25 yuan, while cache-miss input stayed at 3 yuan and output at 6 yuan.
At the time, analysts were buzzing that deepseek-v4 was the hottest domestic model event of Q2. Its full domestic adaptation signaled that China’s “model-chip-cloud” ecosystem was finally clicking into place, which could help build a self-sustaining AI commercial loop. The stronger the model and the faster the iteration, the more it ripples down to AI chips, devices, and cloud infrastructure, giving a direct boost to domestic computing power players.
Then on May 23rd, DeepSeek announced that after the 2.5% discount promo ended on May 31st, 2026, API prices would formally settle at one-quarter of the original rate. That meant input with cache hits at 0.025 yuan per million tokens, cache misses at 3 yuan, and output at 6 yuan—a solid 75% cut from the original. Right after DeepSeek’s move, Xiaomi’s MiMo-V2.5 series API also declared a permanent price drop, with cuts up to 99%. The flurry of discounts was widely seen as a sign that the domestic large-model price war was hitting fever pitch.
But the tide turned at the end of June. DeepSeek sent out upgrade reminder emails to users, saying the official V4 version would launch in mid-July, packed with more feature tweaks and performance boosts.
DeepSeek explained that to allocate resources more wisely and keep service stable, they’d be adjusting the API pricing strategy after the official release, introducing a peak-valley pricing model where peak-hour calls would cost double.
Here’s the breakdown: For deepseek-v4-pro, a million tokens of input (cache hit) normally costs 0.025 yuan, but 0.05 yuan during peak hours. For cache misses, it’s 3 yuan normally and 6 yuan at peak. Output is 6 yuan normally and jumps to 12 yuan during peak times. DeepSeek defines peak hours as 9:00-12:00 and 14:00-18:00 (Beijing time) daily.
The deepseek-v4-flash is more budget-friendly: input (cache hit) is 0.02 yuan normally and 0.04 yuan at peak; cache-miss input is 1 yuan normally and 2 yuan at peak; and output is 2 yuan normally, 4 yuan at peak.
Some industry watchers argue that the peak-valley pricing in the official V4 isn’t simply a price hike, but rather a standardized way to manage resources when computing power is scarce.

The looming steep price increase is also a strategic response to the crushing cost of computing power. Running large-model inference eats up high-end GPUs like crazy, and keeping prices low for volume over a long stretch just keeps draining a company’s cash flow.
According to the weekly global AI model usage ranking from OpenRouter (July 27 to August 2), DeepSeek V4 Flash took the top spot with a jaw-dropping 7.22 trillion tokens processed. Data from the open-source project team OpenCode also shows that DeepSeek V4 Flash’s usage has exploded—on August 1st alone, the model handled 8 trillion tokens in a single day. And on August 4th, deepseek-v4-flash actually hit capacity issues due to unprecedented traffic.
It’s not just DeepSeek. Over the past six months, many large-model providers have tweaked their pricing in different ways. At Zhipu’s 2025 performance briefing in April, CEO Zhang Peng revealed that API pricing would jump 83% in Q1 2026. Even with that, demand is still outstripping supply, with call volumes growing by 400%.
In June, Doubao Professional Edition introduced a three-tier subscription model: Standard, Enhanced, and Premium plans at 68 yuan, 200 yuan, and 500 yuan per month, respectively. Analysts generally agree that paid subscriptions signal that domestic large-model apps are moving into a phase of monetization and value validation.
And on July 19th, Kimi announced that within 48 hours of the Kimi K3 release, user requests had far exceeded projections and were nearing the limits of their current cluster capacity. They immediately paused new consumer subscriptions to focus on expanding computing power, promising to gradually reopen slots once new capacity comes online.
Over the past couple of years, everyone was slashing API prices to grab market share, even offering massive free tiers to lure developers over. But heading into the second half of 2026, with computing power supply-demand gaps, rising chip procurement costs, and ballooning operational expenses, the old playbook of trading low prices for scale just isn’t sustainable anymore. The whole industry is shifting from “racing to the bottom on price” to “commercialization with viable costs.”
As for DeepSeek’s upcoming price hike, some believe this pricing correction shows vendors are finally getting serious about sustainable operations, but it’s also going to put direct pressure on small and mid-sized developers working on AI innovation projects.