On August 17, DeepSeek officially rolled out its updated API pricing—and yes, the numbers are turning heads. The company had teased the changes back on August 13 via its official WeChat account, introducing a peak-and-off-peak pricing model where off-peak rates are half of what you’d pay during busy windows. Peak hours run from 9:00 AM to 12:00 PM and 2:00 PM to 6:00 PM Beijing time; everything else is considered off-peak. The new rates took effect at midnight (Beijing time) on August 17.

Let’s break down the numbers with DeepSeek’s flagship model, V4-Pro, as the example. During peak hours, input costs (cache miss) jump to 9 yuan per million tokens—that’s a hefty 200% increase. Output pricing climbs to 27 yuan per million tokens, a 350% spike. And here’s the kicker: cache-hit input now costs 0.3 yuan per million tokens, which represents a jaw-dropping 1,100% surge from previous levels. If you can shift your workloads to off-peak windows, though, you’ll pay just 4.5 yuan for input (cache miss), 13.5 yuan for output, and a mere 0.15 yuan for cache-hit input.
For context, DeepSeek-V4-Pro’s official release came just days earlier on August 13, bringing enhanced agent capabilities to the table. At launch, its pricing sat roughly three times higher than the lightweight V4 Flash model. Per million tokens, V4-Pro was priced at 0.025 yuan (cache hit) and 3 yuan (cache miss) for input, with output at 6 yuan. V4 Flash, by comparison, was priced at 0.02 yuan, 1 yuan, and 2 yuan respectively.
So why the shake-up? DeepSeek says the core motivation behind this peak-off-peak pricing strategy is tackling daytime compute congestion. By leveraging market-based price signals, the company hopes to nudge enterprise developers into scheduling their large-model tasks during off-peak hours, smoothing out the demand curve and boosting overall platform stability.

It’s been a rocky year for DeepSeek on the reliability front. The service crashed repeatedly on May 8, 21, 24, and 28—blamed largely on a perfect storm: user volume exploding by 66.7% while compute capacity only grew 8.3%, leaving a serious supply-demand gap.
Hualong Securities weighed in earlier, noting that DeepSeek’s API price increase signals the end of the year-long price war among domestic large models. The industry’s pricing baseline is now heading into a recovery phase, which will have structural ripple effects on both the compute and application sides of the ecosystem.
DeepSeek isn’t alone in this move. Before its adjustment, Tencent Cloud, Alibaba Cloud, and Baidu AI Cloud had all raised their API compute product prices. Zhipu AI has actually hiked API prices three times already this year.
Speaking at Zhipu’s 2025 earnings call, CEO Zhang Peng revealed that API pricing jumped 83% in Q1 2026—yet demand still outpaced supply, with call volumes growing 400%. Today, Zhipu ranks among the top domestic vendors by paid token consumption.
The broader trend here is pretty clear: tech companies are raising API prices because AI is advancing at breakneck speed, and compute supply simply can’t keep pace with the explosive growth in token call volumes.
Global model aggregator OpenRouter’s latest weekly token rankings paint a vivid picture. From August 3 to August 9, worldwide AI large-model call volumes hit 69 trillion tokens, up 21.48% week-over-week. Chinese AI models accounted for 34.25 trillion of that total, growing 21.76% sequentially, while US models contributed 9.17 trillion, up 109.36%. Notably, Chinese AI models have now outpaced their American counterparts on OpenRouter for 15 consecutive weeks, holding the global top spot.
The same rankings show that all top four AI models globally by call volume during that week came from China. Leading the pack was DeepSeek-V4-Flash-0731 (the official V4-Flash release) with 8.83 trillion tokens called—a staggering 570% jump week-over-week. Tencent’s Hy3 model took second place with 8.05 trillion tokens, up 67%. DeepSeek-V4-Flash-0423 (the preview version) ranked third at 5.88 trillion tokens, down 19%. Xiaomi’s MiMo-V2.5 rounded out the top four with 5.39 trillion tokens, slipping 14%.
To keep up with this insatiable token appetite, companies like Tencent are throwing serious money at compute infrastructure. Tencent’s Q2 2026 earnings report shows capital expenditures surging 176% to 52.78 billion yuan during the period, with free cash flow sitting at negative 13.8 billion yuan—only turning positive at 37.6 billion yuan once AI compute prepayments are stripped out.
Tencent President Martin Lau explained the strategy: “If we proceed step by step—first using compute for our own model building, then for application development, and finally renting out spare capacity—we can build a native AI business at significant scale that generates very healthy profits.”
East Money Research Center points out that the core bottleneck in compute supply lies in AI chips. According to IDC data, China shipped roughly 4 million AI accelerator cards in 2025, with domestic cards accounting for about 1.65 million units—a share of around 41%. Today, domestic substitution has become the critical path to bridging the compute gap.