Staff Reporter | Song Jianan
According to the latest weekly global AI model usage rankings (July 27 – August 2) released by OpenRouter, a multi-model aggregation platform, DeepSeek V4 Flash claimed the top spot with a staggering 7.22 trillion tokens processed in a single week.
Meanwhile, the open-source project team OpenCode reported that DeepSeek V4 Flash’s usage on their platform skyrocketed—on August 1 alone, the model handled a jaw-dropping 8 trillion tokens in one day. Of that, 5 trillion came from free trial quotas, while the remaining 3 trillion was paid usage by developers through the OpenCode platform.
OpenRouter’s data shows total global AI model calls last week reached 56.8 trillion tokens, a slight 2.07% dip week-over-week. Among the ranked models, Chinese AI models collectively accounted for 28.13 trillion tokens, down 14.76% from the previous week. In contrast, US AI models posted 4.38 trillion tokens, surging 87.18% week-over-week. Notably, Chinese models have now held the global lead for 14 consecutive weeks, comfortably outpacing their US counterparts.
On August 5, NUPIAO checked OpenRouter’s current weekly chart and found that DeepSeek V4 Flash 0423 still sits at No. 1, with 6.92 trillion tokens as of press time. Originally launched in April this year, this model strikes a balance between inference speed, long-context capabilities, and cost efficiency, making it a go-to choice for developers and commercial enterprises seeking high-throughput inference services.

Taking second place is Xiaomi’s MiMo-V2.5, with 5.1 trillion tokens for the week—a sharp 52% drop from the previous week. The traffic pullback signals a cooling of market enthusiasm after its earlier surge.

Xiaomi’s model officially launched its public beta on April 23 and went fully open-source by the end of April. Built on a Mixture-of-Experts (MoE) architecture, it boasts over one trillion total parameters with just 42 billion activated per inference. It comes standard with a 1-million-token ultra-long context window and supports full multimodal interaction across text, voice, and images. Its low inference costs and rock-solid agent performance are the key reasons overseas developers are flocking to integrate it.
Just the week before (July 20–26), MiMo-V2.5 had actually topped the OpenRouter chart with 10.5 trillion tokens, a 12% week-over-week increase.
In third place is Tencent’s self-developed general-purpose model, Hunyuan Hy3, with 5.01 trillion tokens. Its traffic remained essentially flat this period, showing 0% growth. Hy3 was the fastest riser on the chart during the week of July 26, when its weekly calls hit 3.94 trillion tokens—an eye-popping jump of over 999% week-over-week.
Hy3 officially went open-source on July 6, featuring 295 billion total parameters with only 21 billion activated per inference. It uses a hybrid fast-and-slow thinking architecture, supports up to 256K context, and delivers major upgrades in code generation and interactive intelligence. Tencent Hunyuan previously announced that as of July 15, Hy3’s total call volume had grown more than 68-fold compared to its predecessor, Hy2.
Fourth place also belongs to DeepSeek—the newly released V4 Flash 0731 model—which pulled in 3.45 trillion tokens this week. Rounding out the top five is OpenAI’s GPT-5.6 Luna, which exploded onto the scene with 2.99 trillion tokens, an astronomical 738% surge.
Looking at the top five, Chinese-made models lock down four of the five spots, demonstrating overwhelming market traction. But OpenAI’s aggressive push with new iterations shows that the race among global tech giants is far from over.
In the top ten, DeepSeek holds three slots—alongside the two V4 Flash versions, DeepSeek V4 Pro ranks sixth with 2.97 trillion tokens. Anthropic, Google, NVIDIA, MiniMax, StepFun, and Zhipu AI all made the list. Free models are increasingly eating up traffic share, with NVIDIA Nemotron 3 Ultra, Poolside Laguna S 2.1, and Inclusion AI Ling-3.0-flash all posting rapid growth—many free models saw week-over-week gains exceeding 50%.
Worth noting: Kimi K3, which had drawn attention from Elon Musk, dropped out of the top ten this week to land at No. 12. That model packs 2.8 trillion total parameters with a 1-million-token context window, making it the largest open-source AI model globally by parameter count.
Unlike the old playbook of buying traffic with rock-bottom prices, today’s Chinese models are winning over overseas developers through continuous improvements in inference performance and dependable service stability. But challenges remain—OpenAI and Anthropic keep shipping new versions that narrow the performance gap, while Google’s Gemini series holds steady near the top of the charts, giving overseas giants a solid foundation to build on.
Divergence within the race is also worth watching. Several Chinese models saw significant traffic declines this period: Xiaomi’s MiMo-V2.5, MiniMax M3, and Zhipu’s GLM5.2 all experienced varying degrees of pullback. With hype cycles rotating faster than ever, any slowdown in product iteration could quickly translate into lost developer traffic.
One AI industry analyst pointed out that the overseas independent developer market has reached a fever pitch. Relying on a single version update is no longer enough to hold your ground. The ability to keep iterating while managing inference costs efficiently will ultimately determine which models can retain users over the long haul.