Kimi K3 Goes Viral—But GPU Servers Can’t Keep Up

Avatar 0

By | Chaqin Jun

Edited by | Wen Shuqi

It’s been less than a week since Kimi K3 launched, and it’s already so popular that Moonshot AI had to hit pause.

On the evening of July 19, Moonshot AI announced that due to a shortage of compute resources, it would stop accepting new subscribers for Kimi, prioritizing existing users with the limited GPU power available.

Source: Moonshot AI

On July 20, Moonshot AI responded to NUPIAO, clarifying that this change only affects new subscribers. Existing members’ benefits remain untouched, old plans haven’t been canceled, and current subscribers can still use the service normally with auto-renewals working as before.

For users who manually turned off auto-renewal, the company will gradually restore the renewal feature. Upgrading plans for existing users will also be reopened step by step. As for new plans, they’re still unavailable due to compute limitations, and a restart date hasn’t been set yet.

Looking at the official response, this move is more about rebalancing the experience between new and existing users given limited compute, rather than overhauling the membership system.

Ever since K3 dropped, developer communities worldwide have been buzzing with reviews and benchmarks. Many are comparing it head-to-head with giants like Claude and GPT. The conversation has shifted beyond just benchmark scores to real-world issues like inference costs, GPU resources, and service stability.

Just How Hot Is K3?

Right after launch, K3 became the talk of the global AI community.

According to the latest “Intelligence Index” from independent AI model evaluator Artificial Analysis, Kimi K3 scored 57 points overall, landing third globally—just behind Claude Fable 5 (60 points) and GPT-5.6 Sol (59 points), and ahead of Claude Opus 4.8 (56 points).

On top of that, K3’s average cost per task is about $0.94, roughly on par with GPT-5.6 Sol and only half of what Claude Opus 4.8 costs.

What really got the industry’s attention? K3 is the first open-source model to break into the top three of Artificial Analysis’s overall rankings. For a long time, the conventional wisdom was that open-source models lagged behind closed-source ones by about half a generation. K3 is closing that gap fast.

Source: Monolith Capital

Beyond the benchmark scores, K3’s sheer scale is turning heads. It’s Moonshot AI’s first trillion-parameter MoE model, boasting a total of 2.8 trillion parameters.

A source from a domestic AI chip company told NUPIAO that, since early this year, as overseas players like Anthropic keep rolling out bigger parameter models, the competition in China has swung back to “bigger parameters, longer contexts.” Trillion-parameter models with million-token contexts are now the battleground for top-tier models.

In their view, K3 exploded in popularity not just because of its capabilities, but because it marks the first time a domestic open-source model has reached this scale. “For the industry, 2.8T itself is a huge signal,” they said.

That buzz quickly spread to developer communities.

On platforms like X (Twitter), Reddit, OpenRouter, Linux.do, and V2EX, tons of developers jumped in to share their experiences. Many plugged K3 into AI coding tools like Cursor, Cline, and OpenCode for testing. From code generation to agent tasks to tool calling, K3 has been one of the most talked-about models in the AI space over the past few days.

The discussions boil down to three big questions: First, can it really replace Claude for coding? Second, is it truly on par with the top global models? And third, is it worth switching to?

Many developers feel K3’s real strength isn’t chat, but long-chain agent tasks and code generation. Some testers report that for complex code modifications and tool calls, K3 can already hold its own against international leaders like Claude and GPT.

But there’s also feedback that, in real-world engineering tasks, while the model is impressive, it still misses some details that require manual tweaking. Plus, its performance on complex statistical reasoning can be inconsistent.

Another hot topic is cost. Some users argue that, although K3’s per-unit price is competitive, the heavy token consumption for complex tasks means actual usage costs might not be as low as expected—especially in long-chain agent scenarios.

K3’s heat has also reached overseas AI circles.

Gavin Baker, founder of US tech investment firm Atreides Management, called K3 an “inflection point” for AI development. He believes high-quality open-source models are accelerating AI capability diffusion, benefiting not just model makers but the entire AI application ecosystem.

But some researchers are more cautious.

Ethan Mollick, a professor at the Wharton School of the University of Pennsylvania who studies AI’s impact on work, shared on X that he asked Kimi K3 to perform a complex statistical audit on one of his earlier academic studies. The model “messed it up in many ways,” including misusing statistical methods.

Mollick believes K3 is very close to the world’s leading models, but its reliability in complex professional tasks needs more real-world validation—not just benchmark scores.

This sentiment reflects a broader view in the AI community: on one hand, K3’s performance in code generation and agent tasks has won widespread praise from developers; on the other hand, its stability and reliability in complex, specialized tasks still need more proof from real-world use, not just leaderboard rankings.

Compute Crunch Is Still the Bottleneck

It’s precisely as developers rushed in and usage skyrocketed that Moonshot AI hit another wall—compute power.

Over the past two years, the typical narrative in the large model industry has been about “how many GPUs are needed to train a model” or “how big the parameters are,” not about having to pause new subscriptions because of too many users.

But the truth is, training doesn’t mean the compute demand ends. Especially with the rapid rise of AI coding and agent scenarios, models need to constantly call tools, read contexts, and run multi-step reasoning. A single request can eat up far more GPU time than a traditional chat session. The more powerful the model, the longer the context, and the more frequent the calls, the higher the inference compute consumption.

“The biggest problem today is still the shortage of compute, especially high-quality inference compute,” Chen Yuqin, investment director at Shanghai Guotou Futeng Capital, told NUPIAO earlier. “Even top model companies have started canceling unlimited usage and raising prices. That all points to the fact that inference resources are still tight. If the infrastructure isn’t sorted out, many application companies won’t dare to truly scale their products.”

“From a supply-demand perspective, domestic AI compute has long been in short supply, and with the wave of major model releases, this crunch is only going to get worse,” the chip industry source told NUPIAO.

They explained that independent model companies like Moonshot AI, Zhipu, and MiniMax mainly get their compute in two ways: renting from public cloud providers or building their own clusters. But public cloud resources are limited, and building your own is a heavy capital investment with long lead times. So there’s always pressure on compute supply.

“It’s widely accepted in the industry that a 10,000-card cluster is the baseline for training and inferring trillion-parameter models,” they said. “From 2024 to now, the number of companies that can actually deploy at that scale is still very small. Many big cluster projects that have been announced take years to come online, so there’s a long-term gap between supply and demand.”

At this year’s WAIC conference, one detail really stuck with the chip industry source.

They told NUPIAO that when various domestic chip makers showcased their supernode products, the two most common questions were: “Can this support a 10,000-card cluster?” and “Can it handle trillion-parameter model training?”

Compared to the old days of debating single-card performance and chip specs, the industry now cares more about ultra-large cluster capabilities and the infrastructure to support continuous training and inference of top-tier models.

This shift reflects a change in the competitive logic of large models.

In the past, model companies competed on parameter size, training capability, and benchmark scores. Today, as more models enter real production environments, inference efficiency, service stability, cost control, and infrastructure capabilities are becoming the new battlegrounds.

K3’s pause on new subscriptions might just be a temporary resource shuffle, but it’s given the world its first real glimpse into how the competition has entered a new phase. Model capabilities determine whether a product can attract users; inference compute, infrastructure, and sustained service determine whether it can handle real-world, large-scale applications.

Advertisement

Leave a Reply

Your email address will not be published. Required fields are marked *

Log In / Sign Up

Enter your email to receive a secure code. No password needed.