Kimi K3 Goes Open Source: The Dawn of Top-Tier Model Open Sourcing

Avatar 0

Journalist | Wu Yuyang

Editor | Wen Shuqi

After Kimi K3 went open source, the industry welcomed the first 3T-scale model with fully open weights — currently the largest open-source model globally by parameter count.

Over the past three years, a major shift in the global AI large model space has been Chinese model makers taking a stronger lead in the open-source community compared to their US counterparts.

K3 amplifies this effect. In its official announcement, Kimi stated that while K3’s overall performance still lags behind the strongest closed-source models like Claude Fable 5 and GPT-5.6 Sol, it demonstrates cutting-edge capabilities across its entire evaluation suite and consistently surpasses all other models.

Compared to Kimi K2, K3’s total parameters grew by 167%, active parameters increased by 220%, and context length expanded from 128K to 1M — an 8x improvement. Theoretically, its computational complexity scales quadratically.

The reason K3 can stack up to 2.8 trillion parameters and achieve a 2.5x efficiency boost (higher model capability with the same compute) lies in two core optimizations to the traditional Transformer architecture: KDA (Kimi Delta Attention / Mixed Linear Attention Mechanism) and Attention Residuals.

The former compresses the ever-growing KV Cache into a fixed-size matrix state, while the latter effectively reduces information decay and gradient vanishing during training, keeping information flow smooth.

At the same time, K3 leverages technologies like Stable LatentMoE (Stable Latent Mixture of Experts), MoonEP (low-level communication library), and FlashKDA (high-performance computing operators) to significantly optimize inference costs, communication capabilities, and computational efficiency.

Upon landing on Hugging Face, K3 topped the trending chart with over 4,000 likes in just 30 minutes. Hugging Face CEO Clem Delangue posted that this is the fastest release growth rate ever seen.

OpenAI President Greg Brockman, in a recent interview, called K3 an undeniably competitive new AI model and revealed, based on relevant evaluations, that China’s gap with the US in model development might be only about 4 months.

Additionally, Elon Musk has publicly praised the “Attention Residuals” research underlying K3’s architecture as “impressive.”

Beyond the impact of open sourcing itself, an AI large model developer told NUPIAO that the more direct commercial shock from K3’s open source — combined with lower prices and narrowing model gaps — fundamentally challenges closed-source vendors like OpenAI and Anthropic’s dominance over API pricing, which involves massive potential commercial interests.

Before K3, another sensational closed-source model was Anthropic’s Fable 5. While no official parameters were released, the industry widely speculated it to be a 5T-10T scale model.

The explosive popularity of both Fable 5 and K3 has, without exception, dispelled long-standing doubts in the industry: is it still meaningful to keep scaling up model parameters?

The aforementioned developer told NUPIAO that for model makers to stay in the game, “Scaling Up” is almost the only choice. “At least for now, post-training techniques like RL (Reinforcement Learning) are mainly ‘polishing.’ The better the base model, the higher the starting point for post-training.”

Chinese model makers are already following suit, developing models with over 2T parameters. Alibaba released the Qwen3.8 Max preview on July 19, a multimodal model with a massive 2.4 trillion parameters; in early July, reports emerged that MiniMax is developing a 2.7 trillion parameter next-gen LLM, internally called “M3 Pro.”

Both models are reportedly open-sourced upon release. In fact, this has become a key strategy for Chinese model makers to expand their technical influence.

Baidu’s Wenxin 5.0, officially launched in January this year, was described as a 2.4T parameter model, possibly making Baidu the earliest domestic company to reach this scale. However, this model didn’t generate much buzz domestically or internationally.

On one hand, it uses an ultra-sparse mixture-of-experts architecture, with active parameter ratio less than 3%, or about 70B. In contrast, K3’s active parameters are about 104B. To some extent, the scale of active parameters also affects model performance.

On the other hand, while Wenxin 4.5 joined the open-source ecosystem last September, Wenxin 5.0 remains a closed-source model. Industry insiders point out this is a coordinated approach of open-sourcing base models and keeping flagship models closed, aiming for a balance between technology popularization and commercial benefits — using the former to attract users and the latter to convert them.

But China’s open-source model ecosystem is growing stronger. Wenxin 5.0’s closure means many developers can’t deploy it themselves or analyze its internal structure through observation, losing some buzz from the geek community, which is increasingly crucial for a model’s current influence.

Beyond a series of comprehensive hardware and software challenges, the more pressing issue for makers chasing larger parameter models is computing power shortage.

According to industry analysts, just deploying K3 requires an 8-card B300 server, while training needs a top-tier AI GPU cluster at the ten-thousand-card level. On just the second day after launch, Kimi urgently issued a computing power shortage alert and paused new memberships.

In this regard, startups once again show weakness against big tech. A smart computing center insider said most of their computing power is already locked by large clients like top-tier companies. For startups like Kimi, this means their purchasing power lags behind, making it hard to replenish resources in emergencies like K3’s demand surge.

Another cloud vendor insider revealed their client composition: the top two are a major hardware giant and a major internet giant, with other model makers’ combined monthly revenue not matching either one. Moreover, due to the scarcity of top-tier GPUs, cloud vendors face delays in fulfilling customer orders. Their biggest client’s order of a thousand units was only 10% delivered in the first half of the year.

Even so, the drive for model makers to push better models and open-source them is, to some extent, urgent.

Almost the same week Kimi K3 was open-sourced, Jensen Huang, along with 25 tech giants including Microsoft, Meta, IBM, and Intel, issued a public letter supporting the open-source ecosystem, with OpenAI and Google subsequently signing on.

Whether due to public opinion or commercial pressure, Anthropic CEO Dario Amodei, always firm on technology blockade, recently made a public statement that Anthropic has never advocated banning open-weight models.

With more influential voices, the global open-source model trend is becoming unstoppable. From a technology popularization perspective, this is undoubtedly positive. But for companies, from technology to costs, full open-source will only make them more transparent to users. Then, breakthroughs and survival won’t be any easier than today.

Advertisement

Leave a Reply

Your email address will not be published. Required fields are marked *

Log In / Sign Up

Enter your email to receive a secure code. No password needed.