Kimi K3 officially announces the Burning Art, now it truly becomes a living Bodhisattva.

I declare that Kimi is now officially at the same table with DeepSeek, becoming a true living Buddha of the open-source world.
Because last night, Kimi K3 was officially open-sourced on Hugging Face.
Within half an hour, it gained 4000 likes on Hugging Face, making it arguably the most anticipated open-source model ever. I estimate many foreigners and the reviewer were waiting right on time.
Moreover, what Kimi open-sourced this time is not simple. Besides the model weight files themselves, it also included a detailed technical report and released many training model tools together.
Now, just like DeepSeek, it truly took out the secret technique for the community.
I won't go into details on various benchmark tests; basically, all scores are only lower than Fable 5 and GPT-5.6 Sol.
It can be said that K3 is now the world's third-best model, and among the top 3, it's the only one you can download to run locally and fine-tune yourself.
Moreover, it comes with multimodal capabilities natively, with no weaknesses.
The model parameters are even larger than larger. This time, K3's total model parameters directly reached 2.8T, making it the largest open-source model.
And not only the total parameters are large, the model's activation parameters are also absurdly large,
When actually working, K3 can directly activate 104.2 billion parameters, which is three times that of its previous generation, twice that of DeepSeek V4 Pro, and even larger than the entire size of many open-source models.
To handle this massive amount of data, K3 made many architectural innovations this time. Of course, they didn't hide anything and openly wrote everything in the documentation.
First, regarding the attention aspect that everyone cares about. Everyone knows the traditional attention mechanism: every time the model reads a new word, it needs to compute it with all previous tokens. So when the content read doubles, the computational cost directly quadruples.
To save computational power, Kimi adopted a method called mixed attention this time. When information comes in, it first passes through a layer of KDA (Kimi Delta Attention). KDA reads the text while simultaneously writing key information into a continuously updated state.
After new content arrives, it also determines which old information needs to be retained and which can be gradually forgotten.
At the same time, to prevent KDA from missing key information, Kimi also arranged a Gated MLA after three layers of KDA, which is equivalent to having the model take a few pages of notes and then go back to check the original text.
With this combination, K3 significantly reduces the computational cost for extremely long contexts.
In addition, Kimi also designed Attention Residuals connections to reduce information loss during transmission.
It also designed Stable LatentMoE, which compresses the original 7168-dimensional information to 3584 dimensions, then simultaneously distributes it to 16 experts for processing.
It also arranged rate-limiting and scheduling mechanisms for these experts to suppress abnormal activation values and reasonably allocate tasks.
To avoid situations where some work 996 every day while others slack off all day.
Through these upgrades in architecture, data, and training methods, despite the increase in total parameters, K3's training efficiency has improved by 2.5 times compared to the previous generation K2.
And the release and open-sourcing of Kimi K3 is no longer just a self-congratulation in a small tech circle.
It's like a fuse that has brought out all kinds of people, good and bad, all sorts of characters.
For example, some American politicians accused K3 of being distilled from American large models.
Kimi employees also made a beautiful sarcasm, saying that Fable was only released not long ago, and Kimi distilled a model in just tens of days, which could apply for a Guinness World Record. Are Chinese people all superpowered, distilling the model that Anthropic plans to release in the future?
I get it now. Kimi invented a time machine, reversed time to start distilling Fable. How could Kimi be so evil...
In addition, the US government recently prepared to sanction and investigate our country's AI companies.
Our Ministry of Commerce also made a perfect counterattack.
On the other side, the high-EQ Jensen Huang also registered on X, posted an open letter explaining the importance of open source, and rallied nearly 100 tech giants including Microsoft, Google, and OpenAI to jointly sign.
Publicly opposing the ban on Chinese open-source models.
Under pressure, the staunch closed-source warrior Anthropic once again posted a long article, saying, "Oh, we don't oppose open source, we're just afraid that powerful models will fall into the hands of bad people."
To be honest, it's quite surprising that open-source models could catch up so quickly to these world-class closed-source models this time...
The reviewer still remembers the first time seeing Mythos and Fable. Although Anthropic never does anything good, the models they make are indeed quite impressive.
At that time, everyone generally thought the gap between open-source and closed-source models would grow larger and larger.
Because these two types of models are in a completely unequal race.
Closed-source models can freely learn techniques from open-source models. For example, DeepSeek once painstakingly taught everyone how to make models learn reasoning and thinking.
But conversely, open-source models cannot learn much from closed-source models.
But fortunately, there are still many opportunities for open-source models to learn from each other and support each other.
Open Kimi's paper this time, and you'll find that DeepSeek's shadow is everywhere in it.
From multi-head attention MLA, mixture of experts MoE, to training stability and inference efficiency, many techniques contain the accumulation of the open-source community over the past few years.
But Kimi did not simply copy; instead, on this basis, it continued to create new things like KDA, Gated MLA, AttnRes, and Stable LatentMoE.
This culture of mutual learning and mutual respect in the open-source community is the most valuable thing about open-source models.
The experience left by the previous generation of models directly becomes the starting point for the next generation.
These open-source models thus leverage each other, building layer upon layer, and forcefully caught up with those closed-source giants that have more computing power, funding, and data.
Thanks to the continuous efforts of these open-source models, at least now after the release of K3, no one will say such hot takes as “open-source models will fall further behind.”
And the development of AI is no longer just a conflict between companies, open source vs. closed source. It is not only related to technological equality but also inseparable from national security, geopolitics, and so on.
I also hope that in the future, there will be more and more excellent models like K3 in our country.
After all, reasoning is certainly useful, but many times, showing strength works better.
Related Articles

Huawei phone users have finally got what they've been waiting for! The HarmonyOS trial beta version of NetEase Cloud Music is now available.
about 4 hours ago

The price after discounts is 6544 yuan! Lenovo ThinkPad E14 2026 laptop is now available: Core 5 320 + 512GB
about 5 hours ago















