Kimi K3 Puts a 2.8-Trillion-Parameter Model on Hugging Face
Moonshot has shipped the largest open-weight model yet, and the interesting part is not the parameter count.
Moonshot AI has published the weights for Kimi K3, the model it launched on its own app and API in mid-July. The weights, posted to Hugging Face on 27 July alongside a technical report on GitHub, belong to a 2.8-trillion-parameter mixture-of-experts model with 104 billion parameters active per token, native image and video understanding and a one-million-token context window. That makes it the largest open-weight model anyone has released, and it arrives under Moonshot's own Kimi K3 Licence, which permits research and commercial deployment with conditions attached.
Moonshot's headline numbers are strong. It reports 67.5 on DeepSWE, 91.2 on BrowseComp and 90.0 on Video-MME. Independent scoring is thinner so far, but Artificial Analysis placed K3 at an Elo of 1,547, behind only Claude Fable 5, according to MLQ's write-up. On the API it costs $3 per million input tokens and $15 per million output tokens.
The architecture is the story
Anyone can announce a big number. What makes K3 worth reading about is how Moonshot got there. The model routes each token through 16 experts selected from a pool of 896, a scheme the company calls Stable LatentMoE. It introduces a new attention mechanism it calls Kimi Delta Attention, adds what it terms Attention Residuals, and ships with MXFP4 quantisation support. Moonshot claims the combination delivers roughly 2.5 times the scaling efficiency of the Kimi K2 generation.
I would treat that efficiency figure as a claim rather than a result until someone outside Moonshot reproduces it. But the technical report is public, and so are three pieces of infrastructure the company released alongside the weights: MoonEP for expert parallelism, FlashKDA for its attention kernel, and AgentEnv for agent training environments. That is more of the machinery than most labs, open or closed, tend to hand over. The model card also ships deployment recipes for vLLM and SGLang from day one, which matters more to practitioners than any leaderboard.
Open weights, but who can run them?
Here is the sobering part. A 2.8-trillion-parameter model is open in the legal sense and closed in the practical one for nearly every organisation reading this. Even with a sparse design and four-bit quantisation, serving K3 means a multi-node GPU cluster. The realistic audience for the weights is cloud providers, inference start-ups, national labs and the research community, not a team with a couple of workstations.
That still matters. Open weights at this scale let researchers study a frontier-class model's internals, let inference providers compete on price for the same model, and let regulated organisations host it inside their own perimeter. It is a different kind of value from the "run it on your laptop" promise that open models used to carry.
There is also a token-efficiency caveat. MLQ notes that K3 uses 21 per cent fewer output tokens than K2.6, which is welcome, but also that it spent more than 13,000 tokens on a simple SVG drawing task, about $0.25 for a single answer. Per-token price is only half the cost picture; reasoning models that think at length can be cheap per token and expensive per task. Anyone comparing K3 against closed models should measure cost per completed task, not cost per million tokens.
What to watch
- Independent evals. Moonshot's self-reported scores are competitive with the best closed models. Third-party agentic and coding evaluations over the next few weeks will show whether that holds outside its own harness.
- Hosting prices. Once several providers serve the same weights, the market price of K3 inference will tell us more about its real efficiency than any paper.
- The response from other Chinese labs. Alibaba is already previewing a 2.4-trillion-parameter Qwen model. The race to the top of the open-weight tables is now being run almost entirely by Chinese labs, and the gap to closed frontier models is narrower than it has ever been.
The honest summary: K3 is a genuine engineering contribution wrapped in a parameter count that most people cannot use directly. Read the report, ignore the size.
Sources