Kimi · model № 003
Kimi K3
A 2.8T-parameter model that scales delta-rule attention, expert routing, and long-horizon training together
total params
2.8T
largest open ever
active / token
104B
16 of 896 experts
context
1M
agentic RL at full length
scaling efficiency
2.5×
over Kimi K2
Chapters
- 1The 2.8T StatementWhat a 2.8T-parameter open model is trying to prove, and where the public results support the claim.
- 2The Bet That PaidHow delta-rule attention edits a fixed memory state, and why Kimi scaled the idea to 2.8T parameters.
- 3Depth, RewiredHow residual attention paths help information survive when a model becomes very deep.
- 416 of 896How latent routing keeps expert traffic balanced without the usual auxiliary loss.
- 5Million-Token HomeworkHow reinforcement learning can train a model on long-running agent tasks without resetting the work each step.
- 6Reading the ScoreboardHow K3 compares on public scoreboards, which scores matter, and where its tradeoffs remain visible.