Kimi · model № 003

Kimi K3

A 2.8T-parameter model that scales delta-rule attention, expert routing, and long-horizon training together

total params
2.8T
largest open ever
active / token
104B
16 of 896 experts
context
1M
agentic RL at full length
scaling efficiency
2.5×
over Kimi K2

Chapters

  1. 1The 2.8T StatementWhat a 2.8T-parameter open model is trying to prove, and where the public results support the claim.
  2. 2The Bet That PaidHow delta-rule attention edits a fixed memory state, and why Kimi scaled the idea to 2.8T parameters.
  3. 3Depth, RewiredHow residual attention paths help information survive when a model becomes very deep.
  4. 416 of 896How latent routing keeps expert traffic balanced without the usual auxiliary loss.
  5. 5Million-Token HomeworkHow reinforcement learning can train a model on long-running agent tasks without resetting the work each step.
  6. 6Reading the ScoreboardHow K3 compares on public scoreboards, which scores matter, and where its tradeoffs remain visible.