A benchmark is a fixed set of tasks with a stated setup and scoring rule. It tells us how a model did on that test. It cannot collapse every kind of work into one number. Moonshot's paper says Kimi K3 “consistently outperforms other open and proprietary models” in most categories, “while its overall performance still trails the most powerful proprietary models.” The bars below unpack that sentence without pretending that one bar settles the question. Start with the card title. It tells you the kind of work being measured. Then read the unit. A percentage is a result on that particular task collection, not a percentage of all coding or reasoning a model can do. A higher percentage is useful within that card, but 88 on one benchmark and 63 on another do not share a difficulty scale. Next, ask who ran the comparison and under what conditions. The paper's tables are Moonshot-reported results. Its Kimi K3 evaluations use maximum reasoning effort, and its coding evaluations use named agent setups such as Kimi Code, Claude Code, or Codex. For Terminal-Bench 2.1, the paper reports the best score across those setups. The paper also flags potential fallbacks in Claude Fable 5 results and potential cyberguards in GPT-5.6 Sol results. Those details do not erase a score. They tell us what had to be held fixed before comparing it. Artificial Analysis (AA) is shown separately because it is an outside index with its own suite and measurement method. Read each row as a result under its named benchmark, effort level, and source, then use several rows before deciding what kind of work the model appears to handle well.
Long-horizon coding, the SWE-Marathon row
SWE-Marathon % (paper §6, self-reported, max reasoning effort)
- Kimi K3: open weights, downloadable
Terminal-Bench 2.1, a close cluster
% (paper §6)
ProgramBench, a relative weakness in this table
% (paper §6, 'trails the most powerful proprietary models', in their words)
- Kimi K3: 13+ points back on this one
Independent check and the output-speed tradeoff
Artificial Analysis (verified)
- AA Intelligence Index: #1 of 111 models tracked
- Output speed (output tokens per second): #56 of 111, smart but slow