GLM · model № 001

GLM-5.3-Flash

A 320B model that serves 18B parameters per token and can run beyond one hardware ecosystem

total params
320B
experts on the shelf
active / token
18B
the only ones working
context
1M
tokens, one sitting
price in / out
$0.15/$0.50
per M tokens

Chapters

  1. 1Meet the FlashWhat this model is, what it costs, and which architectural choices make the price possible.
  2. 2The 18B Trick320B parameters are available, but only 18B run for a token. The router decides which ones.
  3. 3Attention, RewiredA running memory and a selective lookup divide the attention work without losing the whole context.
  4. 4The Memory SqueezeWhy the key-value cache grows with ordinary attention, and how Flash limits that growth.
  5. 5Two Tokens Per StepThe model drafts more than one next token, then verifies the draft before it commits.
  6. 6Anatomy of $0.15The active-parameter, attention, memory, and hardware decisions that add up to the serving price.
  7. 7Any Chip Will DoWhy the released weights can run across several chip stacks instead of one vendor's ecosystem.
  8. 8Does It Deliver?What the public scores measure, what they omit, and where Flash sits among comparable models.