GLM · model № 001
GLM-5.3-Flash
A 320B model that serves 18B parameters per token and can run beyond one hardware ecosystem
total params
320B
experts on the shelf
active / token
18B
the only ones working
context
1M
tokens, one sitting
price in / out
$0.15/$0.50
per M tokens
Chapters
- 1Meet the FlashWhat this model is, what it costs, and which architectural choices make the price possible.
- 2The 18B Trick320B parameters are available, but only 18B run for a token. The router decides which ones.
- 3Attention, RewiredA running memory and a selective lookup divide the attention work without losing the whole context.
- 4The Memory SqueezeWhy the key-value cache grows with ordinary attention, and how Flash limits that growth.
- 5Two Tokens Per StepThe model drafts more than one next token, then verifies the draft before it commits.
- 6Anatomy of $0.15The active-parameter, attention, memory, and hardware decisions that add up to the serving price.
- 7Any Chip Will DoWhy the released weights can run across several chip stacks instead of one vendor's ecosystem.
- 8Does It Deliver?What the public scores measure, what they omit, and where Flash sits among comparable models.