Eleven models in eight weeks
Since Kimi K3, we have put eleven open-weight models in front of our users. The flash models caught the flagships, memory became the spec that matters, and coding models learned to see.
On July 31 we put Kimi K3 into production. At 2.8 trillion parameters, it is the biggest open model we have ever served. Eight weeks later, our default coding model is a flash model that puts 18 billion parameters to work on each token, and costs a tenth as much, or less.
In between, a new model reached our users about every five days. Eleven went through: ten from five labs (DeepSeek, Moonshot, Z.ai, Qwen and Xiaomi), and one we built ourselves. Six made it to production, and one of those retired a month later. Our production lineup today shares a single model with July’s.
Four things changed in those eight weeks.
1. The flash models caught up
Most labs now release two sizes: a flagship, and a smaller, faster model usually called flash. Both are mixtures of experts: for each token, only a slice of the network does the work. That slice, the active parameters, sets most of what a token costs to compute.
This summer, the gap between flash and flagship nearly closed. GLM 5.3 Flash activates 18 billion parameters per token, a sixth of Kimi K3’s 104 billion. On the three independent scores we track, it lands within about five points of K3: 42 against 44 on the Artificial Analysis Intelligence Index, 84.3 against 85 on Terminal-Bench 2.1, and 63.4 against 68.5 on DeepSWE. It costs a twentieth of K3’s price for input, and a thirtieth for output.
DeepSeek’s own lineup tells the same story. V4 Pro, a 1.6-trillion-parameter model, reached production in mid-August and retired a month later. By then, two flash models released after it scored higher, for less than a sixth of its price.
The flagships still own the hard end. Benchmarks aside, Kimi K3 is the model we find strongest on complex, multi-step work and security reviews, and GLM 5.3 has the highest score in our production lineup. But for everyday coding, the quality gap is down to a few points, while the price gap runs from nine to thirty times. In September, we moved umans-coder, our default, from a one-trillion-parameter Kimi to GLM 5.3 Flash.
2. Memory became the spec that matters
To write each new token, a model looks back over the whole session. To avoid reprocessing everything for every token, it keeps notes on what it has already read; engineers call them the KV cache. The notes live in the GPU’s own memory, which is fast, small and expensive. The more room one session’s notes take, the fewer sessions a GPU can serve at once, and the more each token costs.
All the models below advertise a context window of a million tokens. What it takes to actually hold one such session varies by a factor of 53:
On GLM 5.3, six of those sessions would fill an entire 288 GB B300 GPU before a single weight is loaded. On DeepSeek V4.1 Flash, more than 300 would fit. DeepSeek shrank its notes 3.5 times in six weeks, from V4 Flash to V4.1 Flash, partly by storing them in 4 bits, a format the new model is designed for. For the same session, Z.ai’s new flash base needs about an eighth of its flagship’s memory.
The ideas travel fast between labs. GLM 5.3 Flash mixes sparse and linear attention, and uses mHC, a technique from a DeepSeek paper. Qwen’s Flash Next preview adds sparse attention to its linear layers. Linear attention folds old context into a fixed-size summary; sparse attention keeps compact notes on everything and reads only the few that matter at each step. Our bet, for now, is on sparse: V4.1 Flash’s notes are about 16 times smaller per token than Kimi K3’s, which uses linear attention in most of its layers. Borrowed memory shows how we stretch GPU memory further on top of that.
3. Agents need eyes
In July, DeepSeek V4 Flash had one obvious gap: it could not see. The more autonomy you give a coding agent, the more it has to check its own work, and a lot of that work is visual: the page it just built, a failing UI test, a diagram in a design doc, a screenshot in a bug report. An agent that cannot see has to ask you to look.
So we gave V4 Flash eyes. We took Kimi K3’s vision encoder, the part that turns an image into something a language model can read, and connected it to V4 Flash through a small adapter. Both models stayed frozen. All the learning happened in the bridge between them, which we started from the words the two models already share rather than from random noise. Text requests skip the new parts entirely, so on text it is exactly V4 Flash. We published it on Hugging Face and opened a lab for it on August 11. It scores 54.6% on MMMU, a test of college-level questions about images, between two well-known small vision models, Pixtral 12B and Qwen2.5-VL 7B.
It was a research project, to see how far we could take a model we serve, and one of the most fun things we built this summer, training included. It also grew our stack: serving a model we changed ourselves, with images flowing through it, is a different job from serving one as released.
As far as we know, it was the first DeepSeek V4 that could see. Twenty days later, DeepSeek released its own vision version of V4 Flash, which we ran in Labs too, and ten days after that, V4.1 Flash shipped with vision built in. GLM 5.3 Flash is the first GLM-5 model that reads images. Today, four of our six production models do.
4. Training leaves fingerprints
The biggest quality jump we saw came from training, not size. GLM 5.3 sits on the same base model as GLM 5.2; only its post-training is new, the stage where a lab teaches a pretrained model to follow instructions, use tools and see long tasks through. That alone took its DeepSWE score, a benchmark of long, real coding tasks, from 44 to 69.
Training also leaves fingerprints. MiMo V2.6 Pro, from Xiaomi, was the pleasant surprise of September. It had the best Intelligence Index of any open model when we tested it, and our testers liked it a lot: thorough, effective, precise. It also shows what look like side effects of its agentic reinforcement learning: it thinks at length, second-guesses itself, and sometimes builds more than you asked for. One tester asked for four download links on a web page and got a complete downloads module, with throttling and analytics, tested end to end.
DeepSeek’s flash models have a fingerprint of their own. Once in a while, a long run falls into a loop and repeats itself until it runs out of room. Both V4 Flash and V4.1 Flash do it, V4.1 Flash far more often. We traced it to the model itself, not to how it is served, and our best guess is that it comes from training. If you hit one, restart the run, or use V4 Flash for long sessions.
Every lab, one line
| Model | Lab | Outcome | Our take |
|---|---|---|---|
| Kimi K3 | Moonshot | In production | The strongest on hard problems, in our experience. Also by far the biggest. |
| DeepSeek V4 Flash | DeepSeek | In production | Our most dependable everyday agent, and the cheapest model we serve. |
| Our V4 Flash Vision | Umans AI | Research model | Kimi K3’s eyes on V4 Flash, open on Hugging Face. |
| DeepSeek V4 Pro | DeepSeek | Retired after a month | Came too late: newer flash models passed it at a fraction of its price. |
| Qwen3.8 27B | Qwen | Not released | Testers loved it for its size, small enough to run at home. Dense and heavy on memory, it didn’t earn a slot next to the flash models. |
| GLM 5.3 Flash | Z.ai | In production | Our default, umans-coder: near-flagship scores at a flash price, and it sees. |
| Qwen3.8 Flash Next | Qwen | Not released | An early look at Qwen’s next architecture, impressive for its size. Testers hit mangled edits in its one day in Labs, and it stayed a preview. |
| GLM 5.3 | Z.ai | In production | Our highest score in production, for complex, long-horizon work. Text only. |
| DeepSeek V4 Flash Vision Exp | DeepSeek | Not released | DeepSeek’s first vision V4. V4.1 Flash replaced it within two weeks. |
| DeepSeek V4.1 Flash | DeepSeek | In production | Fast, sees, and the top agentic scores in our lineup, on DeepSeek’s own numbers. Loops more than V4 Flash. |
| MiMo V2.6 Pro | Xiaomi | Lab closed | Testers liked it a lot. Thinks long, and sometimes builds more than asked. |
What we would pick today
- Everyday coding: GLM 5.3 Flash, the model behind umans-coder.
- Long agent runs on a budget: DeepSeek V4 Flash. It is older than V4.1 Flash, and steadier.
- The hardest problems, especially with screenshots: Kimi K3.
Every model has a page with our take, its scores and its prices at app.umans.ai/models.
Two months ago, a coding model that sees, holds a million tokens in a few gigabytes of notes and costs well under a dollar per million output tokens was a wish list. Today we serve two. The next lab is never far off: we announce each one on Discord. Come break it with us.
Sources and further reading
- Our models and past models. Our take, scores and prices for every model we serve, and what we retired.
- DeepSeek-V4-Flash-0731-Vision. Our vision model on Hugging Face, with its MMMU result and limits.
- Artificial Analysis (Intelligence Index v4.3.2 and its Terminal-Bench 2.1 runs) and DeepSWE (v1.1). The independent scores quoted here, as of early October.
- kvcache.ai KV cache calculator. The memory figures: one sequence, notes in FP8 except DeepSeek V4.1 Flash, whose notes are 4-bit by design.
- Borrowed memory and Tokenomics, behind the scenes. How GPU memory and serving choices turn into the price of a token.