Eleven models in eight weeks

Since Kimi K3, we have put eleven open-weight models in front of our users. The flash models caught the flagships, memory became the spec that matters, and coding models learned to see.

On July 31 we put Kimi K3 into production. At 2.8 trillion parameters, it is the biggest open model we have ever served. Eight weeks later, our default coding model is a flash model that puts 18 billion parameters to work on each token, and costs a tenth as much, or less.

In between, a new model reached our users about every five days. Eleven went through: ten from five labs (DeepSeek, Moonshot, Z.ai, Qwen and Xiaomi), and one we built ourselves. Six made it to production, and one of those retired a month later. Our production lineup today shares a single model with July’s.

Timeline from July 20 to October 5, 2026. In production in July: Kimi K2.7 Code and GLM 5.2, both retired September 10, and Qwen3.6 35B, still in production. New since July 27: Kimi K3 (Labs July 27, production July 31), DeepSeek V4 Flash (Labs July 31, production August 3), our V4 Flash Vision (Labs August 11 to 13), DeepSeek V4 Pro (production August 15, retired September 14), Qwen3.8 27B (Labs August 18 to 21), GLM 5.3 Flash (Labs August 26, production September 9), Qwen3.8 Flash Next (Labs August 27 to 28), GLM 5.3 (production August 29), DeepSeek V4 Flash Vision Exp (Labs September 1 to 5), DeepSeek V4.1 Flash (Labs September 10, production September 18), MiMo V2.6 Pro (Labs September 22 to 24).

Four things changed in those eight weeks.

1. The flash models caught up

Most labs now release two sizes: a flagship, and a smaller, faster model usually called flash. Both are mixtures of experts: for each token, only a slice of the network does the work. That slice, the active parameters, sets most of what a token costs to compute.

This summer, the gap between flash and flagship nearly closed. GLM 5.3 Flash activates 18 billion parameters per token, a sixth of Kimi K3’s 104 billion. On the three independent scores we track, it lands within about five points of K3: 42 against 44 on the Artificial Analysis Intelligence Index, 84.3 against 85 on Terminal-Bench 2.1, and 63.4 against 68.5 on DeepSWE. It costs a twentieth of K3’s price for input, and a thirtieth for output.

DeepSeek’s own lineup tells the same story. V4 Pro, a 1.6-trillion-parameter model, reached production in mid-August and retired a month later. By then, two flash models released after it scored higher, for less than a sixth of its price.

Scatter of the Artificial Analysis Intelligence Index against the price of a million output tokens, on a log scale. Flash models: DeepSeek V4 Flash, 34 at 0.28 dollars; GLM 5.3 Flash, 42 at 0.50; DeepSeek V4.1 Flash, 39 at 0.60. Flagships: GLM 5.3, 45 at 4.40; Kimi K3, 44 at 15. Retired: DeepSeek V4 Pro, 36 at 3.96. GLM 5.3 costs 9 times GLM 5.3 Flash for 3 more points.

The flagships still own the hard end. Benchmarks aside, Kimi K3 is the model we find strongest on complex, multi-step work and security reviews, and GLM 5.3 has the highest score in our production lineup. But for everyday coding, the quality gap is down to a few points, while the price gap runs from nine to thirty times. In September, we moved umans-coder, our default, from a one-trillion-parameter Kimi to GLM 5.3 Flash.

2. Memory became the spec that matters

To write each new token, a model looks back over the whole session. To avoid reprocessing everything for every token, it keeps notes on what it has already read; engineers call them the KV cache. The notes live in the GPU’s own memory, which is fast, small and expensive. The more room one session’s notes take, the fewer sessions a GPU can serve at once, and the more each token costs.

All the models below advertise a context window of a million tokens. What it takes to actually hold one such session varies by a factor of 53:

Horizontal bars of GPU memory for the notes, or KV cache, of one 1-million-token session: DeepSeek V4.1 Flash 0.89 GB, DeepSeek V4 Flash 3.1 GB, DeepSeek V4 Pro 4.5 GB, GLM 5.3 Flash 6.1 GB, Kimi K3 14.3 GB, MiMo V2.6 Pro 25.6 GB, GLM 5.3 47.6 GB.

On GLM 5.3, six of those sessions would fill an entire 288 GB B300 GPU before a single weight is loaded. On DeepSeek V4.1 Flash, more than 300 would fit. DeepSeek shrank its notes 3.5 times in six weeks, from V4 Flash to V4.1 Flash, partly by storing them in 4 bits, a format the new model is designed for. For the same session, Z.ai’s new flash base needs about an eighth of its flagship’s memory.

The ideas travel fast between labs. GLM 5.3 Flash mixes sparse and linear attention, and uses mHC, a technique from a DeepSeek paper. Qwen’s Flash Next preview adds sparse attention to its linear layers. Linear attention folds old context into a fixed-size summary; sparse attention keeps compact notes on everything and reads only the few that matter at each step. Our bet, for now, is on sparse: V4.1 Flash’s notes are about 16 times smaller per token than Kimi K3’s, which uses linear attention in most of its layers. Borrowed memory shows how we stretch GPU memory further on top of that.

3. Agents need eyes

In July, DeepSeek V4 Flash had one obvious gap: it could not see. The more autonomy you give a coding agent, the more it has to check its own work, and a lot of that work is visual: the page it just built, a failing UI test, a diagram in a design doc, a screenshot in a bug report. An agent that cannot see has to ask you to look.

So we gave V4 Flash eyes. We took Kimi K3’s vision encoder, the part that turns an image into something a language model can read, and connected it to V4 Flash through a small adapter. Both models stayed frozen. All the learning happened in the bridge between them, which we started from the words the two models already share rather than from random noise. Text requests skip the new parts entirely, so on text it is exactly V4 Flash. We published it on Hugging Face and opened a lab for it on August 11. It scores 54.6% on MMMU, a test of college-level questions about images, between two well-known small vision models, Pixtral 12B and Qwen2.5-VL 7B.

It was a research project, to see how far we could take a model we serve, and one of the most fun things we built this summer, training included. It also grew our stack: serving a model we changed ourselves, with images flowing through it, is a different job from serving one as released.

As far as we know, it was the first DeepSeek V4 that could see. Twenty days later, DeepSeek released its own vision version of V4 Flash, which we ran in Labs too, and ten days after that, V4.1 Flash shipped with vision built in. GLM 5.3 Flash is the first GLM-5 model that reads images. Today, four of our six production models do.

4. Training leaves fingerprints

The biggest quality jump we saw came from training, not size. GLM 5.3 sits on the same base model as GLM 5.2; only its post-training is new, the stage where a lab teaches a pretrained model to follow instructions, use tools and see long tasks through. That alone took its DeepSWE score, a benchmark of long, real coding tasks, from 44 to 69.

Training also leaves fingerprints. MiMo V2.6 Pro, from Xiaomi, was the pleasant surprise of September. It had the best Intelligence Index of any open model when we tested it, and our testers liked it a lot: thorough, effective, precise. It also shows what look like side effects of its agentic reinforcement learning: it thinks at length, second-guesses itself, and sometimes builds more than you asked for. One tester asked for four download links on a web page and got a complete downloads module, with throttling and analytics, tested end to end.

DeepSeek’s flash models have a fingerprint of their own. Once in a while, a long run falls into a loop and repeats itself until it runs out of room. Both V4 Flash and V4.1 Flash do it, V4.1 Flash far more often. We traced it to the model itself, not to how it is served, and our best guess is that it comes from training. If you hit one, restart the run, or use V4 Flash for long sessions.

Every lab, one line

ModelLabOutcomeOur take
Kimi K3MoonshotIn productionThe strongest on hard problems, in our experience. Also by far the biggest.
DeepSeek V4 FlashDeepSeekIn productionOur most dependable everyday agent, and the cheapest model we serve.
Our V4 Flash VisionUmans AIResearch modelKimi K3’s eyes on V4 Flash, open on Hugging Face.
DeepSeek V4 ProDeepSeekRetired after a monthCame too late: newer flash models passed it at a fraction of its price.
Qwen3.8 27BQwenNot releasedTesters loved it for its size, small enough to run at home. Dense and heavy on memory, it didn’t earn a slot next to the flash models.
GLM 5.3 FlashZ.aiIn productionOur default, umans-coder: near-flagship scores at a flash price, and it sees.
Qwen3.8 Flash NextQwenNot releasedAn early look at Qwen’s next architecture, impressive for its size. Testers hit mangled edits in its one day in Labs, and it stayed a preview.
GLM 5.3Z.aiIn productionOur highest score in production, for complex, long-horizon work. Text only.
DeepSeek V4 Flash Vision ExpDeepSeekNot releasedDeepSeek’s first vision V4. V4.1 Flash replaced it within two weeks.
DeepSeek V4.1 FlashDeepSeekIn productionFast, sees, and the top agentic scores in our lineup, on DeepSeek’s own numbers. Loops more than V4 Flash.
MiMo V2.6 ProXiaomiLab closedTesters liked it a lot. Thinks long, and sometimes builds more than asked.

What we would pick today

  • Everyday coding: GLM 5.3 Flash, the model behind umans-coder.
  • Long agent runs on a budget: DeepSeek V4 Flash. It is older than V4.1 Flash, and steadier.
  • The hardest problems, especially with screenshots: Kimi K3.

Every model has a page with our take, its scores and its prices at app.umans.ai/models.

Two months ago, a coding model that sees, holds a million tokens in a few gigabytes of notes and costs well under a dollar per million output tokens was a wish list. Today we serve two. The next lab is never far off: we announce each one on Discord. Come break it with us.

Sources and further reading