<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Umans AI Blog</title><description>Notes on collaborative AI teammates, consistency, and the tooling that keeps teams aligned.</description><link>https://blog.umans.ai/</link><item><title>Eleven models in eight weeks</title><link>https://blog.umans.ai/blog/eleven-models-in-eight-weeks/</link><guid isPermaLink="true">https://blog.umans.ai/blog/eleven-models-in-eight-weeks/</guid><description>Between late July and late September, a new open-weight model reached our users about every five days. Flash models caught the flagships, memory became the spec that matters, and coding models learned to see. What we shipped, what we turned down, and why.</description><pubDate>Mon, 05 Oct 2026 00:00:00 GMT</pubDate></item><item><title>The token gap</title><link>https://blog.umans.ai/blog/the-token-gap/</link><guid isPermaLink="true">https://blog.umans.ai/blog/the-token-gap/</guid><description>Last October, ten billion lifetime tokens earned developers a public tribute at OpenAI DevDay. This June, one Umans AI user ran nearly eleven billion tokens in a single day. What that says about token demand, compute supply, and the serving efficiency that has to close the gap between them.</description><pubDate>Tue, 25 Aug 2026 00:00:00 GMT</pubDate></item><item><title>Tokenomics, behind the scenes</title><link>https://blog.umans.ai/blog/tokenomics-behind-the-scenes/</link><guid isPermaLink="true">https://blog.umans.ai/blog/tokenomics-behind-the-scenes/</guid><description>Put the exact same model on GPUs and its token production cost can vary by multiples depending on how you serve it. A walk through the economics of a token: the frontier, the dollars it turns into, the levers that move it, and the claim our measurements let us make.</description><pubDate>Sat, 15 Aug 2026 00:00:00 GMT</pubDate></item><item><title>DeepSeek V4 Pro DSpark: the model isn&apos;t ready, the architecture is</title><link>https://blog.umans.ai/blog/deepseek-v4-pro-dspark-the-architecture-is-ready/</link><guid isPermaLink="true">https://blog.umans.ai/blog/deepseek-v4-pro-dspark-the-architecture-is-ready/</guid><description>We served DeepSeek V4 Pro for two days. What we wanted was hands on the architecture: attention that scales almost linearly instead of quadratically, and a speculative decoder that adapts to load. Both delivered. The preview checkpoint did not.</description><pubDate>Sat, 25 Jul 2026 00:00:00 GMT</pubDate></item><item><title>GLM-5.2 NVFP4: fast, cheap, and not worth serving</title><link>https://blog.umans.ai/blog/glm-5-2-nvfp4-not-worth-serving/</link><guid isPermaLink="true">https://blog.umans.ai/blog/glm-5-2-nvfp4-not-worth-serving/</guid><description>We put an NVFP4 build of GLM-5.2 in front of real users for four days, at 200+ tokens per second. Users called it intoxicating. We retired it anyway: if the tokens are not useful, they are not worth serving, no matter how efficiently we can produce them.</description><pubDate>Mon, 13 Jul 2026 00:00:00 GMT</pubDate></item><item><title>How We Actually Ship with AI</title><link>https://blog.umans.ai/blog/how-we-ship-with-ai/</link><guid isPermaLink="true">https://blog.umans.ai/blog/how-we-ship-with-ai/</guid><description>The concrete workflow we use to go from a vague idea to code running in production, with AI agents doing the heavy lifting. Three moves, one increment file, everything in git.</description><pubDate>Mon, 30 Mar 2026 00:00:00 GMT</pubDate></item><item><title>Stop Being Your AI Agent&apos;s Assistant</title><link>https://blog.umans.ai/blog/stop-being-your-ai-agents-assistant/</link><guid isPermaLink="true">https://blog.umans.ai/blog/stop-being-your-ai-agents-assistant/</guid><description>When we launched code.umans.ai at the beginning of February, my cofounder Naji and I saw it as an opportunity to challenge ourselves. We had unlimited LLM access now. No more token anxiety. So we asked: what happens if we really let our agents work autonomously?</description><pubDate>Thu, 26 Feb 2026 00:00:00 GMT</pubDate></item><item><title>GLM-5 vs Kimi-K2.5: Long-context serving at scale</title><link>https://blog.umans.ai/blog/glm-5-vs-kimi-k2-5-long-context-serving/</link><guid isPermaLink="true">https://blog.umans.ai/blog/glm-5-vs-kimi-k2-5-long-context-serving/</guid><description>Why GLM-5 is a real step forward compared to GLM-4.7 for serving long-context coding agents, and why we still keep Kimi-K2.5 as the default for the best experience.</description><pubDate>Thu, 12 Feb 2026 00:00:00 GMT</pubDate></item><item><title>Shipping Solo With AI Agents Without Reading the Code</title><link>https://blog.umans.ai/blog/building-software-iteratively-alone-with-ai-in-jan-2026/</link><guid isPermaLink="true">https://blog.umans.ai/blog/building-software-iteratively-alone-with-ai-in-jan-2026/</guid><description>A new abstraction is emerging for solo building: iterate on intent with an AI agent, run, refine, and read little code. What enables it now, and where it breaks.</description><pubDate>Fri, 02 Jan 2026 00:00:00 GMT</pubDate></item><item><title>The Claude Code experience, self-hosted</title><link>https://blog.umans.ai/blog/host-claude-code/</link><guid isPermaLink="true">https://blog.umans.ai/blog/host-claude-code/</guid><description>How to run Claude Code against a self-hosted DeepSeek V3.2–class model using vLLM + LiteLLM, so agentic coding stays inside your perimeter.</description><pubDate>Mon, 29 Dec 2025 00:00:00 GMT</pubDate></item><item><title>Can coding agents actually follow your codebase’s rules?</title><link>https://blog.umans.ai/blog/agent-apply/</link><guid isPermaLink="true">https://blog.umans.ai/blog/agent-apply/</guid><description>We ran a reality check to see how well different coding agents follow a repo’s AGENTS.md in practice.</description><pubDate>Sat, 29 Nov 2025 00:00:00 GMT</pubDate></item><item><title>Your AI Coding Agent Doesn’t Care About Your Conventions</title><link>https://blog.umans.ai/blog/consistency-matters/</link><guid isPermaLink="true">https://blog.umans.ai/blog/consistency-matters/</guid><description>Exploring the challenges and strategies for maintaining codebase consistency in the age of AI coding agents.</description><pubDate>Mon, 10 Nov 2025 00:00:00 GMT</pubDate></item></channel></rss>