Kimi K3, Qwen 3.8, and a Week That Reset the Open-Weight Frontier
Canonical version: Kimi K3, Qwen 3.8, and a Week That Reset the Open-Weight Frontier.
July 2026 was a LOT. Two open-weight models at frontier scale were announced within three days of each other, a 27B model now runs on a phone, and there was plenty of tooling news.
Kimi K3: 2.8 trillion parameters
Moonshot AI released Kimi K3 on 16 July, with weights on Hugging Face by the 27th. 2.8 trillion parameters, 16 of 896 experts active per token, 1M context, natively multimodal.
Moonshot claims to trail only Fable 5 and GPT-5.6 Sol. Fireworks published an analysis titled "Kimi K3 is competitive with Fable", and several practitioners report that they run K3 alongside Claude on normal coding work and cannot tell them apart.
Compared with K2.6, nearly every conventional component was swapped for a cheaper one: Stable LatentMoE, multi-head latent attention, Kimi Delta Attention, and no positional embeddings at all. Models with similar designs usually keep RoPE in the local layers; Moonshot dropped it entirely, because the recurrent gating in KDA encodes position implicitly.
This undercuts the claim that Chinese open models are just distillations of Western ones.
Pricing is $3 / $15 per million tokens, the same as Anthropic's Sonnet series. That is high for a Chinese open-weight model, and it was the most debated number in the discussion threads. GLM 5.2 is under a third of the price.
Qwen 3.8, three days later
Alibaba announced Qwen 3.8 on 19 July: 2.4 trillion parameters, multimodal, claimed to be "second only to Fable 5".
What actually shipped is qwen3.8-max-preview, available through Alibaba's Token Plan and its Qoder coding platforms at 10% of standard pricing. As I write this, thirteen days later, there are still no weights, no model card, no license, and not a single published benchmark. Nothing from the Qwen org on Hugging Face since June. The active parameter count (the number that actually determines what a sparse MoE costs to run) is undisclosed.
If you see coverage saying the weights landed on 27 July, that's the Kimi K3 date being copied across. Alibaba said "soon" and named no date.
The announcement came three days after K3. Alibaba kept its flagship Max tier closed for two straight generations. 3.6-Max and 3.7-Max were API-only. Opening the biggest model it has ever trained would reverse that policy.
For now there is an announcement and no repository, so I wouldn't plan around it.
Both releases are a change of direction for Chinese labs, which built their reputation on small, fast, cheap, good-enough models.
Bonsai 27B runs on your phone
This is my favourite release of the month.
Bonsai 27B from PrismML is an extremely low-bit compression of Qwen3.6-27B, Apache 2.0. Two variants:
- Ternary at 1.71 effective bits per weight: 5.9 GB, retains 95% of the full-precision benchmarks
- 1-bit at 1.125 bits: 3.9 GB, retains 90%
For reference, the full-precision model is about 54 GB, and a normal 4-bit build is 18 GB.
That puts a 27B-class model, with reasoning, a 262K context and multimodal input, on an iPhone.
If you're building agents, note that tool calling degrades most (74.0 ternary, 66.0 for 1-bit, against an 80.0 baseline). Chat and reasoning hold up better; test tool calling before you put the model in a loop.
Gemini 3.6 Flash and 3.5 Flash-Lite
Gemini 3.6 Flash landed on 21 July at $1.50 / $7.50 per million tokens, 17% more token-efficient than 3.5 Flash, and better at coding and tool use. Alongside it: 3.5 Flash-Lite at $0.30 / $2.50, and a security specialist called Flash Cyber.
Google spent its earnings call being asked by four banks in a row what it plans to do about not having a state-of-the-art model. Each time, Google answered that everyone uses Flash anyway.
I think they're half right: Google's cloud revenue was up 82% year over year. But their announcement benchmarks these models only against previous Gemini versions, with no comparison to Claude, GPT, Kimi or GLM. For a release that argues on cost-effectiveness, that omission matters.
Cursor published the numbers on agent swarms
Cursor Agent Swarms documents implementing SQLite from scratch with different model combinations, and publishes the bill:
| Setup | Cost |
|---|---|
| Opus 4.8 planner + Composer 2.5 workers | $1,339 |
| GPT-5.5 for both roles | $10,565 |
GPT-5.5 workers alone cost $9,373. Opus planning plus cheap workers cost $411 for the same layer. The worker tier is where the money goes, and it's the tier where model choice matters least.
They also compared their old orchestration to the new one, with the same model and the same task:
- Commits: 68,000 → 1,000
- Merge conflicts: 70,000+ → under 1,000
- Final engine code: 64,305 lines → 9,908 lines
The new orchestration produced about six times less code for the same working result. Every fix they applied is topological (impartial reconciler agents, decorrelated reviewers, shared design docs), which matches what graph engineering predicts.
DSLs are how you make LLMs reliable
Unmesh Joshi published a piece on Martin Fowler's site: DSLs Make LLM Output Reliable.
Joshi argues that a general-purpose language offers a hundred valid ways to express one intent, and a DSL removes that variation. Choices the language already makes are no longer left to the model. Add a validator (a parser, a schema, a compiler) and your agent can self-correct without you.
I'd add that a DSL also works as the handoff contract between agent steps. Pipe natural language through five model calls and the fifth receives something quite different from what its prompt expects. Make the DSL the contract at each boundary and the drift stops compounding.
I also wrote up the concept itself in Domain Specific Languages (DSLs), with book references.
The tooling round-up
- JetBrains Context: a semantic index of your repositories that agents query by concept instead of grepping around. Claims up to 68% fewer agent turns and 48% lower cost. Works with Claude Code and OpenAI Codex, not just Junie, so JetBrains treats the index as the durable asset and the agent as swappable
- GitHub Copilot Vision went generally available on 1 July. Attach images and PDFs to a Copilot chat, on every plan including Free
- LM Studio Bionic: an agent app built specifically for open models. Skeptics point out that every open harness already talks to LM Studio. The open question is whether a harness tuned for small local models can make local agents useful daily
- NotebookLM is now Gemini Notebook. Same product, folded into the Gemini brand
- FFASR Leaderboard: the first open far-field speech recognition benchmark. Word error rate in a real room at low signal-to-noise is several times higher than in a clean booth
- Codex Micro: OpenAI made a $230 macro pad with a dial for reasoning effort and RGB keys showing live agent status. It's a rebranded Work Louder pad, and I'd buy a Stream Deck instead. The needs it targets make sense: seeing what every agent is doing, approving or rejecting fast, adjusting how hard the model thinks
What I take away from all this
Open-weight models now reach the frontier, and they are no longer cheap: K3 costs the same as Sonnet.
The money goes to the worker tier. In Cursor's test, the worker layer cost $9,373 with GPT-5.5 and $411 with Composer 2.5. If you run agents and haven't separated planning from execution, I'd start there. Simon Willison's Fable 5 follow-up suggests delegating the decision about which model does the work, too.
Local models keep getting more capable: Bonsai runs a 27B model on a phone. More tasks that used to need a frontier model now run on local hardware, so I'd treat the local tier as a real option rather than a fallback.
Related
- Kimi K3
- Qwen 3.8
- Gemini 3.6 Flash
- Bonsai 27B
- Cursor Agent Swarms
- Codex Micro
- LM Studio Bionic
- JetBrains Context
- GitHub Copilot Vision
- FFASR Leaderboard
- NotebookLM
- DSLs Make LLM Output Reliable
- Domain Specific Languages (DSLs)
- Graph Engineering
- Moonshot AI
- Qwen
- Claude Fable 5
- GPT-5.6
- AI Open Weight Models
- Running AI Models Locally
- AI Agent Swarms
- Agentic Engineering
- Simon Willison
About Sébastien
Ready to get to the next level?
Found this valuable? Share it with someone who needs it.