The All In models we recommend, and the ones that burn the allowance
Sonnet 5, GPT-5.6 Terra, Luna, Kimi K2.7 Code, Kimi K3, and Qwen 3.8 Max are now on every All In plan. Use the cheap lane first.
Every All In plan now includes Claude Sonnet 5, GPT-5.6 Terra, GPT-5.6 Luna, Kimi K2.7 Code, Kimi K3, and Qwen 3.8 Max. That is the interesting part of the picker, not the default. These routes draw from the same monthly inference dollars as DeepSeek V4.1 Flash. The expensive ones will empty that allowance while you still have cloud sessions left.
Start on DeepSeek V4.1 Flash. Move to GLM-5.3 Flash when the change needs more shape. Use GLM-5.3 when the work is actually hard. Keep Terra, Sonnet 5, Kimi K3, and Qwen 3.8 Max for the jobs that earn the burn.
/code, choosing Sonnet 5, Terra, Kimi K3, or Qwen 3.8 Max shows a warning: those models use All In inference faster than the recommended lineup. Luna and K2.7 Code stay in the efficient band. The warning is there because the first cap you hit is usually dollars, not sessions.The lineup we actually want you on
All In $10 includes $20 of inference at provider cost. Counting prompts does not explain why one afternoon of Sonnet 5 can cost more than a week of Flash. The recommended stack is the one that still leaves room for a second and third session.
| Model | Use it for | All In burn | |
|---|---|---|---|
| DeepSeek V4.1 Flash | DeepSeek V4.1 Flash | Default author. Fast edits, tests, and everyday hosted sessions. | Lowest |
| GLM-5.3 Flash | GLM-5.3 Flash | The step up when Flash-class DeepSeek needs more coding shape, still cheap. | Low |
| GLM-5.3 | GLM-5.3 | Hard coding. People keep comparing it to Opus 5; a lot of them prefer GLM and it costs a lot less. | Low to mid |
| GPT-5.6 Luna | GPT-5.6 Luna | Planning and decomposition. More efficient than Terra, still an OpenAI reasoning model. | Mid |
| Kimi K2.7 Code | Kimi K2.7 Code | Coding-first Moonshot lane. Long-horizon implementation without jumping to K3. | Mid |
| GPT-5.6 Terra | GPT-5.6 Terra | Strong OpenAI model when Luna is not enough. Burns faster. | High |
| Kimi K3 | Kimi K3 | Planning and long-horizon reasoning. Pair with K2.7 Code, do not live here. | High |
| Claude Sonnet 5 | Claude Sonnet 5 | When you specifically want Sonnet. It is now on All In, and it is not cheap. | High |
| Qwen 3.8 Max | Qwen 3.8 Max | Alibaba’s higher-end Max route. Useful, faster to drain the allowance. | High |
GLM-5.3 is the Opus-class coding model most people can actually afford
Opus 5 remains on Founders Edition. GLM-5.3 is on every All In plan. The comparison people keep making is not subtle: for coding agents, a lot of operators now treat GLM-5.3 as the model that does Opus-5 work at a price that does not end the month. We are not going to dress that up as a private leaderboard. It is what we hear, and it matches how the credit floors sit (GLM-5.3 at 3, Opus 5 at 30).
GLM-5.3 Flash is the cheap sibling. Use it when you want GLM’s coding habits without paying the full GLM-5.3 token burn. DeepSeek V4.1 Flash stays the default because it is the fastest way to get a hosted session moving.
Terra is strong. Luna is the OpenAI model you should live on.
GPT-5.6 Terra is a serious model. If you want OpenAI at full strength, it is in the picker now. It also drains All In inference quickly, which is why /code warns when you select it.
GPT-5.6 Luna is the efficient OpenAI route. It is already good at planning: break the task up, name the files, decide the order, then hand implementation to a cheaper coder. That split is how you keep speed without lighting the dollar cap.
K2.7 Code writes. K3 plans.
Kimi K2.7 Code is the coding specialist. Keep it on the implementation loop. Kimi K3 is the larger planning model. Using K3 to think and K2.7 Code to edit is the Moonshot version of Luna-then-Flash: more reasoning where it pays, more throughput where it does not.
K3 used to sit on Founders Edition. It is now on All In $10 and $30. That is the whole reason the warning exists. A planning model at K3’s rate can finish the monthly inference allowance in a handful of long sessions.
Reasoning is now a slider, not a hidden default
Hosted Code now has a reasoning control next to the model picker: Low → Medium → High → Extra high. OpenAI and Anthropic routes expose Extra high. Everyone else stops at High. Default is still Medium.
Turn it up when the task is ambiguous. Leave it at Medium for ordinary edits. Extra high on Terra or Sonnet 5 is how you discover the inference cap before lunch.