Stop putting Kimi on a rename
critique/auto starts cheap, stays sticky, and escalates once if the loop is stuck.
Set model to critique/auto. The first turn picks an efficient Inference specialist. Later turns reuse it. If tools keep failing, the session moves once onto a frontier coding model. Credits bill the routed id, not the word auto.
Most “auto” routers are not coding products. A generic market-spend auto route is a seven-day popularity index over mixed traffic. That is not “what should run this agent turn.”
The other failure mode is worse: pin the strongest model you sell, because you are afraid of looking slow. Then a one-line rename and a tenancy design review pay the same bill. Factory’s public router write-up is honest about that curve. The savings only exist on the flat part: cheaper models take the work they can finish, and frontier stays on the work they cannot.
Critique Auto is that policy on the Inference API. It picks from a fixed coding portfolio — listed Inference specialists plus a few catalog ids we scored for SWE work (Hy4, GLM-5.3 Flash, Mercury 2.5, GPT-5.6 Luna on efficient; GPT-5.6 Sol and Claude Opus 5 on frontier). It will not search OpenRouter’s full catalogue. If you want a model outside the rungs, pin it.
Two rungs, one session
A new critique/auto session starts on the efficient rung: Kimi K2.7 Code, MiniMax M3, KAT-Coder Pro, GLM-5.3 Flash, Mercury 2.5, Hy4, GPT-5.6 Luna, Seed 2.0 Code, DeepSeek V4 Flash 0731, Gemini 3.8 Flash, Qwen3.8 Flash, Muse Glimmer, Qwen3.8 27B, Laguna S 2.1, and Ling 3.0 Flash. That is the default for edits, FIM, screenshots that a VLM can see, and healthy tool loops.
critique/auto-best starts on the frontier rung: GPT-5.6 Sol, Claude Opus 5, Kimi K3, Muse Spark 1.3, GLM-5.3, Qwen3.8 2.4T, Grok 4.6, and Fugu Ultra. Trinity and Kimi K2.6 are out of Auto. Kimi K2.7 Code, MiniMax M3, and KAT-Coder Pro sit on efficient now.
The first response includes X-Critique-Router-Session. Send it back. Auto reuses the same specialist and the same OpenRouter provider so the prefix cache is not a casualty of cleverness. Switching models every message is how a router loses money while claiming to save it. If a client drops the header, Auto still recovers the session from a fingerprint of the user, first prompt, and tools.
Escalate when the cheap model is actually failing
If the same tool loop throws twice, or the transcript is already long and still erroring, Auto moves that session onto the frontier rung once. It does not thrash. A Flash classifier may help pick *inside* a rung. It is not allowed to promote Kimi on turn one of a healthy critique/auto session. That would recreate “everything is Opus.”
Hard vetoes still apply. A screenshot cannot land on a text-only model. A million-token dump cannot land on a 128K specialist. An agent schema cannot land on a model we score as a weak tool caller.
What we are not claiming
We are not publishing a Terminal-Bench 2 number until we have run that slice against pinned Seed and pinned Kimi on the same tasks. A router without cost-per-successful-session is a vibe. Auto is the policy we will measure. The Inference dashboard already records the routed model, the session id, and whether the session escalated.
Point the harness at Auto
Echo X-Critique-Router-Session on later turns of the same agent run.
curl https://critique.sh/api/v1/chat/completions \
-H "Authorization: Bearer crt_YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "critique/auto",
"messages": [{ "role": "user", "content": "Ship the billing pause flow" }],
"tools": [{ "type": "function", "function": { "name": "read_file" } }]
}'