Skip to content
Skip to content
Models & pricing / September 202628 min read

Critique Auto Gets a New Model Roster

A practical guide to the refreshed Critique Auto portfolio, current token rates, one shared benchmark view, and hosted coding in the CritiqueCode app.

Critique
The short version

Critique Auto now starts with current efficient coding specialists, keeps a model sticky through a healthy session, and escalates once when the loop is genuinely struggling. The refreshed roster includes DeepSeek V4 Flash 0731, Gemini 3.8 Flash, Qwen3.8 Flash, Muse Spark 1.3, GLM-5.3, Kimi K3, Grok 4.6, Fugu Ultra, and the latest direct Inference API offers. In the CritiqueCode app, every user also gets four hosted routes with a shared weekly allowance of 15M input and 5M output tokens. Accounts with review credits unlock the broader managed model shelf shown on Pricing and Models.

0
models on the managed catalog
0
hosted routes in CritiqueCode
0M
shared weekly input tokens
0M
shared weekly output tokens

Model catalogs age in a surprisingly specific way. The router still works, but the names around it start telling the wrong story: a “3.7” label that should be “3.8,” a former “Flash” route that has become a dated alias, or a transport suffix that leaks into a product selector where a person only needs a clear name. That drift is costly because model choice is part of the product contract. It affects latency expectations, the amount of context a turn can accept, token pricing, benchmark interpretation, and the fallback a coding agent reaches when its first plan does not survive contact with a repository.

This update brings the current model names into Critique Auto, the direct Inference API offer list, the hosted CritiqueCode selector, Remedy, the model catalog, and Pricing. Internal route identifiers remain implementation details. The UI shows the model name, provider mark, access level, and the relevant credit or token-rate context. Existing stored aliases still resolve forward, so a policy that references a previous release does not become a dead configuration overnight.

Auto is a session policy, not a random model carousel. On the first turn, Critique reads the request shape: tool use, editing, completion, long context, vision, video, reasoning depth, and the expected agent horizon. It ranks models that can actually accept that shape. Once a model is selected, the session keeps it sticky while the loop is healthy. Stickiness gives the agent a stable behavioral surface and preserves the practical value of context and cached prefixes.

Healthy Auto session
Read task shapePick an efficient specialistKeep the model stickyReturn the reviewed result
When the loop is stuck
Observe repeated tool failureEscalate once to the frontier rungKeep the same sessionStop escalating and preserve evidence

The refresh moves the Auto portfolio forward without changing that policy. Gemini 3.7 Flash is now Gemini 3.8 Flash. Qwen3.7 Flash is now Qwen3.8 Flash. DeepSeek V4 Flash is now DeepSeek V4 Flash 0731. Muse Spark 1.2 is now Muse Spark 1.3. GLM 5.2 is now GLM-5.3. Grok is on 4.6. The frontier side also includes DeepSeek V4 Pro 0813, Kimi K3, Kimi K2.7 Code, Qwen3.8 2.4T, MiniMax M3, KAT Coder Pro v2.5, Trinity Large Thinking, Kimi K2.6, Fugu Ultra, and the refreshed specialist set.

The important behavior is the boundary around those names. Auto chooses from Critique-resold Inference API offers. It does not silently pull an arbitrary model from the managed review catalog. That is intentional: an Auto request should have a known route, a known billing surface, and an explainable fallback list. If you want a specific managed model, pin it from the Models shelf. If you want Auto, let the session policy do the selection and inspect the routed model in the response metadata.

Critique.shLive · Updated just now

The current Auto vocabulary

Names shown in the product. The selector does not expose raw transport route IDs or provider-specific suffixes.

10 models88.3% top benchmark score

Kimi K3

Terminal-Bench 2.1

88.3%

GLM-5.3

Terminal-Bench 2.1

81.0%

DeepSeek V4 Flash 0731

SWE-Bench Verified

79.0%

Gemini 3.8 Flash

Terminal-Bench 2.1

78.0%

Qwen3.8 Flash

Benchmark

API rate $0.15 / $0.47

Muse Spark 1.3

Benchmark

API rate $1.25 / $4.25

Grok 4.6

Benchmark

API rate $2 / $6

Fugu Ultra

Benchmark

API rate $5 / $30

Seed 2.0 Code

Benchmark

API rate $0.50 / $3

Hy4 Preview

Benchmark

API rate $0.834 / $2.501

SWE-bench scores reflect best observed performance on the toughest real-world coding tasks.

All scores are relative.

There is a difference between an Auto candidate and a catalog listing. A candidate is scored against a task vector and is eligible for a specific rung. A catalog model is a product offer with a credit floor, role support, and a published display identity. The two sets overlap, but they are not the same list. Keeping that distinction visible prevents a common billing surprise: a user thinks “Auto” means “any model in the product,” then discovers that a paid managed model was never eligible for the direct Auto policy in the first place.

The table below is a directory snapshot of provider-reported input and output rates, normalized to dollars per one million tokens. It is here to make model selection legible, not to imply that a token rate predicts successful repository work. Output is often the expensive side of an agent loop; context length is not the same thing as useful context; and a low per-token price can still lose if the model needs repeated repairs. The managed credit floor shown elsewhere in Critique is the product rate for a review run, while these numbers are the underlying token-rate reference for direct model calls.

Provider rate snapshot · USD per 1M tokens

New names, visible economics

Human-facing model names with the current directory rate, context class, and the access path inside Critique.

DeepSeek V4 Flash 0731
Inference API · Auto
Input / 1M
$0.065
Output / 1M
$0.18
Context 1.31M

Efficient default for healthy coding loops.

DeepSeek V4 Pro 0813
Inference API · frontier
Input / 1M
$1.1154
Output / 1M
$3.3462
Context 1.05M

Reasoning escalation for difficult repairs.

DeepSeek V4 Flash Vision Exp
Managed credits
Input / 1M
$0.44
Output / 1M
$1.32
Context 1.05M

Vision-specialist route; published rates can vary by window.

Gemini 3.8 Flash
Inference API · Auto
Input / 1M
$0.75
Output / 1M
$3.75
Context 1.05M

Multimodal candidate for visual and long-context work.

Qwen3.8 Flash
Inference API · Auto
Input / 1M
$0.15
Output / 1M
$0.47
Context 1M

Low-cost throughput specialist.

Qwen3.8 Max
Managed credits
Input / 1M
$2
Output / 1M
$6
Context 1M

Higher-capability catalog route for deliberate pinning.

GLM-5.3
Inference API · frontier
Input / 1M
$1.40
Output / 1M
$4.40
Context 1.31M

Frontier agentic route with a strong terminal signal.

GLM-5.3 Flash
Managed credits
Input / 1M
$0.075
Output / 1M
$0.25
Context 1.31M

Fast GLM lane for cost-sensitive work.

Muse Spark 1.3
Inference API · Auto
Input / 1M
$1.25
Output / 1M
$4.25
Context 1.05M

Multimodal reasoning and agentic context.

Kimi K3
Inference API · frontier
Input / 1M
$3
Output / 1M
$15
Context 1.05M

Long-horizon specialist; expensive output rewards selectivity.

Hy4 Preview
Managed credits
Input / 1M
$0.834
Output / 1M
$2.501
Context 1.05M

Preview route for tool-heavy agent work.

Seed 2.0 Code
Inference API · Auto
Input / 1M
$0.50
Output / 1M
$3
Context 262K

Code-focused route with a smaller context class.

Fugu Ultra
Inference API · frontier
Input / 1M
$5
Output / 1M
$30
Context 1M

Base rates; higher context pricing may apply after 272K.

Mercury 2.5 Preview
CritiqueCode hosted
Input / 1M
Hosted usage
Output / 1M
Hosted usage
Context 260K

Hosted route; Critique usage is billed through the hosted product.

Laguna S 2.1
CritiqueCode hosted
Input / 1M
Hosted usage
Output / 1M
Hosted usage
Context 262K

Hosted route; Critique usage is billed through the hosted product.

MiniMax M3
CritiqueCode hosted
Input / 1M
Hosted usage
Output / 1M
Hosted usage
Context 1.05M

Hosted route; the managed catalog also has a credit path.

Claude Fable 5.1
Managed credits
Input / 1M
$10
Output / 1M
$50
Context 1M

Updated from Fable 5 in the managed catalog.

Rates are a point-in-time directory snapshot and can change independently of Critique credit floors. Exact context thresholds and special routing windows belong on the live model page. Prices are shown without raw provider route identifiers so the product name stays stable.

A model catalog can contain dozens of benchmark claims and still have no clean comparison. The answer is not to mash every score into one leaderboard. For this roster, the cleanest shared coding benchmark is SWE-bench Verified: real GitHub issues, a recognizable unit, and enough overlap across the catalog to show a useful shape. The chart includes only models for which the catalog has a published SWE-bench Verified reading. It intentionally leaves out newer models whose available public scores use another suite, a different harness, or no comparable score at all.

The useful frontier lives in the lower-left

Published SWE-bench Verified score against Critique credits per run. Lower cost is left; higher score is up.

SWE-bench Verified score compared with Critique credits per runEach dot is a model with a published SWE-bench Verified score. Lower on the horizontal axis means fewer Critique credits per run.60%70%80%90%05102040Critique credits / run →SWE-bench VerifiedStepFun 3.7 Flash: 74.4% at 1 creditsKAT Coder Pro v2.5: 69.4% at 2 creditsTrinity Large Thinking: 63.2% at 1 creditsDeepSeek V4 Flash 0731: 79.0% at 0.5 creditsDeepSeek V4 Pro 0813: 80.6% at 1 creditsQwen3.7 Plus: 78.8% at 1.5 creditsQwen3.8 Max: 80.4% at 6 creditsMiMo v2.5 Pro: 78.9% at 1 creditsKimi K2.6: 80.2% at 4 creditsClaude Sonnet 5: 85.2% at 15 credits
StepFun 3.7 Flash
74.4% · 1 cr
KAT Coder Pro v2.5
69.4% · 2 cr
Trinity Large Thinking
63.2% · 1 cr
DeepSeek V4 Flash 0731
79.0% · 0.5 cr
DeepSeek V4 Pro 0813
80.6% · 1 cr
Qwen3.7 Plus
78.8% · 1.5 cr
Qwen3.8 Max
80.4% · 6 cr
MiMo v2.5 Pro
78.9% · 1 cr
Kimi K2.6
80.2% · 4 cr
Claude Sonnet 5
85.2% · 15 cr

Scores are published vendor/model-card readings in the Critique catalog, not a Critique-run head-to-head. The test harness, agent scaffold, trial count, and date can differ. Use the chart to see trade-offs and shortlist a run; use your own repository evaluations to decide.

The lower-left is where Auto earns its keep. DeepSeek V4 Flash 0731 is a good example of the shape we want from an efficient rung: a sub-credit floor in the managed catalog, a strong shared benchmark reading, and a context window large enough for realistic repository prompts. The upper-right is still important. Claude Sonnet 5 and Qwen3.8 Max are not “bad values” because they cost more; they are deliberate escalation or pinned-choice models when the expected cost of another failed loop is higher than the token bill. The point is to make that trade visible before the run starts.

Notice what the chart does not say. It does not rank Gemini 3.8 Flash below GLM-5.3 because one has a SWE-bench Pro row and the other has a Terminal-Bench row. It does not convert Kimi K3’s Terminal-Bench 2.1 score into a fake SWE-bench value. A benchmark visual is only useful when its axes share a definition. Missing data is a meaningful state, not an invitation to interpolate.

CritiqueCode is Critique’s author harness: an interactive coding agent that implements in a session, then enters the existing review and verified-repair machinery. The CritiqueCode app is the hosted version of that harness. It gives the author an isolated worker, a repository branch, a live session, and a model selector that is honest about what can start on the current deployment. The four hosted routes are available directly inside that app.

Critique.shLive · Updated just now

Four hosted routes, one weekly allowance

Every user sees names and provider marks. The UI does not expose raw route IDs or provider-specific suffixes.

4 models79.0% top benchmark score

DeepSeek V4 Flash 0731

SWE-Bench Verified

79.0%

Mercury 2.5 Preview

Benchmark

CritiqueCode hosted · 15M input / 5M output per week

Laguna S 2.1

Benchmark

CritiqueCode hosted · 15M input / 5M output per week

MiniMax M3

Benchmark

CritiqueCode hosted · 15M input / 5M output per week

SWE-bench scores reflect best observed performance on the toughest real-world coding tasks.

All scores are relative.

The hosted allowance is shared across those four routes for each account: 15 million input tokens and 5 million output tokens per week. It resets weekly for all users. A model switch does not reset the meter, and opening another session does not create another allowance. That shared pool makes the offer predictable and prevents a user from multiplying capacity just by starting parallel sessions. The hosted page shows the model name, access level, and deployment capability; provider credentials and sandbox policy remain server-side concerns.

If you need more than the weekly hosted allowance, add Inference API credits to the account. Credits unlock the latest managed catalog on the account’s model selector, including the current model pages and the new routes listed above. That is the paid path for teams that want to pin a stronger model, run more review volume, or make a deliberate choice outside the four hosted routes. The credit shelf is there when the workload becomes a real operating lane.

Decision guide

Pick the control surface that matches the job

There is no virtue in hiding a product choice behind a single “best model” badge. Use the least complicated control that gives the repository enough capability and the team enough predictability.

Metric
When it fits
What you get
Critique Auto
Mixed coding work and changing task shapes
Efficient first, sticky session, one evidence-based escalation
Pinned Inference route
Stable evals, reproducible agent behavior, direct API calls
A named route and its current token rate
Managed model with credits
Review volume or a deliberate high-capability choice
The latest Models/Pricing shelf and its per-run credit floor
CritiqueCode hosted
Using the author harness inside the hosted app
Four hosted routes with a shared weekly allowance

For a normal change request, start with Auto and let the session stay coherent. For a benchmark or a regression investigation, pin a route so the comparison has a stable independent variable. For large review programs, use the managed model shelf and treat credits as a budget you can reason about per run. For a first CritiqueCode session, use one of the hosted routes and learn the harness contract before deciding which paid model deserves a permanent place in your workflow.

The UI is intentionally name-first. “GLM-5.3 Flash” is useful to a person; a provider-prefixed route with a suffix is useful to a transport layer. Both can exist, but they should not compete for the same visual space. LobeHub provider marks now cover the refreshed catalog in the hosted selector, model pages, legacy Remedy surfaces, and the article components. Sakana has a small matching local mark because that provider is not present in the icon package; it follows the same sizing and accessibility contract as the other marks.

Old IDs remain aliases where a forward mapping is safe: Fable 5 resolves to Fable 5.1, Gemini 3.6 and 3.7 resolve to Gemini 3.8 Flash, Grok 4.3 and 4.5 resolve to Grok 4.6, GLM 5.1 and 5.2 resolve to GLM-5.3, and the previous DeepSeek Flash names resolve to V4 Flash 0731. The removed base MiMo v2.5 is not advertised as a current route; MiMo v2.5 Pro remains available. Aliasing protects stored policies without allowing stale names to leak back into the current product shelf.

A price page answers “what will this route cost under this product’s billing rule?” A provider directory answers “what token rate was reported for this model at the time of the snapshot?” A benchmark card answers “how did this model score in this suite, under this harness, on this date?” Those are three different facts. We show them together because engineers need the comparison, but we do not collapse them into a single quality score. The most responsible choice is usually the model that completes the task with the fewest failed loops at a cost your team can defend.

Browse the refreshed shelf
See the live catalog, compare credit floors, or start a CritiqueCode session with the hosted routes.
Independent reviewVerified repairReal repositoriesBuilt for developersLoved by agents