Skip to content
Skip to content
INFERENCE API · 25 MODELS

OpenAI-compatible
models at $/1M input and output.

Use Critique as raw token inference for Cursor, Windsurf, Zed, VS Code, and JetBrains, your agent harness, eval loops, or any OpenAI client. Set critique/auto when you want the coding selector. Private list rates, or opt into log retention on eligible models and cut the bill by 75%.

Sign in for a crt_ key. Billed at published $/1M input and output rates.

GET
/api/v1/models
POST
/api/v1/chat/completions
Auto
critique/auto
crt_ keyOpenAI clientPer-token bill
WHY CRITIQUE

Inference built for agent loops

OpenAI-compatible completions priced in dollars per million input and output tokens — private by default, with an optional log-retention discount.

Cheap raw inference

Pin a model when you want one specialist. Use critique/auto when the agent should start on the efficient rung and escalate only if the loop is stuck. Pay the published $/1M input and output rates — or send X-Critique-Billing: byok with a saved OpenAI, Anthropic, or OpenRouter key.

75% off with log retention

Need the bill lower? Opt in to short-term log retention on eligible models and pay 25% of list price. Private-by-default stays at standard rates — no retention, no discount.

Logs improve the stack

Retained prompts and completions help us tune models, routing, and the Critique platform. You ship faster on cheaper tokens; we get signal to make the system better.

  • OpenAI-compatible
  • Bearer crt_ keys
  • Western-hosted
  • Pay per token
  • Cursor
  • Windsurf
  • Zed
  • VS Code
  • JetBrains
CODING SELECTOR

critique/auto

A session-aware coding selector over the Inference catalogue. Efficient models handle the easy work. Frontier specialists take over when the loop is stuck. Token rates follow the routed model.

Docs →

Start cheap

A new Auto session lands on the efficient rung: Laguna S 2.1, Gemini 3.8 Flash, Seed 2.0 Code, DeepSeek V4 Flash 0731, DeepSeek V4.1 Flash, Muse, Qwen 27B, Hy3, Ling. Enough for most edits, FIM, vision, and healthy tool loops.

Stay sticky

The same specialist is reused for the rest of the session so the prompt cache stays warm. Echo X-Critique-Router-Session on later turns.

Escalate once

If tools keep failing, Auto moves the session onto the frontier rung (Kimi K3, Grok 4.6, Fugu Ultra, Kimi K2.7 Code, GLM-5.3, Qwen 2.4T, KAT-Coder, MiniMax). It does not hop every message.

Efficient rung

Default for critique/auto

  • Gemini 3.8 Flashgoogle/gemini-3.8-flash
  • Muse Glimmer 30Bmeta/muse-glimmer-30b
  • Qwen3.8 27Bqwen/qwen3.8-27b
  • Qwen3.8 Flashqwen/qwen3.8-flash
  • DeepSeek V4 Flash 0731deepseek/deepseek-v4-flash-0731
  • DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash
  • Laguna S 2.1poolside/laguna-s-2.1
  • Seed 2.0 Codebytedance-seed/seed-2.0-code
  • Hy3tencent/hy3
  • Ling 3.0 Flashinclusionai/ling-3.0-flash
Frontier rung

critique/auto-best, or after struggle

  • Grok 4.6x-ai/grok-4.6
  • Kimi K3moonshotai/kimi-k3
  • Fugu Ultrasakana/fugu-ultra
  • Muse Spark 1.3meta/muse-spark-1.3
  • Kimi K2.7 Codemoonshotai/kimi-k2.7-code
  • GLM 5.3z-ai/glm-5.3
  • KAT-Coder Prokwaipilot/kat-coder-pro-v2.5
  • Qwen3.8 2.4Tqwen/qwen3.8-2.4t-a95b
  • MiniMax M3minimax/minimax-m3
  • Trinity Large Thinkingarcee-ai/trinity-large-thinking
  • Kimi K2.6moonshotai/kimi-k2.6

Auto only routes among Critique-resold Inference offers. It will not pick Anthropic, OpenAI, or Google catalog ids. How Auto works → · Essay →

RATE CARD

Catalogue rates

Private list rates for fast iteration, or log-retention pricing at 75% off. Every model is priced as $/1M input and $/1M output.

BUILDER PRICING

World-class rates for the models agents burn all day.

OpenAI-compatible endpoints for Cursor, Windsurf, Zed, VS Code, and JetBrains, LangChain, Vercel AI SDK, or your own harness. List rates for fast iteration — or opt into log retention and cut the bill by 75%.

Qwen3.8 FlashNew · flash lane
Private / 1M
$0.150 in
$0.470 out
Log retention / 1M
$0.0375 in
$0.117 out
Laguna S 2.1New · coding MoE
Private / 1M
$0.0900 in
$0.180 out
Log retention / 1M
$0.0225 in
$0.0450 out
Gemini 3.8 FlashNew · Google flash
Private / 1M
$0.750 in
$3.75 out
DeepSeek V4.1 FlashNew · CED multimodal
Private / 1M
$0.150 in
$0.600 out
Log retention / 1M
$0.0375 in
$0.150 out
Qwen3.8 27BNew · dense VLM
Private / 1M
$0.450 in
$3.20 out
Log retention / 1M
$0.113 in
$0.800 out
Ling 3.0 FlashNew · token-efficient MoE
Private / 1M
$0.0750 in
$0.220 out
Log retention / 1M
$0.0187 in
$0.0550 out
MiniMax M3New · multimodal
Private / 1M
$0.300 in
$1.20 out
Log retention / 1M
$0.0750 in
$0.300 out
KAT-Coder Pro V2.5New · agentic coding
Private / 1M
$0.740 in
$2.96 out
Log retention / 1M
$0.185 in
$0.740 out
Kimi K2.7 CodeCoding specialist
Private / 1M
$0.730 in
$3.50 out
Log retention / 1M
$0.182 in
$0.875 out
NVIDIA Nemotron 3 UltraWorld-class MoE
Private / 1M
$1.00 in
$5.00 out
Log retention / 1M
$0.250 in
$1.25 out
Tencent Hy310% below market
Private / 1M
$0.126 in
$0.522 out
Log retention / 1M
$0.0315 in
$0.131 out
DeepSeek V4 FlashDefault · fast lane
Private / 1M
$0.0650 in
$0.180 out
Log retention / 1M
$0.0163 in
$0.0450 out

Opt into log retention on eligible models and pay 75% less per token. Private tier stays at list rates with no retention.

ModelArchitectureContextInput / 1MOutput / 1M
DeepSeek V4 Flash 0731deepseek/deepseek-v4-flash-0731Default
284B MoE (13B active) · 1.31M context1M$0.0650$0.180
Qwen3.8 Flashqwen/qwen3.8-flashNew
Vision-language · 1M context1M$0.150$0.470
Ling 3.0 Flashinclusionai/ling-3.0-flashNew
124B MoE (5.1B active)131K$0.0750$0.220
Tencent Hy3tencent/hy310% below market
295B MoE (21B active)262K$0.126$0.522
NVIDIA Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55bWorld-class MoE
550B MoE (55B active)1M$1.00$5.00
MiniMax M3minimax/minimax-m3New
Multimodal · 1M context1M$0.300$1.20
KAT-Coder Pro V2.5kwaipilot/kat-coder-pro-v2.5New
Agentic coding · 256K context256K$0.740$2.96
Kimi K2.7 Codemoonshotai/kimi-k2.7-codeCoding
1T MoE (32B active)262K$0.730$3.50
Kimi K2.6moonshotai/kimi-k2.6Multimodal agent
1T MoE (32B active)262K$0.750$3.75
GLM-5.3z-ai/glm-5.3Long-horizon coding
753B MoE · 1.31M context1M$1.40$4.40
Trinity Large Thinkingarcee-ai/trinity-large-thinkingOpen reasoning
400B MoE (13B active)262K$0.250$1.00
DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flashNew · CED flash
552B CED MoE (8B/16B active) · 1.05M context1M$0.150$0.600
Seed 2.0 Codebytedance-seed/seed-2.0-codeNew · coding
Agentic coding · 256K · multimodal262K$0.500$3.00
Qwen3.8 2.4T A95Bqwen/qwen3.8-2.4t-a95bNew · open Max
2.4T MoE (95B active) · 1M context1M$2.00$6.00
Muse Glimmer 30Bmeta/muse-glimmer-30bNew
30B dense · multimodal · 128K131K$0.350$1.50
Muse Spark 1.3meta/muse-spark-1.3New · reasoning
Multimodal reasoning · 1M context1M$1.25$4.25
Qwen3.8 27Bqwen/qwen3.8-27bNew
27B dense VLM · 256K context262K$0.450$3.20
Laguna S 2.1poolside/laguna-s-2.1New · coding
118B MoE (8B active) · 256K262K$0.0900$0.180
Grok 4.6x-ai/grok-4.6List only
Frontier · 500K context500K$2.00$6.00
Kimi K3moonshotai/kimi-k3List only
2.8T multimodal · 1M context1M$3.00$15.00
Gemini 3.8 Flashgoogle/gemini-3.8-flashList only
Multimodal flash · 1M context1M$0.750$3.75
Fugu Ultrasakana/fugu-ultraList only
Multi-agent orchestration · 1M1M$5.00$30.00

Rates are $/1M input and $/1M output. Responses include X-Critique-Estimated-Usd

Private by default — log retention is optional for the discount. Full pricing docs →

  • Critique Autocritique/auto
  • DeepSeek V4 Flash 0731deepseek/deepseek-v4-flash-0731
  • Qwen3.8 Flashqwen/qwen3.8-flash
  • Ling 3.0 Flashinclusionai/ling-3.0-flash
  • Tencent Hy3tencent/hy3
  • NVIDIA Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b
  • MiniMax M3minimax/minimax-m3
  • KAT-Coder Pro V2.5kwaipilot/kat-coder-pro-v2.5
  • Kimi K2.7 Codemoonshotai/kimi-k2.7-code
  • Kimi K2.6moonshotai/kimi-k2.6
  • GLM-5.3z-ai/glm-5.3
  • Trinity Large Thinkingarcee-ai/trinity-large-thinking
  • DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash
  • Seed 2.0 Codebytedance-seed/seed-2.0-code
  • Qwen3.8 2.4T A95Bqwen/qwen3.8-2.4t-a95b
  • Muse Glimmer 30Bmeta/muse-glimmer-30b
  • Muse Spark 1.3meta/muse-spark-1.3
  • Qwen3.8 27Bqwen/qwen3.8-27b
  • Laguna S 2.1poolside/laguna-s-2.1
  • Grok 4.6x-ai/grok-4.6
  • Kimi K3moonshotai/kimi-k3
  • Gemini 3.8 Flashgoogle/gemini-3.8-flash
  • Fugu Ultrasakana/fugu-ultra
  • Critique Autocritique/auto
  • DeepSeek V4 Flash 0731deepseek/deepseek-v4-flash-0731
  • Qwen3.8 Flashqwen/qwen3.8-flash
  • Ling 3.0 Flashinclusionai/ling-3.0-flash
  • Tencent Hy3tencent/hy3
  • NVIDIA Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b
  • MiniMax M3minimax/minimax-m3
  • KAT-Coder Pro V2.5kwaipilot/kat-coder-pro-v2.5
  • Kimi K2.7 Codemoonshotai/kimi-k2.7-code
  • Kimi K2.6moonshotai/kimi-k2.6
  • GLM-5.3z-ai/glm-5.3
  • Trinity Large Thinkingarcee-ai/trinity-large-thinking
  • DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash
  • Seed 2.0 Codebytedance-seed/seed-2.0-code
  • Qwen3.8 2.4T A95Bqwen/qwen3.8-2.4t-a95b
  • Muse Glimmer 30Bmeta/muse-glimmer-30b
  • Muse Spark 1.3meta/muse-spark-1.3
  • Qwen3.8 27Bqwen/qwen3.8-27b
  • Laguna S 2.1poolside/laguna-s-2.1
  • Grok 4.6x-ai/grok-4.6
  • Kimi K3moonshotai/kimi-k3
  • Gemini 3.8 Flashgoogle/gemini-3.8-flash
  • Fugu Ultrasakana/fugu-ultra
$ / 1M

Billing

Prompt tokens bill at the model input rate. Completion tokens bill at the output rate. Both are published as dollars per million tokens.

  1. 1
    Count tokens

    prompt_tokens bill at the model input rate. completion_tokens (including reasoning tokens when reported) bill at the output rate.

  2. 2
    Charge USD

    inputUsd = (prompt_tokens ÷ 1M) × inputRate

    outputUsd = (completion_tokens ÷ 1M) × outputRate

    totalUsd = inputUsd + outputUsd

    Inference does not convert this into review credits.

BYOK

Bring your own keys

Route OpenAI, Anthropic, OpenRouter, Crof, or LLM Gateway with keys saved in Settings — OpenRouter-style multiplexing on the same crt_ endpoint.

Provider keys →
  1. 1
    Save a provider key

    Add OpenAI, Anthropic, OpenRouter, Crof, or LLM Gateway under Settings. Keys stay encrypted and are used only server-side.

  2. 2
    Send BYOK billing

    Header X-Critique-Billing: byok. OpenAI ids (openai/gpt-4o, gpt-4o) route to api.openai.com. Anthropic ids use your Anthropic key. Everything else uses LLM Gateway, then Crof, then OpenRouter — same multiplex shape as OpenRouter.

  3. 3
    Tokens bill the provider

    BYOK completions do not debit prepaid Critique Inference credits. Managed traffic still uses the rate card and prepaid balance.

QUICKSTART

Point any OpenAI client at Critique.

Use https://critique.sh/api/v1 with a crt_ key. Works in Cursor, Windsurf, Zed, VS Code, and JetBrains, LangChain, Vercel AI SDK, CI evals, or a sidecar next to your agent loop. CritiqueCode (@critiquedotsh/harness) can default to critique/auto on this API.

CritiqueCode

npm install --global @critiquedotsh/harness
critique-code
curl https://critique.sh/api/v1/chat/completions \
  -H "Authorization: Bearer crt_..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek/deepseek-v4-flash-0731",
    "messages": [
      { "role": "user", "content": "Summarize this webhook retry design in three bullets." }
    ]
  }'

OpenAI SDK

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.CRITIQUE_API_KEY,
  baseURL: "https://critique.sh/api/v1",
});

const response = await client.chat.completions.create({
  model: "deepseek/deepseek-v4-flash-0731",
  messages: [{ role: "user", content: "Draft a TypeScript interface for idempotent job enqueue." }],
});

console.log(response.choices[0]?.message?.content);

OpenAI-compatible client

// OpenAI-compatible — same shape in any client with a custom base URL
import OpenAI from "openai";

export const inference = new OpenAI({
  apiKey: process.env.CRITIQUE_API_KEY,
  baseURL: "https://critique.sh/api/v1",
});

const res = await inference.chat.completions.create({
  model: "critique/auto",
  messages: [{ role: "user", content: "Refactor this handler for idempotency." }],
});
ACCESS

API keys

Scopes: read:inference / write:inference

Create your Inference API key

Keys use the crt_ prefix with read:inference and write:inference scopes. Sign in to generate a key — the full secret is shown once.

Sign in to generate a key
FAQ

Questions

What is critique/auto?

A session coding selector over Critique-resold Inference models. New sessions start on the efficient rung (Laguna S 2.1, Gemini 3.8 Flash, Seed, DeepSeek V4 Flash 0731, DeepSeek V4.1 Flash, Muse, Qwen 27B, Hy3, Ling). Later turns reuse the same specialist. Repeated tool failures escalate once onto the frontier rung (Kimi K3, Grok 4.6, Fugu Ultra, Kimi K2.7 Code, GLM-5.3, Qwen 2.4T, and others). It does not pick from the full catalog. Token rates follow the routed model.

What is CritiqueCode vs the Critique CLI?

CritiqueCode is the author agent, package @critiquedotsh/harness, binary critique-code. It implements, then the controller forces review and verified repair. Run critique-code login and approve on the website to connect Critique Inference with no pasted key. /voice transcribes speech with qwen/qwen3-asr-0.6b so you can talk to the author and reviewer. The Critique CLI is the sidecar reviewer, package @critiquedotsh/cli, binary critique. Same npm org, different products.

Why use Critique as my inference layer?

Same crt_ keys as finish checks and the CLI, OpenAI-compatible endpoints, Western-region routing, and list rates in $/1M input and output for agents that burn tokens all day. Point Cursor, Windsurf, Zed, VS Code, and JetBrains, your harness, or any OpenAI client at https://critique.sh/api/v1 — no separate provider account on managed billing.

Do you store or train on my prompts?

Private tier (default): no log retention for model training — requests are metered for billing only. Log retention tier (optional, 75% off on eligible models): you opt in explicitly; retained prompts and completions help improve models, routing, and the Critique platform. Toggle in Settings → Connections or per request with X-Critique-DeepSeek-Training-Opt-In.

What is log retention pricing (75% off)?

On DeepSeek V4 Flash 0731, DeepSeek V4.1 Flash, Qwen3.8 Flash, Qwen3.8 2.4T A95B, Qwen3.8 27B, Seed 2.0 Code, Muse Glimmer 30B, Muse Spark 1.3, Laguna S 2.1, Ling 3.0 Flash, Hy3, Nemotron 3 Ultra, MiniMax M3, KAT-Coder Pro, Kimi K2.7 Code, Kimi K2.6, GLM-5.3, and Trinity Large Thinking you can opt into short-term log retention and pay 25% of list price (75% off). Grok 4.6, Kimi K3, Gemini 3.8 Flash, and Fugu Ultra stay at list rates. Accept the conditions, enable account-wide in Settings, or send X-Critique-DeepSeek-Training-Opt-In: true per request.

Read the full API reference →

POLICY

Privacy & acceptable use

Payloads are processed to return completions and meter token spend — not sold. Critique does not train foundation models on your prompts unless you explicitly enable log-retention pricing on eligible models.

Traffic stays on Western-hosted capacity. Do not use the API for illegal activity, malware, or attempts to extract other customers' data. Abuse may suspend keys.

Full reference: Inference API docs →

Same org, two binaries

@critiquedotsh/harness is CritiqueCode, the author agent. @critiquedotsh/cli is the sidecar reviewer. Do not mix them.