Journal · independent review
Field notes on
finish checks.
Product launches, verifier research, and essays on why coding agents need an independent second opinion before they claim a task is done.
Browse comparisonsDeepSeek V4.1 Flash in CritiqueCode: A Completed Review Run
We ran DeepSeek V4.1 Flash locally in CritiqueCode: architecture, inference rates, coding and vision tests, and the schema fix that made the structured review complete.
Coding Agent Completion Benchmark: Methodology Before Scores
An open methodology for comparing author-only, author-plus-review, and verified-repair coding agent workflows without inventing product superiority.
OpenCode Code Review Workflow for Local Changes
Review OpenCode changes before commit with explicit intent, complete working-tree scope, executable checks, and an independent finish result.
Codex Code Review for Local and Uncommitted Changes
A practical Codex code review workflow for staged, unstaged, deleted, and untracked changes before commit or pull request.
How to Verify AI-Generated Code Before You Ship
A practical checklist for verifying AI-generated code with author checks, independent review, tests, scanners, CI, and human judgment.
Claude Code Review: Self-Review vs an Independent Second Pass
Compare Claude Code local review, managed pull-request review, and an independent finish pass for uncommitted changes.
AI Code Review CLI: Review Uncommitted Changes Before You Push
Use an AI code review CLI to inspect staged, unstaged, deleted, and untracked work before you push. Compare self-review, independent review, CI, and local inference.
Bringing the Frontier to Critique
Critique adds OpenAI GPT-6 Astra and Anthropic Claude Fable 5.1 to the managed model catalog, with transparent frontier routing and published per-million-token pricing.
CritiqueCode 2.0: The Harness Decides What Is True
CritiqueCode 2.0 turns coding-agent claims into state-bound engineering records: requirements, controlled checks, evidence, review, repair, and honest incomplete outcomes.
One Model, Seven Harnesses, 140 Trials
We ran 140 Harbor trials in fresh E2B sandboxes with z-ai/glm-5.3-flash across seven coding harnesses, retaining task-level outcomes, costs, trajectories, tool activity, timing, patches, and verifier evidence.
Same Model. Four Harnesses. What Actually Survives the Verifier?
We held Mercury 2.5 constant and changed the coding harness. This open study reports 40 Harbor/E2B trials, independent F2P/P2P verification, task-level failures, costs, telemetry, and the complete reproduction trail.
Critique Auto Gets a New Model Roster
Critique Auto now routes across refreshed coding models, with current token pricing, shared SWE-bench context, four hosted routes in CritiqueCode, and credit access to the full catalog.
CritiqueCode Hosted Runtime: An Early, Truthful Release
The first hosted CritiqueCode runtime is an early implementation with isolated E2B sandboxes, GitHub branches, and explicit availability checks.
CritiqueCode 0.1.0 Shipped a Contract. This Is Everything After.
CritiqueCode 0.1.3 vs 0.1.0: website login, critique/auto, voice, imported skills, critique_task workers, and a local web UI. Forced review did not move.
CritiqueCode Leaves the Terminal Without Leaving Your Machine
critique-code web serves CritiqueCode in a local browser on 127.0.0.1. Same author kernel, same /review gate, no OpenCode fork and not Critique Cloud.
Talk to Code with CritiqueCode Voice Mode
Talk to CritiqueCode instead of typing. /voice transcribes with Qwen3 ASR 0.6B; spoken review, repair, and ship still run as commands. Start --voice.
Cursor vs Codex
Cursor vs Codex in 2026: pick the IDE Agent if you live in the editor, or OpenAI’s CLI and optional cloud if you work in the terminal. Fair comparison.
Best CLI Tools for Coding in 2026
Compare git, gh, Claude Code CLI, Codex CLI, Aider, OpenCode, and CritiqueCode. Best CLI tools for coding in 2026 depend on the job: VCS vs AI author.
CritiqueCode Is the Coding Agent That Cannot Grade Its Own Homework
CritiqueCode is Critique’s author agent: a Claude Code-style coding CLI that implements in session, then the controller forces independent review and verified repair. Install @critiquedotsh/harness and run critique-code.
Best Terminal Coding Agents in 2026
Compare the best terminal coding agents in 2026: Claude Code, Codex CLI, OpenCode, Aider, Gemini CLI, and CritiqueCode. Pick by surface, models, and review contract, not invented benchmarks.
Best Open Source Coding Agents in 2026
Compare the best open source coding agents in 2026: OpenCode, Aider, Cline, Continue, Goose, OpenHands, Gemini CLI. Licenses, surfaces, and where CritiqueCode fits.
Best Coding Agent in 2026: Claude, Codex, Cursor
Compare Claude Code, Codex, Cursor, OpenCode, and CritiqueCode in one 2026 guide. Pick the best coding agent by jobs: IDE, CLI, or author-plus-review.
OpenCode vs Claude Code
OpenCode vs Claude Code: MIT model-agnostic terminal agent at opencode.ai versus Anthropic’s Claude Code. Fair, sourced comparison of license, models, and review.
Aider vs Claude Code
Aider vs Claude Code: open-source git-centric pair programming at aider.chat versus Anthropic’s Claude Code. Fair, sourced comparison of license, models, and review.
Claude Code vs Codex
Claude Code vs Codex CLI in 2026: Anthropic terminal agent vs OpenAI Codex. Install, review, automation, when to pick each, and a sidecar if you keep either.
Claude Code vs Cursor
Claude Code vs Cursor in 2026: Anthropic’s terminal agent versus Cursor’s AI IDE. Fair comparison of surfaces, tools, pricing, and when to keep each.
CritiqueCode vs Claude Code: When to Keep, Switch, or Combine
CritiqueCode vs Claude Code: keep Anthropic's terminal agent, switch to critique-code, or pair Claude Code with Critique CLI sidecar. Review vs evidence.
CritiqueCode vs OpenCode
CritiqueCode vs OpenCode: open-source model-agnostic terminal agent versus critique-code with forced review. Fair, sourced, MIT OpenCode vs @critiquedotsh/harness.
CritiqueCode vs Codex
CritiqueCode vs Codex: keep OpenAI Codex CLI, switch to critique-code, or pair Codex with the Critique CLI sidecar. Review vs evidence, install, and when to switch.
CritiqueCode vs Cursor
CritiqueCode vs Cursor: keep Cursor Agent and Composer in the IDE, switch to critique-code, or pair Cursor with Critique CLI sidecar. Review vs evidence.
Best Claude Code Alternatives in 2026
The best Claude Code alternatives in 2026 depend on CLI, IDE, or forced independent review. Compare Codex, Cursor, OpenCode, Aider, Gemini CLI, CritiqueCode.
GitHub Copilot vs Cursor
GitHub Copilot vs Cursor in 2026: VS Code and GitHub seats versus Cursor’s AI IDE, Tab, Agent, and Composer. Official list prices, a fair two-column table, and when to keep each.
Cursor vs Codex vs CritiqueCode
Compare Cursor vs Codex vs CritiqueCode in 2026: IDE agent, OpenAI CLI/cloud loop, or local authoring whose controller forces review and verified repair.
Stop putting Kimi on a rename
Critique Auto is a session coding selector on the Inference API. It routes among Critique-resold models only: efficient specialists first, frontier coding models after struggle, billed at the routed model.
From git diff to First Customer
Your AI coding agent finished the code. Here’s the stack for reviewing AI-generated code, deploying to production, presenting your product and finding your first customers.
Introducing CritiqueCode: the hosted Code workspace
CritiqueCode Beta is the Critique-owned hosted Code workspace. It gives code changes explicit task contracts, guarded tools, durable artifacts, and evidence-backed outcomes.
Critique stopped wrapping the agent and started owning the review
Critique CLI now bundles a pinned OpenCode runtime behind a typed SDK, keeps the main coding agent out of reviewer judgment, captures local evidence through a Critique plugin, and provides a collaborative task spine and second brain.
Critique CLI Gets a Conversational Second Brain
Critique CLI adds read-only ask and chat sessions for coding-agent second opinions, plus stricter budgets, events, and verified repair honesty.
Critique 0.2: Review Less, Test the Right Risk
Critique CLI 0.2 adds focused AI code review modes for quick checks, task-aware verification, security testing, and resilience testing, with clearer terminal guidance for people and coding agents.
Critique vs Built-In Agent Checks
Compare Critique with built-in Codex, Claude Code, and OpenCode checks: fast self-verification versus an independent, evidence-backed finish pass.
critique.sh vs CI: Different Layers of Proof
Critique is a pre-CI or alongside-CI verification layer for coding agents, not a CI replacement. See where each belongs in an engineering workflow.
Best Coding Agent Verification Tools in 2026
An opinionated 2026 guide to coding agent verification: Critique, built-in agent checks, CI, security scanners, and human review.
Your Coding Agent Is Lying to You (It Just Doesn’t Know It)
Why coding agents can confidently report completion without proving it, and how an independent finish loop replaces self-report with evidence.
Agents Grade Their Own Homework. That’s Broken.
Why self-grading coding agents create avoidable blind spots, and how a narrow independent verifier improves agent-written code without slowing the primary loop.
We Killed the Dashboard. It Wasn’t the Workflow.
Why Critique shifted from a broad human-facing code review surface to an agent-facing completion layer that verifies a coding agent’s finish claim.
Critique Is Now the Teammate Your Agent Calls
Critique is now an agent-facing CLI for independent verification and repair. Meet the new finish loop for Codex, Claude Code, and OpenCode.
LLM Gateway BYOK on Critique: Same $8 Harness, Broader Model Surface
Critique adds LLM Gateway BYOK for review, Remedy, Builder, and chat model calls: same $8 harness, encrypted key storage, broader OpenAI-compatible model gateway, and LLM Gateway takes precedence when multiple BYOK keys are saved.
The senior developer shortage is coming. We're engineering it ourselves.
Entry-level tech hiring collapsed while AI productivity soared. The five-to-seven year mentorship lag means the senior developer drought hits around 2029–2031 — and every company that ran the spreadsheet will be competing for a pool they collectively drained.
Critique Coding Agent API: How Teams Are Actually Using Cloud Agents Over HTTP
Critique’s June 2026 Coding Agent API update adds lifecycle fields, tags, metadata, intent classification, better SSE status, and cleaner follow-up behavior — plus a clearer view of how teams are wiring cloud coding agents into CI, support, and internal platforms.
Critique Is Going Open Source
Critique is launching a public community edition focused on self-hosted GitHub pull request review. This post explains what is opening first, what is staying out of scope, and why we are using a separate public repository.
Critique Intake: Feedback That Arrives Already Debugged
Introducing Critique Intake, an agentic bug intake workflow for turning vague user reports into triaged engineering packets, agent prompts, Review runs, and Change Passports.
Critique v6.1.1: Agent Stack Integrations — When the Merge Gate Meets RWX, Swytchcode, and Iron Book
Product update: Critique v6.1.1 ships upstreamSignals on POST /api/v1/reviews, partner cookbooks for RWX / Swytchcode / Identity Machines, passport provenance, and docs for the full agent stack loop — build, judge, merge, execute.
The Merge Gate API: Why Agent Teams Need a Judge That Is Not the Writer
A deep essay on Critique’s Merge Gate API — why writer and judge must be separate roles, what v6.1 ships (REST, MCP, webhooks, findings=all), and how merge governance beats self-review loops as coding agents multiply PR volume.
GLM-5.2 Lands in Critique: 1M Context, Design Arena #1, Same 3-Credit GLM-5.1 Shelf
GLM-5.2 replaces GLM-5.1 and GLM-5V-Turbo in Critique’s catalog at the same 3-credit review floor. Z.ai benchmark table vs GPT-5.5 and Claude Opus 4.8, Design Arena Elo 1360, long-horizon FrontierSWE / PostTrainBench / SWE-Marathon charts, and how the inference API still exposes `z-ai/glm-5.1` routing to GLM-5.2 upstream.
You Merged That. Why.
You knew it wasn't ready. You merged anyway. Why AI velocity outpaced review, what actually happens to your PRs, and how Critique closes the gap with sandbox verification and merge policy.
Kimi K2.7 Code Lands in Critique: Open-Source Coding at 4.5 Credits, With Moonshot’s Full Benchmark Table
Kimi K2.7 Code joins Critique’s Moonshot lane at 4.5 credits per PR review run (K2.6 stays at 4). Vendor-reported coding and agentic benchmarks vs GPT-5.5 and Opus 4.8, MCP Atlas cross-reads against Gemini 3.5 Flash, MiniMax M3, and Qwen3.7 Max, and reasoning-efficiency gains over K2.6.
Why Your AI-Generated Next.js App Works Locally but Dies on Vercel
Next.js works locally but fails on Vercel? Case-sensitive imports, env var gaps, and next build vs next dev — plus a 10-minute pre-push ritual to catch AI import mistakes.
Breaking the Dead Loop: What to Do When Cursor or Windsurf Refuses to Learn
Cursor or Windsurf stuck on the same broken fix? Loop taxonomy, two-strike rule, circuit breaker prompts, TDD ground truth, and Critique sandbox verification.
How AI Code Generators Create Circular Imports in Next.js (and How to Spot Them)
Five AI patterns that create circular imports in Next.js, why local dev lies, how to detect cycles with madge and dependency-cruiser, and the fix order that actually works.
Code and Pray Is Not a Workflow: Why Vibe Coders Need Sandbox Verification
Static review vs runtime proof, the Plan-Work-Verify loop, E2B and Docker sandbox patterns for GitHub PRs, and how Critique productizes sandbox verification for vibe coders.
The Vibe Coder's Security Checklist: 5 Ways Your AI IDE Is Leaking Secret Keys
Five common ways Cursor, Copilot, and agent mode leak API keys into git diffs, model context, client bundles, and CI logs — plus a PR merge-gate checklist to prevent environment variable leaks.
Silent TypeScript Failures: The Danger of 'Fixed' by Any
Why AI agents silence TypeScript errors with `any`, `as`, and `@ts-ignore`, how Next.js dev vs build diverge, and a PR checklist for real type fixes — not check-engine-light patches.
Why We Built the Critique Inference API: Western Servers, NVIDIA-Powered Capacity, and Credits You Already Own
Critique Inference API launches at /inference-api with DeepSeek V4 Flash, Tencent Hy3 Preview, and NVIDIA Nemotron 3 Ultra on Western sweetener servers — stress-tested, privacy-first by default, and billed from Critique credits. Background from Leemer Labs on why we did not wait for enterprise timing.
Critique v5.1: The Review System Starts Reviewing Itself
Critique v5.1 ships a review evaluation harness, repository learnings, a quality budget for findings, an independent critic pass, E2B/OpenCode runtime drift checks, and a clearer Coding Agent API surface.
Critique v5.1: We Hit Go, Walked Away for Five Hours, and Shipped While Reviewing Ourselves
Critique v5.1.0 (5 June 2026) ships operator-first review runs, one Coding Agent product across Builder and API, marketplace attribution, repo-home defaults, and Platform connections beta — documented in a ship log written largely by cloud agents that reviewed their own changes for nearly five hours after a human pressed go.
Critique vs Devin for Cloud Coding Agents: API, Pricing, and Model Freedom
Compare Critique Coding Agent API and Devin on API access, pricing (quota vs credits vs OpenRouter BYOK), model choice including OpenRouter free routes, and when each cloud coding agent fits.
Critique v5 Beta: Marketplace, Merge Policy in Plain English, Passport Exports, and the Biggest Ship Since v4
Critique v5 beta (v5.0.0) ships the Agent Skill Marketplace, NL merge policy compiler, Change Passport HMAC exports, MiniMax M3 welcome pricing, Cursor Agent SDK BYOA, Coding Agent API persistent sessions, repo-first PR dashboard, finding feedback, opt-in model-feedback sharing, Workspace agent queue, and Insights velocity/risk/cost/compliance — the largest platform update since v4 Change Control.
Coding Agent API: Persistent OpenCode Sessions for Multi-Turn Automation
How Critique Coding Agent API persistent sessions work: warm E2B sandboxes, OpenCode continuity, idle status, SSE streaming, and follow-up messages for CI bots and internal agent platforms.
Cursor as a Top-Tier Agent Harness: Composer 2.5, Cloud BYOA, and How It Compares to the Models on Critique
Why Cursor is a first-class agent harness, what Composer 2.5 is (specs, SWE-Bench Multilingual, Terminal-Bench 2.0, pricing), how Critique queues cloud agents via the Cursor Agent SDK on composer-2.5, and how that compares to models on the Critique review catalog.
MiniMax M3 and Qwen3.7 Plus on Critique: Coding Benchmarks and a Two-Week M3 Welcome Price
MiniMax M3 coding performance on SWE-Bench Pro and Terminal-Bench 2.1, M3 vs Opus 4.8 and GPT-5.5, Qwen3.7 Plus vs Qwen3.6 Plus on terminal and UI benchmarks, and a 50% welcome credit window on M3 PR review runs through June 17, 2026. Critique Chat stays Ling and DeepSeek V4 Flash.
Best Code Review Skill for Claude Code, Hermes, Codex, and Opencode
Best code review skill for Claude Code, Codex, Hermes, and Opencode, with a real Moonshot Kimi K2.6 PR review comparison and setup guide.
Critique v4.1: Your Merge Verdict, Where Your Team Already Works
Critique v4.1 connects to Linear, Slack, Zapier, and the tools you already use. One review verdict across GitHub, chat, and your stack—without asking engineers to learn a new destination.
Getting Hundreds of PRs and No Time to Review Them? Open Source Needs PR Control.
Open source maintainers and foundations facing hundreds of pull requests need PR control: gate, evidence-backed review, and Change Passports. Pro and Team plans for volume; verified OSS lane at Solo-equivalent pricing. OSS credit grants available.
Critique v4: The AI Change Control Platform (and What Happened to “Just Review”)
A deep guide to Critique v4 vs v1–v3: Change Passports, the Control Board, gate → review → merge, AI-powered sandbox evidence runs, and why Critique is building the best AI management layer for code — not another comment bot on your diff.
v3.6.0: Bring Your CrofAI Key — Same $8 Harness, Direct OSS Billing
Critique v3.6.0 adds CrofAI BYOK alongside OpenRouter: save a Crof API key, bill nahcrof directly at https://crof.ai/v1, and keep the same $8/month Critique harness for orchestration.
Critique Welcomes Claude Code: Review Here, Execute on Claude
How Critique hands a scoped review blueprint to Claude Managed Agents—encrypted Anthropic keys, GitHub-mounted PR checkout, the managed-agents-2026-04-01 beta, and a worker that streams until the session goes idle. Review stays on Critique; execution bills on your Anthropic account.
Critique Welcomes Codex: Review Here, Execute on Your OpenAI Account
How Critique hands review findings to OpenAI Codex: a deterministic JSON handoff envelope with allowed write paths, validation commands, and stop conditions; an encrypted OpenAI key; and a Responses API worker. Review stays on Critique; execution bills on your OpenAI account.
Critique Welcomes Cursor: Review Here, Fix There, Ship Together
Critique now queues scoped fix handoffs to Cursor Cloud Agents with your API key. Why review and execution belong in different products—and how to get the best developer experience by using both.
Introducing Workspace: one place for review, chat, build, and repair
Critique Workspace merges PR review, Chat, Builder, and Remedy into one signed-in surface at /workspace — one history rail, shared repo context, and one credit ledger.
DeepSeek & MiMo Are 0.5 Credits Forever — Plus Opus 4.8, Qwen3.7-Max, and 2026's Cheapest Frontier Review Stack
Critique’s May 2026 model spring: permanent 0.5cr DeepSeek V4 Flash and MiMo v2.5, 1cr DeepSeek V4 Pro and MiMo Pro, Ling at 0.5cr, MiniMax M2.7 at 1.5cr, Qwen3.7-Max at 6cr, Claude Opus 4.8 and Fast, Gemini 3.5 Flash, Grok Build 0.1, and Ring-2.6-1T.
Is critique.sh Safe for Private Repositories?
Is critique.sh safe for private repos? A code-level review of GitHub access, E2B sandboxes, data retention, chat storage, and provider logging.
AI Code Review Pricing Is Getting Weird: What Teams Actually Pay in 2026
The 2026 AI code review pricing trap: seats, usage billing, Actions minutes, and shared credits all make pull request review cost behave differently.
Before You Merge AI Code, Run This Security Checklist
A security checklist for AI-generated code that catches the scary stuff before merge: auth gaps, data leaks, secrets, dependencies, prompt injection, and unsafe fixes.
CodeRabbit vs Cursor Bugbot Is the Wrong Fight
CodeRabbit vs Cursor Bugbot sounds simple until pricing, IDE lock-in, benchmarks, and GitHub workflow fit change the answer.
Your Startup Probably Needs AI Code Review Before It Hires QA
A startup-focused AI code review guide for teams shipping too fast to manually inspect every risky pull request.
Stop Letting the Agent That Wrote the Code Review It Too
Why coding agents and PR review agents should be separate jobs, and how to wire them into GitHub without trusting the same agent twice.
The Honest Reason AI Code Review Cannot Be Flat-Rate Forever
Why Critique moved to clearer AI code review pricing: bigger credit pools, transparent review costs, low-balance gates, and an $8/month OpenRouter BYOK harness.
A fairer credit system for Critique
Critique now prices one standard review unit at roughly 1M input and 150k output tokens instead of 100k and 15k. That makes large, context-heavy reviews much cheaper without flattening the model ladder. We also refreshed the pricing page and are giving every user +1,000 bonus credits.
Critique Checkpoint: stop slop at the gate
Critique Checkpoint is the new pre-review trust layer for GitHub pull requests: contributor, account, activity, language, and slop-pattern rules before deep Critique review starts.
The Review That Refuses to Ghost You
Inside Critique's two-layer PR review watchdog: DeepSeek V4 Flash supervision, recoverable stalls, GDPval-AA vs Claude Sonnet 4.6 and Opus 4.6, vendor list pricing anchors, limited-time fifty-percent credit lanes on Flash and Pro, and why speed-per-cognition is what merge-gate infra needs.
One month. A billion tokens. What it actually took.
A founder note on the first serious beta month: roughly a billion tokens processed, hundreds of thousands of lines written, five distinct surfaces now talking to each other, and what comes next when they finish becoming one.
Qwen 3.6 and Grok 4.3 in Critique: cheaper routing, stronger coding signal, and one absurd xAI price cut
Catalog refresh for Critique and Remedy: Grok 4.3 replaces Grok 4.2 with a steep credit drop, Qwen3.6-35B-A3B replaces Qwen3.5-27B, InclusionAI Ling-2.6-Flash joins at 1 credit, Qwen3.6-Max-Preview joins at 8 credits, GLM-5.1 gets a 1-credit discount, KAT Coder Pro V2 returns to its normal 2-credit price, and MiMo v2.5 increases by 1 credit.
From ManusAI, v0, and Lovable to code you can merge
If you ship with ManusAI, v0.app, or Lovable, you already have velocity. Here is how Critique fits after export: GitHub-native PR review with repo context, specialist lanes, and optional verified fixes so prototypes mature into production discipline.
Referral rewards are live: credits and cash for early advocates
The Critique referral program is up. Get your link on the referrals page, earn bonus review credits for every successful signup, and race for early milestones: 5,000 credits + $50 at 10 referrers, 10,000 credits + $150 at 100.
Critique PR Review v5: Diff-First Runs & Multi-Turn OpenCode
PR Review v5 leads with the PR diff and model findings, adds multi-turn OpenCode in one session with bounded follow-ups, quiets the default live feed, and sketches the next layer: a cheap thinker (e.g. DeepSeek V4 Flash) for meta-review.
GPT-5.5 and GPT-5.5 Pro in Critique: Benchmarks, Pricing, and When to Spend the Credits
GPT-5.5 and GPT-5.5 Pro are now visible in Critique: benchmark tables for Terminal-Bench 2.0, SWE-Bench Pro, BrowseComp, GDPval, long context, pricing, OpenRouter IDs, and practical routing guidance for PR review and Remedy.
DeepSeek V4 Flash and V4 Pro in Critique: 1M context, EU inference, and open-weights leadership on GDPval-AA
DeepSeek V4 Flash (1 cr) and V4 Pro (3 cr) replace V3.2 Speciale in Critique’s runtime catalog: MoE scale, 1M-token context, EU-hosted routing, GDPval-AA and token-budget benchmarks vs V3.2 and frontier peers, DeepSeek API pricing vs Critique credits, and integration notes for lead, specialist, and Remedy roles.
Introducing Critique Builder: where English compiles
Critique Builder is a browser-first cloud agent surface for starting repo-backed coding jobs, watching live execution, and inspecting results without turning the product into a browser IDE. Managed sandboxes, OpenCode, subagents, and a stronger bias toward shipping real software.
Critique PR Review v4.1: Execution Depth, Live Observability, and Operator Secrets
PR Review v4.1 deepens the OpenCode sandbox: exploratory testing, granular GitHub inline comments, live OpenCode activity streaming to the dashboard, and AES-256-GCM encrypted per-repository secrets for real integration runs.
Critique PR Review v4: The Beta Is Getting Real
Critique PR Review v4 is live in beta: nearly 4,000 PRs reviewed, sandbox-native OpenCode execution getting tighter, and Critique now showing up as the #2 public app this month on OpenRouter’s Embed V1 4B activity view just behind AnythingLLM.
Arcee Trinity-Large-Thinking Lands in Critique at 1 Credit
Arcee AI’s Trinity-Large-Thinking is now available in Critique as `arcee-ai/trinity-large-thinking` at a 1-credit floor. We break down the Apache 2.0 release, MoE architecture, reasoning-trace constraints, benchmark profile, and why this model matters for long-horizon agentic review and Remedy.
Kimi K2.6 Lands in Critique: PR Review, Remedy, and a Week of Half-Price Credits
Kimi K2.6 replaces Kimi K2.5 across Critique’s PR review stack and Remedy chat models: architecture, benchmarks vs K2.5 and peers, integration notes, and April 21–27 2026 introductory pricing at 50% off former Kimi K2.5 credits.
Welcome Gemma-4-31B and MiMo-V2-Flash: The Permanent 0.5cr Review Tier
Gemma-4-31B and MiMo-V2-Flash join critique.sh at a permanent 0.5cr floor. We compare SWE-bench Verified scores and critique.sh credit costs across Gemini 3 Flash, MiniMax-M2.5, Qwen3.5-27B, GPT-5.4 mini, Claude Haiku 4.5, GPT-5.1 Codex Max, and Claude Sonnet 4.6 — with a price-to-performance ratio chart.
Critique PR Review v3.1: The Sandbox Now Finishes the Review
Critique PR Review v3.1 moves final review synthesis into the sandbox. The sandbox now authors the structured review artifact with summary, findings, command timeline, runtime checks, and sub-agent output, while the app stays the control plane.
Hello to v3 of PR Review
Critique PR Review v3 is live one day after the V2 release. Sandbox-backed AI code review, dedicated review-analysis workers, canonical review-run pages, GitHub App publishing, Remedy handoff, fix prompts, and an owner-only review assistant are now part of the product.
Critique is now conversational
Critique now supports interactive PR chat via @critique mentions in GitHub pull request comments. Ask questions, run slash commands for review, explain, fix, security, tests, and more. Powered by Qwen 3.7 Plus — a frontier-level model with a 1M-token context window that beats GPT-5.1-Codex, GPT-5.2, and GPT-5.3-Codex on long-horizon benchmarks.
Critique just got a whole lot better
We shipped ten deep improvements to the Critique review engine: a semantic index, subsystem clustering for large PRs, intrinsic-risk-based drill-down targeting, cross-file relationship analysis, a CODE_QUALITY specialist, two-phase reasoning, domain heuristic packs, contract-diff analyzers, better evidence packaging, and expanded model support across GPT-5.4, Grok-4.2, GLM-5V-Turbo, MiMo v2.5, and more.
The AI Code Review Trust Gap
AI code review is now mainstream, but developer trust still lags dangerously behind adoption. Here is why engineering leaders should evaluate AI review tooling as governance infrastructure — and what the real buying criteria look like in 2026.
Remedy, Now in Critique Chat
Remedy — Critique's autonomous code-fixing engine — is now available directly inside Critique Chat. Start from a review, choose a model, and watch fixes land on your PR branch in real time. Your machine never runs a line of this code.
Hybrid Code Retrieval: Lexical Precision + Semantic Recall
A technical deep dive into hybrid retrieval for codebases: exact string search, semantic search, warmed snapshots, reciprocal rank fusion, and the operational details that make code search trustworthy.
5 Questions Before Installing an AI Code Review GitHub App
Staff engineer or platform lead evaluating AI PR review automation? These five questions — permissions, token model, data egress, check behavior, and uninstall path — separate a trustworthy GitHub App from a risk review waiting to happen.
Our Faith in Open Source
A detailed critique.sh essay on open-source and open-weight AI models, MiniMax M2.7 vs Claude Opus 4.6, why Chinese labs are leading the cost-capability curve, and the March 2026 credit promos for MiniMax, GLM-5, and Kimi K2.5.
MCP vs CLI for AI Agents: Efficiency, Governance, and When Each Wins
MCP and CLI are often framed as rivals for AI agent tool invocation. Primary benchmarks show CLI can be 9–32× cheaper in tokens for well-known tools, while MCP wins on discoverability, OAuth, and tenant isolation. This article synthesizes the evidence and ends with a decision framework.
Critique Chat: Ask Your GitHub Repo Questions with Frontier Models
Critique Chat is the dashboard assistant grounded in your real GitHub codebase. Connect a repo, ask questions, get answers backed by live code search. Free for all signed-in users with frontier models like KAT Coder Pro V2, MiniMax M2.7, GLM-5V-Turbo, and MiMo v2.5.
Cloud Coding Agents: What Changed, What Stayed the Same, and Where Remedy Fits
An honest map of cloud coding agents across GitHub, Cursor, Jules, Devin, OpenCode, E2B, and Vercel Sandbox, plus where Critique Remedy fits as a review-attached execution layer with BYOA support.
Relace Search: The 256K-Context Subagent Behind Serious Codebase Retrieval
Deep dive on Relace Search (relace/relace-search): 256K context, parallel tool calls, report_back handoff to oracle agents, OpenRouter pricing and telemetry, benchmarks vs naive retrieval, and how Critique Chat uses it today.
Xiaomi MiMo-V2-Flash & Pro: What a Phone Company’s LLMs Mean for AI Code Review
Xiaomi’s MiMo-V2-Flash and MiMo-V2-Pro: costs, coding benchmarks, hybrid attention, 1M context, and how critique.sh uses them for critique, fixes, and agentic review at every tier.
If AI Is Writing the Code, Who's Fixing It?
For the last two years, most of the conversation around AI coding has focused on the wrong layer. In 2026, the more important reality is much bigger: AI is increasingly helping maintain, inspect, debug, and repair software too.
STATE OF AI ENGINEERING · 2026: The Claude Code Inflection Point & What It Means for the PR Review Crisis
Inside the fastest adoption curve in developer tooling history — and why the bottleneck has moved from writing code to reviewing it. Claude Code, PR review crisis, and the infrastructure critique.sh was built to close.