Skip to content
Skip to content

Journal · independent review

Field notes on
finish checks.

Product launches, verifier research, and essays on why coding agents need an independent second opinion before they claim a task is done.

Browse comparisons
Comparisons
ComparisonsCursor vs Codex, Claude Code vs Cursor, Copilot vs Cursor, and more.

DeepSeek V4.1 Flash in CritiqueCode: A Completed Review Run

We ran DeepSeek V4.1 Flash locally in CritiqueCode: architecture, inference rates, coding and vision tests, and the schema fix that made the structured review complete.

12 min

Coding Agent Completion Benchmark: Methodology Before Scores

An open methodology for comparing author-only, author-plus-review, and verified-repair coding agent workflows without inventing product superiority.

8 min

OpenCode Code Review Workflow for Local Changes

Review OpenCode changes before commit with explicit intent, complete working-tree scope, executable checks, and an independent finish result.

6 min

Codex Code Review for Local and Uncommitted Changes

A practical Codex code review workflow for staged, unstaged, deleted, and untracked changes before commit or pull request.

6 min

How to Verify AI-Generated Code Before You Ship

A practical checklist for verifying AI-generated code with author checks, independent review, tests, scanners, CI, and human judgment.

7 min

Claude Code Review: Self-Review vs an Independent Second Pass

Compare Claude Code local review, managed pull-request review, and an independent finish pass for uncommitted changes.

6 min

AI Code Review CLI: Review Uncommitted Changes Before You Push

Use an AI code review CLI to inspect staged, unstaged, deleted, and untracked work before you push. Compare self-review, independent review, CI, and local inference.

8 min

Bringing the Frontier to Critique

Critique adds OpenAI GPT-6 Astra and Anthropic Claude Fable 5.1 to the managed model catalog, with transparent frontier routing and published per-million-token pricing.

7 min

CritiqueCode 2.0: The Harness Decides What Is True

CritiqueCode 2.0 turns coding-agent claims into state-bound engineering records: requirements, controlled checks, evidence, review, repair, and honest incomplete outcomes.

9 min

One Model, Seven Harnesses, 140 Trials

We ran 140 Harbor trials in fresh E2B sandboxes with z-ai/glm-5.3-flash across seven coding harnesses, retaining task-level outcomes, costs, trajectories, tool activity, timing, patches, and verifier evidence.

16 min

Same Model. Four Harnesses. What Actually Survives the Verifier?

We held Mercury 2.5 constant and changed the coding harness. This open study reports 40 Harbor/E2B trials, independent F2P/P2P verification, task-level failures, costs, telemetry, and the complete reproduction trail.

12 min

Critique Auto Gets a New Model Roster

Critique Auto now routes across refreshed coding models, with current token pricing, shared SWE-bench context, four hosted routes in CritiqueCode, and credit access to the full catalog.

28 min

CritiqueCode Hosted Runtime: An Early, Truthful Release

The first hosted CritiqueCode runtime is an early implementation with isolated E2B sandboxes, GitHub branches, and explicit availability checks.

6 min

CritiqueCode 0.1.0 Shipped a Contract. This Is Everything After.

CritiqueCode 0.1.3 vs 0.1.0: website login, critique/auto, voice, imported skills, critique_task workers, and a local web UI. Forced review did not move.

16 min

CritiqueCode Leaves the Terminal Without Leaving Your Machine

critique-code web serves CritiqueCode in a local browser on 127.0.0.1. Same author kernel, same /review gate, no OpenCode fork and not Critique Cloud.

11 min

Talk to Code with CritiqueCode Voice Mode

Talk to CritiqueCode instead of typing. /voice transcribes with Qwen3 ASR 0.6B; spoken review, repair, and ship still run as commands. Start --voice.

8 min

Cursor vs Codex

Cursor vs Codex in 2026: pick the IDE Agent if you live in the editor, or OpenAI’s CLI and optional cloud if you work in the terminal. Fair comparison.

12 min

Best CLI Tools for Coding in 2026

Compare git, gh, Claude Code CLI, Codex CLI, Aider, OpenCode, and CritiqueCode. Best CLI tools for coding in 2026 depend on the job: VCS vs AI author.

13 min

CritiqueCode Is the Coding Agent That Cannot Grade Its Own Homework

CritiqueCode is Critique’s author agent: a Claude Code-style coding CLI that implements in session, then the controller forces independent review and verified repair. Install @critiquedotsh/harness and run critique-code.

12 min

Best Terminal Coding Agents in 2026

Compare the best terminal coding agents in 2026: Claude Code, Codex CLI, OpenCode, Aider, Gemini CLI, and CritiqueCode. Pick by surface, models, and review contract, not invented benchmarks.

14 min

Best Open Source Coding Agents in 2026

Compare the best open source coding agents in 2026: OpenCode, Aider, Cline, Continue, Goose, OpenHands, Gemini CLI. Licenses, surfaces, and where CritiqueCode fits.

15 min

Best Coding Agent in 2026: Claude, Codex, Cursor

Compare Claude Code, Codex, Cursor, OpenCode, and CritiqueCode in one 2026 guide. Pick the best coding agent by jobs: IDE, CLI, or author-plus-review.

14 min

OpenCode vs Claude Code

OpenCode vs Claude Code: MIT model-agnostic terminal agent at opencode.ai versus Anthropic’s Claude Code. Fair, sourced comparison of license, models, and review.

11 min

Aider vs Claude Code

Aider vs Claude Code: open-source git-centric pair programming at aider.chat versus Anthropic’s Claude Code. Fair, sourced comparison of license, models, and review.

11 min

Claude Code vs Codex

Claude Code vs Codex CLI in 2026: Anthropic terminal agent vs OpenAI Codex. Install, review, automation, when to pick each, and a sidecar if you keep either.

10 min

Claude Code vs Cursor

Claude Code vs Cursor in 2026: Anthropic’s terminal agent versus Cursor’s AI IDE. Fair comparison of surfaces, tools, pricing, and when to keep each.

12 min

CritiqueCode vs Claude Code: When to Keep, Switch, or Combine

CritiqueCode vs Claude Code: keep Anthropic's terminal agent, switch to critique-code, or pair Claude Code with Critique CLI sidecar. Review vs evidence.

11 min

CritiqueCode vs OpenCode

CritiqueCode vs OpenCode: open-source model-agnostic terminal agent versus critique-code with forced review. Fair, sourced, MIT OpenCode vs @critiquedotsh/harness.

10 min

CritiqueCode vs Codex

CritiqueCode vs Codex: keep OpenAI Codex CLI, switch to critique-code, or pair Codex with the Critique CLI sidecar. Review vs evidence, install, and when to switch.

11 min

CritiqueCode vs Cursor

CritiqueCode vs Cursor: keep Cursor Agent and Composer in the IDE, switch to critique-code, or pair Cursor with Critique CLI sidecar. Review vs evidence.

11 min

Best Claude Code Alternatives in 2026

The best Claude Code alternatives in 2026 depend on CLI, IDE, or forced independent review. Compare Codex, Cursor, OpenCode, Aider, Gemini CLI, CritiqueCode.

14 min

GitHub Copilot vs Cursor

GitHub Copilot vs Cursor in 2026: VS Code and GitHub seats versus Cursor’s AI IDE, Tab, Agent, and Composer. Official list prices, a fair two-column table, and when to keep each.

12 min

Cursor vs Codex vs CritiqueCode

Compare Cursor vs Codex vs CritiqueCode in 2026: IDE agent, OpenAI CLI/cloud loop, or local authoring whose controller forces review and verified repair.

11 min

Stop putting Kimi on a rename

Critique Auto is a session coding selector on the Inference API. It routes among Critique-resold models only: efficient specialists first, frontier coding models after struggle, billed at the routed model.

9 min

From git diff to First Customer

Your AI coding agent finished the code. Here’s the stack for reviewing AI-generated code, deploying to production, presenting your product and finding your first customers.

16 min

Introducing CritiqueCode: the hosted Code workspace

CritiqueCode Beta is the Critique-owned hosted Code workspace. It gives code changes explicit task contracts, guarded tools, durable artifacts, and evidence-backed outcomes.

7 min

Critique stopped wrapping the agent and started owning the review

Critique CLI now bundles a pinned OpenCode runtime behind a typed SDK, keeps the main coding agent out of reviewer judgment, captures local evidence through a Critique plugin, and provides a collaborative task spine and second brain.

10 min

Critique CLI Gets a Conversational Second Brain

Critique CLI adds read-only ask and chat sessions for coding-agent second opinions, plus stricter budgets, events, and verified repair honesty.

7 min

Critique 0.2: Review Less, Test the Right Risk

Critique CLI 0.2 adds focused AI code review modes for quick checks, task-aware verification, security testing, and resilience testing, with clearer terminal guidance for people and coding agents.

6 min

Critique vs Built-In Agent Checks

Compare Critique with built-in Codex, Claude Code, and OpenCode checks: fast self-verification versus an independent, evidence-backed finish pass.

12 min

critique.sh vs CI: Different Layers of Proof

Critique is a pre-CI or alongside-CI verification layer for coding agents, not a CI replacement. See where each belongs in an engineering workflow.

11 min

Best Coding Agent Verification Tools in 2026

An opinionated 2026 guide to coding agent verification: Critique, built-in agent checks, CI, security scanners, and human review.

14 min

Your Coding Agent Is Lying to You (It Just Doesn’t Know It)

Why coding agents can confidently report completion without proving it, and how an independent finish loop replaces self-report with evidence.

10 min

Agents Grade Their Own Homework. That’s Broken.

Why self-grading coding agents create avoidable blind spots, and how a narrow independent verifier improves agent-written code without slowing the primary loop.

10 min

We Killed the Dashboard. It Wasn’t the Workflow.

Why Critique shifted from a broad human-facing code review surface to an agent-facing completion layer that verifies a coding agent’s finish claim.

11 min

Critique Is Now the Teammate Your Agent Calls

Critique is now an agent-facing CLI for independent verification and repair. Meet the new finish loop for Codex, Claude Code, and OpenCode.

11 min

LLM Gateway BYOK on Critique: Same $8 Harness, Broader Model Surface

Critique adds LLM Gateway BYOK for review, Remedy, Builder, and chat model calls: same $8 harness, encrypted key storage, broader OpenAI-compatible model gateway, and LLM Gateway takes precedence when multiple BYOK keys are saved.

9 min

The senior developer shortage is coming. We're engineering it ourselves.

Entry-level tech hiring collapsed while AI productivity soared. The five-to-seven year mentorship lag means the senior developer drought hits around 2029–2031 — and every company that ran the spreadsheet will be competing for a pool they collectively drained.

9 min

Critique Coding Agent API: How Teams Are Actually Using Cloud Agents Over HTTP

Critique’s June 2026 Coding Agent API update adds lifecycle fields, tags, metadata, intent classification, better SSE status, and cleaner follow-up behavior — plus a clearer view of how teams are wiring cloud coding agents into CI, support, and internal platforms.

10 min

Critique Is Going Open Source

Critique is launching a public community edition focused on self-hosted GitHub pull request review. This post explains what is opening first, what is staying out of scope, and why we are using a separate public repository.

7 min

Critique Intake: Feedback That Arrives Already Debugged

Introducing Critique Intake, an agentic bug intake workflow for turning vague user reports into triaged engineering packets, agent prompts, Review runs, and Change Passports.

11 min

Critique v6.1.1: Agent Stack Integrations — When the Merge Gate Meets RWX, Swytchcode, and Iron Book

Product update: Critique v6.1.1 ships upstreamSignals on POST /api/v1/reviews, partner cookbooks for RWX / Swytchcode / Identity Machines, passport provenance, and docs for the full agent stack loop — build, judge, merge, execute.

28 min

The Merge Gate API: Why Agent Teams Need a Judge That Is Not the Writer

A deep essay on Critique’s Merge Gate API — why writer and judge must be separate roles, what v6.1 ships (REST, MCP, webhooks, findings=all), and how merge governance beats self-review loops as coding agents multiply PR volume.

22 min

GLM-5.2 Lands in Critique: 1M Context, Design Arena #1, Same 3-Credit GLM-5.1 Shelf

GLM-5.2 replaces GLM-5.1 and GLM-5V-Turbo in Critique’s catalog at the same 3-credit review floor. Z.ai benchmark table vs GPT-5.5 and Claude Opus 4.8, Design Arena Elo 1360, long-horizon FrontierSWE / PostTrainBench / SWE-Marathon charts, and how the inference API still exposes `z-ai/glm-5.1` routing to GLM-5.2 upstream.

22 min

You Merged That. Why.

You knew it wasn't ready. You merged anyway. Why AI velocity outpaced review, what actually happens to your PRs, and how Critique closes the gap with sandbox verification and merge policy.

6 min

Kimi K2.7 Code Lands in Critique: Open-Source Coding at 4.5 Credits, With Moonshot’s Full Benchmark Table

Kimi K2.7 Code joins Critique’s Moonshot lane at 4.5 credits per PR review run (K2.6 stays at 4). Vendor-reported coding and agentic benchmarks vs GPT-5.5 and Opus 4.8, MCP Atlas cross-reads against Gemini 3.5 Flash, MiniMax M3, and Qwen3.7 Max, and reasoning-efficiency gains over K2.6.

24 min

Why Your AI-Generated Next.js App Works Locally but Dies on Vercel

Next.js works locally but fails on Vercel? Case-sensitive imports, env var gaps, and next build vs next dev — plus a 10-minute pre-push ritual to catch AI import mistakes.

8 min

Breaking the Dead Loop: What to Do When Cursor or Windsurf Refuses to Learn

Cursor or Windsurf stuck on the same broken fix? Loop taxonomy, two-strike rule, circuit breaker prompts, TDD ground truth, and Critique sandbox verification.

8 min

How AI Code Generators Create Circular Imports in Next.js (and How to Spot Them)

Five AI patterns that create circular imports in Next.js, why local dev lies, how to detect cycles with madge and dependency-cruiser, and the fix order that actually works.

9 min

Code and Pray Is Not a Workflow: Why Vibe Coders Need Sandbox Verification

Static review vs runtime proof, the Plan-Work-Verify loop, E2B and Docker sandbox patterns for GitHub PRs, and how Critique productizes sandbox verification for vibe coders.

8 min

The Vibe Coder's Security Checklist: 5 Ways Your AI IDE Is Leaking Secret Keys

Five common ways Cursor, Copilot, and agent mode leak API keys into git diffs, model context, client bundles, and CI logs — plus a PR merge-gate checklist to prevent environment variable leaks.

8 min

Silent TypeScript Failures: The Danger of 'Fixed' by Any

Why AI agents silence TypeScript errors with `any`, `as`, and `@ts-ignore`, how Next.js dev vs build diverge, and a PR checklist for real type fixes — not check-engine-light patches.

9 min

Why We Built the Critique Inference API: Western Servers, NVIDIA-Powered Capacity, and Credits You Already Own

Critique Inference API launches at /inference-api with DeepSeek V4 Flash, Tencent Hy3 Preview, and NVIDIA Nemotron 3 Ultra on Western sweetener servers — stress-tested, privacy-first by default, and billed from Critique credits. Background from Leemer Labs on why we did not wait for enterprise timing.

18 min

Critique v5.1: The Review System Starts Reviewing Itself

Critique v5.1 ships a review evaluation harness, repository learnings, a quality budget for findings, an independent critic pass, E2B/OpenCode runtime drift checks, and a clearer Coding Agent API surface.

16 min

Critique v5.1: We Hit Go, Walked Away for Five Hours, and Shipped While Reviewing Ourselves

Critique v5.1.0 (5 June 2026) ships operator-first review runs, one Coding Agent product across Builder and API, marketplace attribution, repo-home defaults, and Platform connections beta — documented in a ship log written largely by cloud agents that reviewed their own changes for nearly five hours after a human pressed go.

22 min

Critique vs Devin for Cloud Coding Agents: API, Pricing, and Model Freedom

Compare Critique Coding Agent API and Devin on API access, pricing (quota vs credits vs OpenRouter BYOK), model choice including OpenRouter free routes, and when each cloud coding agent fits.

8 min

Critique v5 Beta: Marketplace, Merge Policy in Plain English, Passport Exports, and the Biggest Ship Since v4

Critique v5 beta (v5.0.0) ships the Agent Skill Marketplace, NL merge policy compiler, Change Passport HMAC exports, MiniMax M3 welcome pricing, Cursor Agent SDK BYOA, Coding Agent API persistent sessions, repo-first PR dashboard, finding feedback, opt-in model-feedback sharing, Workspace agent queue, and Insights velocity/risk/cost/compliance — the largest platform update since v4 Change Control.

45 min

Coding Agent API: Persistent OpenCode Sessions for Multi-Turn Automation

How Critique Coding Agent API persistent sessions work: warm E2B sandboxes, OpenCode continuity, idle status, SSE streaming, and follow-up messages for CI bots and internal agent platforms.

11 min

Cursor as a Top-Tier Agent Harness: Composer 2.5, Cloud BYOA, and How It Compares to the Models on Critique

Why Cursor is a first-class agent harness, what Composer 2.5 is (specs, SWE-Bench Multilingual, Terminal-Bench 2.0, pricing), how Critique queues cloud agents via the Cursor Agent SDK on composer-2.5, and how that compares to models on the Critique review catalog.

22 min

MiniMax M3 and Qwen3.7 Plus on Critique: Coding Benchmarks and a Two-Week M3 Welcome Price

MiniMax M3 coding performance on SWE-Bench Pro and Terminal-Bench 2.1, M3 vs Opus 4.8 and GPT-5.5, Qwen3.7 Plus vs Qwen3.6 Plus on terminal and UI benchmarks, and a 50% welcome credit window on M3 PR review runs through June 17, 2026. Critique Chat stays Ling and DeepSeek V4 Flash.

28 min

Best Code Review Skill for Claude Code, Hermes, Codex, and Opencode

Best code review skill for Claude Code, Codex, Hermes, and Opencode, with a real Moonshot Kimi K2.6 PR review comparison and setup guide.

24 min

Critique v4.1: Your Merge Verdict, Where Your Team Already Works

Critique v4.1 connects to Linear, Slack, Zapier, and the tools you already use. One review verdict across GitHub, chat, and your stack—without asking engineers to learn a new destination.

18 min

Getting Hundreds of PRs and No Time to Review Them? Open Source Needs PR Control.

Open source maintainers and foundations facing hundreds of pull requests need PR control: gate, evidence-backed review, and Change Passports. Pro and Team plans for volume; verified OSS lane at Solo-equivalent pricing. OSS credit grants available.

12 min

Critique v4: The AI Change Control Platform (and What Happened to “Just Review”)

A deep guide to Critique v4 vs v1–v3: Change Passports, the Control Board, gate → review → merge, AI-powered sandbox evidence runs, and why Critique is building the best AI management layer for code — not another comment bot on your diff.

24 min

v3.6.0: Bring Your CrofAI Key — Same $8 Harness, Direct OSS Billing

Critique v3.6.0 adds CrofAI BYOK alongside OpenRouter: save a Crof API key, bill nahcrof directly at https://crof.ai/v1, and keep the same $8/month Critique harness for orchestration.

9 min

Critique Welcomes Claude Code: Review Here, Execute on Claude

How Critique hands a scoped review blueprint to Claude Managed Agents—encrypted Anthropic keys, GitHub-mounted PR checkout, the managed-agents-2026-04-01 beta, and a worker that streams until the session goes idle. Review stays on Critique; execution bills on your Anthropic account.

18 min

Critique Welcomes Codex: Review Here, Execute on Your OpenAI Account

How Critique hands review findings to OpenAI Codex: a deterministic JSON handoff envelope with allowed write paths, validation commands, and stop conditions; an encrypted OpenAI key; and a Responses API worker. Review stays on Critique; execution bills on your OpenAI account.

17 min

Critique Welcomes Cursor: Review Here, Fix There, Ship Together

Critique now queues scoped fix handoffs to Cursor Cloud Agents with your API key. Why review and execution belong in different products—and how to get the best developer experience by using both.

14 min

Introducing Workspace: one place for review, chat, build, and repair

Critique Workspace merges PR review, Chat, Builder, and Remedy into one signed-in surface at /workspace — one history rail, shared repo context, and one credit ledger.

10 min

DeepSeek & MiMo Are 0.5 Credits Forever — Plus Opus 4.8, Qwen3.7-Max, and 2026's Cheapest Frontier Review Stack

Critique’s May 2026 model spring: permanent 0.5cr DeepSeek V4 Flash and MiMo v2.5, 1cr DeepSeek V4 Pro and MiMo Pro, Ling at 0.5cr, MiniMax M2.7 at 1.5cr, Qwen3.7-Max at 6cr, Claude Opus 4.8 and Fast, Gemini 3.5 Flash, Grok Build 0.1, and Ring-2.6-1T.

22 min

Is critique.sh Safe for Private Repositories?

Is critique.sh safe for private repos? A code-level review of GitHub access, E2B sandboxes, data retention, chat storage, and provider logging.

16 min

AI Code Review Pricing Is Getting Weird: What Teams Actually Pay in 2026

The 2026 AI code review pricing trap: seats, usage billing, Actions minutes, and shared credits all make pull request review cost behave differently.

11 min

Before You Merge AI Code, Run This Security Checklist

A security checklist for AI-generated code that catches the scary stuff before merge: auth gaps, data leaks, secrets, dependencies, prompt injection, and unsafe fixes.

12 min

CodeRabbit vs Cursor Bugbot Is the Wrong Fight

CodeRabbit vs Cursor Bugbot sounds simple until pricing, IDE lock-in, benchmarks, and GitHub workflow fit change the answer.

12 min

Your Startup Probably Needs AI Code Review Before It Hires QA

A startup-focused AI code review guide for teams shipping too fast to manually inspect every risky pull request.

10 min

Stop Letting the Agent That Wrote the Code Review It Too

Why coding agents and PR review agents should be separate jobs, and how to wire them into GitHub without trusting the same agent twice.

9 min

The Honest Reason AI Code Review Cannot Be Flat-Rate Forever

Why Critique moved to clearer AI code review pricing: bigger credit pools, transparent review costs, low-balance gates, and an $8/month OpenRouter BYOK harness.

10 min

A fairer credit system for Critique

Critique now prices one standard review unit at roughly 1M input and 150k output tokens instead of 100k and 15k. That makes large, context-heavy reviews much cheaper without flattening the model ladder. We also refreshed the pricing page and are giving every user +1,000 bonus credits.

8 min

Critique Checkpoint: stop slop at the gate

Critique Checkpoint is the new pre-review trust layer for GitHub pull requests: contributor, account, activity, language, and slop-pattern rules before deep Critique review starts.

8 min

The Review That Refuses to Ghost You

Inside Critique's two-layer PR review watchdog: DeepSeek V4 Flash supervision, recoverable stalls, GDPval-AA vs Claude Sonnet 4.6 and Opus 4.6, vendor list pricing anchors, limited-time fifty-percent credit lanes on Flash and Pro, and why speed-per-cognition is what merge-gate infra needs.

14 min

One month. A billion tokens. What it actually took.

A founder note on the first serious beta month: roughly a billion tokens processed, hundreds of thousands of lines written, five distinct surfaces now talking to each other, and what comes next when they finish becoming one.

14 min

Qwen 3.6 and Grok 4.3 in Critique: cheaper routing, stronger coding signal, and one absurd xAI price cut

Catalog refresh for Critique and Remedy: Grok 4.3 replaces Grok 4.2 with a steep credit drop, Qwen3.6-35B-A3B replaces Qwen3.5-27B, InclusionAI Ling-2.6-Flash joins at 1 credit, Qwen3.6-Max-Preview joins at 8 credits, GLM-5.1 gets a 1-credit discount, KAT Coder Pro V2 returns to its normal 2-credit price, and MiMo v2.5 increases by 1 credit.

16 min

From ManusAI, v0, and Lovable to code you can merge

If you ship with ManusAI, v0.app, or Lovable, you already have velocity. Here is how Critique fits after export: GitHub-native PR review with repo context, specialist lanes, and optional verified fixes so prototypes mature into production discipline.

8 min

Referral rewards are live: credits and cash for early advocates

The Critique referral program is up. Get your link on the referrals page, earn bonus review credits for every successful signup, and race for early milestones: 5,000 credits + $50 at 10 referrers, 10,000 credits + $150 at 100.

3 min

Critique PR Review v5: Diff-First Runs & Multi-Turn OpenCode

PR Review v5 leads with the PR diff and model findings, adds multi-turn OpenCode in one session with bounded follow-ups, quiets the default live feed, and sketches the next layer: a cheap thinker (e.g. DeepSeek V4 Flash) for meta-review.

12 min

GPT-5.5 and GPT-5.5 Pro in Critique: Benchmarks, Pricing, and When to Spend the Credits

GPT-5.5 and GPT-5.5 Pro are now visible in Critique: benchmark tables for Terminal-Bench 2.0, SWE-Bench Pro, BrowseComp, GDPval, long context, pricing, OpenRouter IDs, and practical routing guidance for PR review and Remedy.

24 min

DeepSeek V4 Flash and V4 Pro in Critique: 1M context, EU inference, and open-weights leadership on GDPval-AA

DeepSeek V4 Flash (1 cr) and V4 Pro (3 cr) replace V3.2 Speciale in Critique’s runtime catalog: MoE scale, 1M-token context, EU-hosted routing, GDPval-AA and token-budget benchmarks vs V3.2 and frontier peers, DeepSeek API pricing vs Critique credits, and integration notes for lead, specialist, and Remedy roles.

29 min

Introducing Critique Builder: where English compiles

Critique Builder is a browser-first cloud agent surface for starting repo-backed coding jobs, watching live execution, and inspecting results without turning the product into a browser IDE. Managed sandboxes, OpenCode, subagents, and a stronger bias toward shipping real software.

9 min

Critique PR Review v4.1: Execution Depth, Live Observability, and Operator Secrets

PR Review v4.1 deepens the OpenCode sandbox: exploratory testing, granular GitHub inline comments, live OpenCode activity streaming to the dashboard, and AES-256-GCM encrypted per-repository secrets for real integration runs.

20 min

Critique PR Review v4: The Beta Is Getting Real

Critique PR Review v4 is live in beta: nearly 4,000 PRs reviewed, sandbox-native OpenCode execution getting tighter, and Critique now showing up as the #2 public app this month on OpenRouter’s Embed V1 4B activity view just behind AnythingLLM.

8 min

Arcee Trinity-Large-Thinking Lands in Critique at 1 Credit

Arcee AI’s Trinity-Large-Thinking is now available in Critique as `arcee-ai/trinity-large-thinking` at a 1-credit floor. We break down the Apache 2.0 release, MoE architecture, reasoning-trace constraints, benchmark profile, and why this model matters for long-horizon agentic review and Remedy.

19 min

Kimi K2.6 Lands in Critique: PR Review, Remedy, and a Week of Half-Price Credits

Kimi K2.6 replaces Kimi K2.5 across Critique’s PR review stack and Remedy chat models: architecture, benchmarks vs K2.5 and peers, integration notes, and April 21–27 2026 introductory pricing at 50% off former Kimi K2.5 credits.

22 min

Welcome Gemma-4-31B and MiMo-V2-Flash: The Permanent 0.5cr Review Tier

Gemma-4-31B and MiMo-V2-Flash join critique.sh at a permanent 0.5cr floor. We compare SWE-bench Verified scores and critique.sh credit costs across Gemini 3 Flash, MiniMax-M2.5, Qwen3.5-27B, GPT-5.4 mini, Claude Haiku 4.5, GPT-5.1 Codex Max, and Claude Sonnet 4.6 — with a price-to-performance ratio chart.

18 min

Critique PR Review v3.1: The Sandbox Now Finishes the Review

Critique PR Review v3.1 moves final review synthesis into the sandbox. The sandbox now authors the structured review artifact with summary, findings, command timeline, runtime checks, and sub-agent output, while the app stays the control plane.

12 min

Hello to v3 of PR Review

Critique PR Review v3 is live one day after the V2 release. Sandbox-backed AI code review, dedicated review-analysis workers, canonical review-run pages, GitHub App publishing, Remedy handoff, fix prompts, and an owner-only review assistant are now part of the product.

9 min

Critique is now conversational

Critique now supports interactive PR chat via @critique mentions in GitHub pull request comments. Ask questions, run slash commands for review, explain, fix, security, tests, and more. Powered by Qwen 3.7 Plus — a frontier-level model with a 1M-token context window that beats GPT-5.1-Codex, GPT-5.2, and GPT-5.3-Codex on long-horizon benchmarks.

9 min

Critique just got a whole lot better

We shipped ten deep improvements to the Critique review engine: a semantic index, subsystem clustering for large PRs, intrinsic-risk-based drill-down targeting, cross-file relationship analysis, a CODE_QUALITY specialist, two-phase reasoning, domain heuristic packs, contract-diff analyzers, better evidence packaging, and expanded model support across GPT-5.4, Grok-4.2, GLM-5V-Turbo, MiMo v2.5, and more.

11 min

The AI Code Review Trust Gap

AI code review is now mainstream, but developer trust still lags dangerously behind adoption. Here is why engineering leaders should evaluate AI review tooling as governance infrastructure — and what the real buying criteria look like in 2026.

14 min

Remedy, Now in Critique Chat

Remedy — Critique's autonomous code-fixing engine — is now available directly inside Critique Chat. Start from a review, choose a model, and watch fixes land on your PR branch in real time. Your machine never runs a line of this code.

7 min

Hybrid Code Retrieval: Lexical Precision + Semantic Recall

A technical deep dive into hybrid retrieval for codebases: exact string search, semantic search, warmed snapshots, reciprocal rank fusion, and the operational details that make code search trustworthy.

12 min

5 Questions Before Installing an AI Code Review GitHub App

Staff engineer or platform lead evaluating AI PR review automation? These five questions — permissions, token model, data egress, check behavior, and uninstall path — separate a trustworthy GitHub App from a risk review waiting to happen.

9 min

Our Faith in Open Source

A detailed critique.sh essay on open-source and open-weight AI models, MiniMax M2.7 vs Claude Opus 4.6, why Chinese labs are leading the cost-capability curve, and the March 2026 credit promos for MiniMax, GLM-5, and Kimi K2.5.

18 min

MCP vs CLI for AI Agents: Efficiency, Governance, and When Each Wins

MCP and CLI are often framed as rivals for AI agent tool invocation. Primary benchmarks show CLI can be 9–32× cheaper in tokens for well-known tools, while MCP wins on discoverability, OAuth, and tenant isolation. This article synthesizes the evidence and ends with a decision framework.

12 min

Critique Chat: Ask Your GitHub Repo Questions with Frontier Models

Critique Chat is the dashboard assistant grounded in your real GitHub codebase. Connect a repo, ask questions, get answers backed by live code search. Free for all signed-in users with frontier models like KAT Coder Pro V2, MiniMax M2.7, GLM-5V-Turbo, and MiMo v2.5.

5 min

Cloud Coding Agents: What Changed, What Stayed the Same, and Where Remedy Fits

An honest map of cloud coding agents across GitHub, Cursor, Jules, Devin, OpenCode, E2B, and Vercel Sandbox, plus where Critique Remedy fits as a review-attached execution layer with BYOA support.

15 min

Relace Search: The 256K-Context Subagent Behind Serious Codebase Retrieval

Deep dive on Relace Search (relace/relace-search): 256K context, parallel tool calls, report_back handoff to oracle agents, OpenRouter pricing and telemetry, benchmarks vs naive retrieval, and how Critique Chat uses it today.

14 min

Xiaomi MiMo-V2-Flash & Pro: What a Phone Company’s LLMs Mean for AI Code Review

Xiaomi’s MiMo-V2-Flash and MiMo-V2-Pro: costs, coding benchmarks, hybrid attention, 1M context, and how critique.sh uses them for critique, fixes, and agentic review at every tier.

21 min

If AI Is Writing the Code, Who's Fixing It?

For the last two years, most of the conversation around AI coding has focused on the wrong layer. In 2026, the more important reality is much bigger: AI is increasingly helping maintain, inspect, debug, and repair software too.

8 min

STATE OF AI ENGINEERING · 2026: The Claude Code Inflection Point & What It Means for the PR Review Crisis

Inside the fastest adoption curve in developer tooling history — and why the bottleneck has moved from writing code to reviewing it. Claude Code, PR review crisis, and the infrastructure critique.sh was built to close.

14 min