skip to content
Rohan
An engineer reads a long pull request diff while five small robots labelled Router, Service, Repository, Style and DB performance inspect it with magnifying glasses, and a sixth, the Verifier, stamps one finding approved and drops another in a rejected tray.

Claude Code Review in Production: How Our Review Rules Evolved Over Five Months

Claude Code review on every PR: how our rules grew from one prompt to a playbook, six specialist agents and a verifier, plus a template to copy.

Table of Contents

Claude Code review started at my company as a single prompt in a GitHub workflow, and five months later it is a playbook, six specialist agents and a verifier that throws out any finding it cannot prove. Almost every step between those two points was a reaction to something that went wrong: false positives while the rules were still settling, findings on code the pull request never touched, severity that depended on the model's mood, and a reviewer that invented problems when we gave it too many rules at once.

This is how the rules evolved, what each change fixed, and a cleaned-up version of where we ended up that another team can copy.

Claude Code's built-in review, and why we wrote our own rules

Claude Code now has a review of its own, and it is worth knowing before you build anything. Anthropic's managed Code Review, in research preview for Team and Enterprise plans, runs through the Claude GitHub App. Several agents look for different classes of issue in parallel, a verification step filters out false positives, and findings arrive as inline comments tagged Important, Nit or Pre-existing. It never approves or blocks a pull request, and you tune it with a REVIEW.md file in the repository. For a local check there is the Claude Code review command, /code-review, which reviews your branch and uncommitted changes, and can post its findings to the pull request with --comment or apply them with --fix.

We looked at it and chose to keep our own, because we already had a customised flow: rules written for our architecture, a severity ladder tied to those rules, and a review that votes on the pull request. Several of the ideas overlap, such as parallel specialists, a verification pass, and keeping bugs in old code apart from bugs the pull request introduced. The difference is purpose. The built-in review hunts for correctness bugs by default. Ours enforces a written architecture, cites a rule for every finding, and casts a vote on the pull request. If your team wants a second pair of eyes for bugs, start with the built-in review. If you have documented rules you want enforced mechanically, with a verdict, you need a playbook of your own, and that is what the rest of this post is about.

Version 1: one prompt, run on request (April)

The first version went in on 20 April. A workflow listened for a pull request comment containing "@claude" and "review", and ran Claude Code with one prompt. The prompt did three things.

  • For files in our layered modules (routers, services, repositories and DTOs), it read the architecture rule documents for each layer (router, service, repository and a cross-cutting rules file) and treated any violation as blocking.
  • For everything else, it did a general review: correctness, edge cases, type gaps, style, security and obvious performance problems. It was told explicitly not to apply the layering rules to legacy code.
  • It grouped findings by file, each with a severity (BLOCKING, WARNING or SUGGESTION), a location, the problem and a concrete fix, and posted them as a PR comment.

That structure survived every later version. What did not survive was the trigger. At first we were still iterating on the rules, and the reviewer produced false positives, so it ran only when someone asked for it. Once the false positives came down, we made it a gate on every pull request to protect code quality.

Version 2: every pull request, with a verdict (May)

Between 12 and 15 May the workflow changed four times in quick succession.

  • It ran on every pull request, on open, on each new push and on reopen, while still honouring a manual "@claude review" comment.
  • It submitted a real GitHub review instead of a comment: request changes if there was any blocking finding, approve otherwise. The reviewer now had a vote.
  • It skipped draft pull requests and ran when a draft was marked ready for review.
  • It stopped running on long-lived project branches, where a stream of intermediate pushes would each have drawn a review.

The lesson of that week: once a reviewer votes, noise stops being a nuisance and becomes a cost. A wrong "request changes" blocks a colleague. Everything that followed was about precision.

Version 3: load only the rules that matter, and review only what changed

The next change rewrote the prompt as numbered steps, and three ideas from it are still the backbone of the system.

Load rules by path, not all at once. The reviewer first listed the changed files, then read only the rule documents those paths needed: router rules only if a router changed, repository rules only if a repository changed, and the reference patterns only when it had to check one. Loading every rule at the same time caused hallucinations: the reviewer reported violations that were not there. That is why we moved to loading only the rules each change needed.

Review only changed lines. The prompt gained a sentence we have never removed: review only changed or added lines, never flag pre-existing code. Without it, a one-line fix in an old file could come back with twenty findings about code written years earlier, none of them the author's to fix in that pull request.

Severity by definition, not judgement. The severity section was headed "apply exactly, no judgment calls". BLOCKING meant a layering-rule violation, a correctness bug, a security issue or a data-loss risk. WARNING meant correct architecture that was still risky. SUGGESTION meant style or naming that broke no rule.

The same week, the turn limit went from 20 to 40 and the reviewer gained file-reading and search tools, because reviews were running out of turns on larger pull requests before they finished.

Version 4: one playbook for local and CI reviews

On 18 May the rules moved out of the workflow and out of CLAUDE.md into a single docs/review.md, which describes itself as the single source of truth. Both the GitHub Action and a local /review command in Claude Code follow the same file, so an engineer running a review before pushing gets the same rules the pull request will face. Several details from that day are worth copying.

  • Ask for the base branch. Locally, the review asks which branch the work will merge into instead of assuming main, because a feature branch that targets a release branch has a different diff.
  • Review what is not committed yet. The local review covers committed changes, working-tree edits and untracked files.
  • Skip data files. SQL, CSV and JSON files are excluded from the review entirely. They are migrations, fixtures and configuration, and reviewing them as code produced findings nobody could act on.
  • Every finding cites a rule. BLOCKING must cite a specific architecture rule file and section. WARNING must cite a specific style rule from CLAUDE.md. SUGGESTION is for anything no rule covers, and by definition can never be a rule violation.
  • Never block legacy code on style. BLOCKING was restricted to the layering rules. General Python style caps at WARNING everywhere.

Tying severity to a citation changed the character of the reviews. A finding either points at a rule the team agreed, or it is a suggestion the author can ignore.

Version 5: specialists and a verifier

On 19 May the single reviewer was split up. The workflow now starts an orchestrator that reviews nothing itself. It lists the changed files, routes each one to the specialists that own it, runs them in parallel and assembles the result.

SpecialistReviewsOwns
Router reviewerAPI and router files in the layered modulesRouter-layer rules
Service reviewerService files in the layered modulesService-layer rules
Repository reviewerRepository files in the layered modulesRepository-layer rules
Architecture reviewerAny file in the layered modulesCross-cutting architecture rules
Python style reviewerEvery Python fileStyle rules from CLAUDE.md, capped at WARNING
VerifierEvery findingNothing: it only checks the others' work

Each specialist reads only its own rule document and ignores anything outside it, so a service file reviewed by four specialists does not get the same finding four times. They return findings in one fixed format: file, line, severity, rule section, the offending code quoted verbatim, the issue and the fix.

The verifier is the Claude Code review agent I would copy first. For every finding it checks four things, and rejects the finding if any fails:

  1. The cited file exists.
  2. The quoted code appears within five lines of the cited line.
  3. That line was added or changed in this diff, not pre-existing context.
  4. The cited rule section exists and says what the finding claims.

Its instructions open with the principle behind the whole design: false positives are far worse than false negatives. It may not re-grade severity or add findings of its own. Rejected findings are dropped silently; only approved ones reach the pull request.

Adding a specialist without touching the workflow (September)

In September we added a sixth reviewer for database performance, because the most expensive production problems we see are queries that are fast in development and slow at scale. It has its own numbered rule document:

  • Blocking: queries inside loops, N+1 access to related fields, per-row writes where a batch write exists, and raw SQL built inside loops or by string interpolation.
  • Warning: evaluating the same query set twice, unbounded fetches, updates that save every field, aggregation done in Python instead of the database, and filters on unindexed fields in migrations.
  • Carve-outs: small fixed-size loops, test code, and batched data migrations. Where a pattern is deliberate, a # db-perf: allow comment that states the reason suppresses the finding, so the exception and its reason live next to the code.

Two decisions stand out. Database performance findings can block legacy code as well as code in the layered modules, the one deliberate exception to "never block legacy", because a query in a loop costs the same wherever it lives. And the workflow file did not change at all, because it reads docs/review.md at run time. Adding a specialist meant adding one rule document, one agent definition and one row in the routing table.

Knowing when not to review

The last change, a few days later, took something away. Our weekly release pull request promotes already-reviewed commits from one long-lived branch to the next. Any inline comment on it becomes an unresolved conversation, and our main branch requires conversations to be resolved, so a review on that pull request could block the automated weekly merge. The reviewer now skips it. Feature pull requests are still reviewed, and a manual "@claude review" still works anywhere.

What we would tell another team

  • Start with the output format and the severity definitions. They survived every version; the trigger and the architecture did not.
  • Make every finding cite a rule. It turns arguments about taste into a conversation about whether the rule is right.
  • Never flag pre-existing code. Nothing destroys trust in an automated reviewer faster.
  • Split by rule document, not by file. Specialists that each own one document do not duplicate each other and stay small enough to reason about.
  • Verify before you post. A second pass that checks the quote, the line, the diff and the rule is cheap compared with a wrong "request changes".
  • Keep one playbook for local and CI. Engineers should see the same review before they push that the pull request will get.

This connects to the wider argument in six feedback loops for AI coding agents: the review stage only improves the next change if its findings are precise enough to act on. It is also a guard against the problem in why generated tests miss bugs, where passing checks say less than they appear to.

A Claude Code review template you can copy

Below is the structure of our playbook with our company-specific rules removed. Replace the rule documents with your own; the routing, severity ladder, finding format and verifier transfer as they are. In Claude Code, the orchestrator can be a slash command (.claude/commands/review.md), the specialists and verifier are sub-agents (.claude/agents/*.md), and the same playbook can drive a GitHub Action.

# Code Review Playbook (single source of truth for local /review and CI)

## 1. Scope
- CI: the PR diff. Local: ask which branch this merges into; review committed,
  uncommitted and untracked changes against it.
- Exclude data and config artefacts (.sql, .csv, .json). If nothing reviewable
  remains, report "No reviewable code files changed" and stop.
- Review ONLY added or changed lines. Never flag pre-existing code.

## 2. Route (the orchestrator reviews nothing itself)
| Specialist            | Files                         | Rule document             |
|-----------------------|-------------------------------|---------------------------|
| review-api            | API / router files            | docs/rules/api.md         |
| review-domain         | service / domain files        | docs/rules/domain.md      |
| review-data           | repository / data-access files| docs/rules/data.md        |
| review-architecture   | all application code          | docs/rules/architecture.md|
| review-db-performance | all non-test code             | docs/rules/db-perf.md     |
| review-style          | all code                      | style guide (caps at WARNING) |
Dispatch every relevant specialist in parallel. Skip a specialist with no files.

## 3. Finding format (every specialist)
### Finding
- file: <path>
- line: <line in the new file>
- severity: BLOCKING | WARNING | SUGGESTION
- section: <rule document> §<n> <heading>
- quote: <offending code, verbatim>
- issue: <which rule, why it is violated>
- fix: <concrete fix>
Or exactly: NO FINDINGS

## 4. Severity (mechanical, no judgement calls)
- BLOCKING: violates a documented blocking rule. Must cite file and section.
- WARNING: violates a documented warning-level or style rule. Must cite it.
- SUGGESTION: a concern no documented rule covers. Never a rule violation.
- Style never blocks.

## 5. Verify (one verifier, after all specialists)
Reject a finding unless ALL hold:
1. The file exists.
2. The quoted code appears within ±5 lines of the cited line.
3. That line is added or changed in this diff (not context, not removed).
4. The cited rule section exists and says what the finding claims.
False positives are far worse than false negatives. Do not re-grade severity.
Do not add findings. Drop rejected findings silently.

## 6. Deliver
- Group approved findings by file; lead with the verdict.
- CI: request changes if any approved BLOCKING, otherwise approve.
- Local: print the review; never post.

For the GitHub side, two settings carried most of the weight for us: run on opened, synchronize, reopened and ready_for_review but skip drafts; and skip branches where a review cannot help, such as release promotions.

Claude Code review: common questions

Can Claude Code PR review run automatically on every pull request? Yes. Claude Code runs in GitHub Actions through Anthropic's official action, and can be triggered on pull request events or by an "@claude" comment. Ours submits a formal approve or request-changes review.

Claude Code review skill, command or agent: which should I use? They fit together. A slash command is a good entry point for a local review; sub-agents let you split the work so each reviewer holds one set of rules; and a skill suits a reusable review procedure you want Claude to apply whenever it fits. Claude Code's own /code-review command is a good default before you write any of your own.

Can I package this as a Claude Code review skill? Yes. The playbook above works as a skill: put the steps in a SKILL.md, keep the rule documents and the finding format beside it, and let Claude load it when you ask for a review. Keep the specialists and the verifier as sub-agents, because each needs its own context and its own rule document. If you have not built one before, how to create Claude Skills walks through the structure.

What makes a good Claude Code review prompt? Specific rules it can cite, a fixed output format, severity defined by rule rather than judgement, and an explicit instruction to review only changed lines. The template above is our current version.

Does Claude reviewing Claude's code actually catch anything? In our experience it catches rule violations reliably when the rules are written down and cited, which is most of what a human reviewer spends time on. It is weaker at judging whether a change is the right change. That is why a human approval is still required on every pull request: the reviewer's vote is a gate, not a replacement for a person.

Does it replace human review? No. It removes the mechanical part of review, so human reviewers can spend their attention on design, intent and whether the change should exist at all. Review effort is also the cost that headline numbers hide: a 90% agent PR merge rate can still describe a weak workflow if people are quietly doing the real work in review.

Where this leaves us

The system we run today is not clever. It is a playbook, a routing table, a handful of reviewers that each know one set of rules, and a verifier that refuses anything it cannot prove. Every piece exists because an earlier, simpler version failed in a specific way. If you are starting now, you can skip most of those failures: write your rules down, make the reviewer cite them, review only what changed, and check its findings before they reach a colleague.

Stay in the loop

Get practical notes on backend systems, databases, and building with AI in your inbox.

Email subscriptions are handled by Substack. Unsubscribe anytime. Form not loading? Subscribe on Substack.