An AI code review agent that clears the easy findings before a human looks
The agent reviews every pull request within minutes: it reads the diff in the context of the surrounding code, checks it against the rules your team enforces, and leaves specific line-level comments. Senior engineers still approve the merge, without spending that approval on missing null checks.
Engineering manager or tech lead who owns review standards for the repository
A pull request is opened, or new commits are pushed to an open one
01the problem and who owns it
Review is the bottleneck nobody budgets for. Pull requests wait a day for the two people who know the payments module, and half their comments are about naming, missing tests, or an unhandled error path. The design questions only they can answer get a skim.
The tech lead owns review standards, but they live in their head and in old pull request comments. A linter catches formatting; it cannot notice that a new endpoint skips the authorization helper, or that a migration locks a large table during business hours.
02what the AI does, step by step
- Collect the diff and its contextA CI job or app webhook hands the agent the diff, full changed files, the pull request description, the linked ticket, and nearby callers and tests. A diff reviewed without its surroundings produces confident, wrong comments.
- Load the team's written rulesThe agent reads a review guide kept in the repository: architecture boundaries, required auth and logging helpers, migration rules, and banned patterns. Changing the guide is a pull request, so the rules stay versioned.
- Review for correctness and riskThe model looks for unhandled errors, missing authorization checks, unsafe input handling, race conditions, unscalable queries, and tests that assert nothing. Each finding cites exact lines and the failure it would cause.
- Filter before postingA second pass drops low-confidence findings, merges duplicates, and removes nits the linter already enforces. Fewer, better comments are what keep people from muting the tool.
- Post inline comments and a summaryFindings land as inline comments with a suggested change where obvious, plus a short summary of where a human should look hardest. The agent comments; it never approves or merges.
- Learn from resolutionsDismissed or unhelpful comments are logged, and a monthly review of them drives edits to the guide or filter, not silent prompt tweaks.
03systems it connects to
- Source control. GitHub or GitLab, through an app or bot account with read access to code and permission to comment on pull requests.
- CI. GitHub Actions, GitLab CI, or similar to run the review job alongside tests and linters.
- Issue tracker. Jira or Linear, so the agent can read the ticket a change claims to address.
- Model provider. An LLM API under terms that exclude your code from training, or a self-hosted model where policy requires it.
04human checkpoints
- Human approval to merge. Branch protection still requires a human approving review; the agent's account is never a required or sufficient reviewer.
- Review guide ownership. The tech lead approves every change to the rules file, because that file now shapes every review in the repository.
- Security-sensitive paths. Changes to auth, payments, or infrastructure code still require review from a named owner through CODEOWNERS, whatever the agent says.
05what to measure
- Time to first review. From pull request opened to first substantive comment, before and after launch.
- Comment acceptance. Share of agent comments that lead to a code change rather than a dismissal.
- Human review rounds. How many back-and-forth cycles a pull request needs before merge.
- Escaped defects. Production bugs in agent-reviewed code, sampled to see what it missed.
06risks and guardrails
- Prompt injection through code. Comments, strings, or pull request descriptions can carry instructions aimed at the model. Give the agent no write access beyond commenting, and never let its output trigger deploys.
- Secrets in context. Strip environment files and anything matching secret patterns before the diff leaves CI, and keep secret scanning as a separate, deterministic check.
- Rubber-stamp drift. When the agent finds nothing, reviewers may assume nothing is there. The summary should state what it did not evaluate, such as product intent or real-load performance.
07build vs buy
Off-the-shelf AI review apps install in minutes and handle generic correctness and style well. For a small repository with conventional patterns, start there and see what reviewers keep.
A custom agent pays off when the valuable rules are yours: internal frameworks, domain invariants, migration policies, and cross-service contracts a generic tool cannot know. On Pinned Golf, AI automation in the engineering process is part of how the engineering need went from five engineers to one, with review discipline keeping it safe.
08related playbooks
Browse every engineering playbook or the full library.
want this running in your business?
We write your review guide with your tech lead, wire the agent into CI on one repository, and tune what it posts against real pull requests before it touches the rest.
See how we deliver it: ai coding orchestration.
book a call drop your number