Every AI code review tool I have used or built reaches the same problem on a busy repository. The individual comments can certainly be useful, but after a few pushes the review becomes harder to use.
The same finding returns, old comments remain beside newer ones, and some findings report a problem without giving the reviewer enough evidence to verify it. Better model output will not fix most of this because the missing part is review state: the history that connects one review of a pull request to the next.
Re-reviewing the same pull request creates noise
GitHub’s documentation says Copilot may repeat previous comments during a re-review, including comments that have already been resolved or downvoted.
It becomes more noticeable when automatic review of new pushes is enabled. A pull request that changes five or six times can accumulate feedback from several versions of the code, leaving the developer to decide which comments still apply.
The same problem appears quickly in a custom reviewer. I have previously shown how to run the GitHub Copilot SDK inside GitHub Actions. A basic implementation can fetch the change, call the model and publish the result. Unless it deliberately preserves state, the next run has little connection to the previous review.
GitHub has been improving this experience. It added High, Medium and Low severity labels and grouped similar comments in May 2026. On 18 September 2026, it released a clearer review overview and smarter auto-resolution. The overview now separates open findings from those resolved since the last review and marks findings introduced by a new commit.
Those changes address the experience around the model rather than assuming another prompt will remove the noise. A custom reviewer needs the same sort of state if it is going to run more than once on a pull request.
One bot should own the current review summary
For a summary or status comment, I prefer the bot to own one comment and update it as the pull request changes. A hidden marker gives the workflow something stable to search for:
<!-- ai-review:summary -->
The workflow finds the existing bot-owned comment and replaces its contents instead of adding another one. I covered the implementation in How to update a pull request comment from GitHub Actions instead of creating duplicates.
I would match both the marker and the expected author. Matching visible text is brittle, while matching only the marker risks editing a copied comment belonging to somebody else.
The summary should also have one writer. If several matrix jobs update it independently, the last job to finish can overwrite newer output from another job. Aggregate the results, then let one job publish the current state.
A single summary solves the pull request status problem, but each finding still needs an identity of its own.
A finding needs an identity that survives another push
A line number tells the developer where a problem currently exists, but it is a poor long-term identity. Lines move and commit-specific information changes with every revision. Matching on either makes an existing issue look new after a small edit.
For custom review automation, I would assign each finding a stable identifier and keep its current location as separate data. The identifier could be stored in the comment as a hidden marker:
<!-- ai-review:security:sql-injection:src/auth.py -->
The exact format is less important than the fields behind it. A rule identifier and file path will usually survive pull request churn better than a line number or commit SHA. Where the same rule can occur more than once in a file, I would add a symbol or a fingerprint of the relevant code.
The next run can then reconcile its findings with the previous result:
- Refresh findings that still reproduce with their current location and evidence.
- Mark findings that have disappeared as resolved.
- Create feedback for new findings.
- Group repeated instances of the same underlying issue.
- Suppress findings dismissed as incorrect, or require a higher-confidence result before reopening them.
This does not require a particularly elaborate data model. The review run can be keyed by repository, pull request and head SHA, while the finding is keyed by its stable identifier. Line numbers, commit SHAs and snippets belong to the current occurrence of that finding rather than its identity.
Severity should have a meaning outside the model response
I would keep the severity scale small and display it consistently:
| Severity | Behaviour |
|---|---|
| High | Eligible to block merge |
| Medium | Requires review but does not automatically block |
| Low | Advisory |
| None | No material finding |
The reviewer classifies the finding, while repository policy decides whether that classification affects CI. Keeping those responsibilities separate means the same analysis can begin as advisory feedback and later become a required check without rewriting the review logic.
It also prevents the wording generated by the model from becoming policy by accident. A forceful explanation should not block a merge unless the finding meets the repository’s defined criteria, and a cautiously worded High finding should not escape the same policy.
I would resist adding more levels because the model can produce them. Arguing over whether a finding is Moderate or Significant adds classification detail without changing what the team needs to do.
A useful finding includes the evidence needed to check it
This comment leaves most of the work with the reviewer:
This function may contain a null reference.
It does not identify the relevant path through the code or explain what makes the value nullable. A useful version carries enough evidence to verify the claim without repeating the whole diff:
High: possible null dereference in src/orders.ts
order.customer.name is accessed after customer can be returned as null by findCustomer().
Handle the null case before accessing name.
I apply the same principle to infrastructure review:
Medium: database backup retention is being reduced
backup_retention_days changes from 14 to 2 on example_database.orders.
Confirm the shorter recovery window matches the requirements for this environment.
The useful parts are concrete:
- the severity and rule being applied;
- the affected file, resource or symbol;
- the change that supports the finding;
- the likely impact;
- a remediation or verification step.
Where the remediation is mechanical, GitHub’s suggested-change format lets the author inspect and apply the edit directly.
Missing context needs a harder boundary. If the reviewer did not receive a field, file or configuration, it should not invent a defect based on the absence. It can report that something could not be verified, but uncertainty is not evidence that the code is wrong.
AI approvals turn poor comments into a governance problem
GitHub put Copilot code review approvals into public preview on 1 September 2026. Every Copilot review now contains an approval assessment, and administrators can allow Copilot to submit an approval that satisfies a repository’s required-approval rule.
The feature is off by default and can be controlled at enterprise, organisation and repository level. When new commits are pushed, Copilot’s approval is dismissed in the same way as a human approval.
Duplicate findings and weak evidence are irritating while the reviewer is advisory. Once its decision can contribute to whether a pull request is allowed to merge, they become governance concerns.
I have written before about review burden as part of the real cost of AI-assisted engineering. Every low-value comment consumes somebody else’s attention, even when generating it was cheap.
I would leave AI approvals disabled initially and compare the assessments with human decisions across a useful sample of pull requests. The evaluation should cover:
- High findings accepted, dismissed or downgraded by engineers;
- resolved findings that return after another push;
- duplicate comments describing the same underlying issue;
- time spent validating findings;
- merge decisions that would differ if the automated approval or gate were enabled.
The useful question is whether the changes the reviewer would block or allow match the engineering decisions the team is prepared to automate. Approval rate and comment volume do not answer that.
The minimum state model I would build
For a custom reviewer, I would start with these boundaries:
- Give each review run an immutable identifier tied to the current head SHA.
- Identify findings using the rule, path and, where needed, a symbol or code fingerprint.
- Store the current line and commit as location data rather than identity.
- Reconcile the new result with the previous run before publishing comments.
- Update one bot-owned summary with new, open and resolved findings.
- Publish inline comments only when they contain evidence or an actionable suggested change.
- Keep merge policy outside the model response.
- Retain the evidence used for any decision capable of blocking a merge.
A small state store keyed by repository and pull request is enough to start. Before spending another round tuning the prompt, I would inspect what happens after the second and third push to the same pull request. If the reviewer cannot identify what remains open, what was fixed and what is new, I would keep every finding advisory regardless of how convincing the individual comments sound.