Skip to content
AtomicReps
Guide

AI code review checklist: approve only the code you can explain

The short answer

Before you approve a pull request that a coding agent wrote, close the diff and answer eight things from memory: what the change does, why each file changed, what one real input returns, what happens when a call fails, why each new dependency is there, whether the code repeats something the codebase already has, which bug each test would catch, and what you did not check. If you cannot answer one of the eight, reopen the code at that point before you approve.

The checklist checks the reviewer's understanding. Linters, tests and AI reviewers check the code, and none of them checks whether the person who approves the change could debug it later. In a randomized trial, developers who learned a new library with an AI assistant scored about 15 percentage points lower on a comprehension quiz than developers who used documentation and web search (Shen & Tamkin, 2026).

What should you check before approving AI-generated code?

Check that you can explain the change without the diff open. Write one sentence on how you would have solved the task before you read the agent's code, then work through this list with the code closed:

  • Say in two sentences what the change does and why it was needed.
  • Name each changed file and the reason that file had to change.
  • Pick one real input and predict what the new code returns for it. Then pick one edge case, such as an empty list or a missing field, and predict that too.
  • Trace one failure: what happens when a network call, database write or parse in the change fails.
  • For each new dependency, say what it does and why the existing code could not do the same job.
  • Say whether the change repeats logic that already exists somewhere in the codebase.
  • For each new test, name a bug that would make the test fail.
  • List what you did not check, and say so in the review.

Then open the code and compare. Every answer you got wrong or could not give is the part of the change to read again, or to ask the author about, before you approve.

What does each item on an AI code review checklist catch?

Each item catches a failure that a passing test run does not show. This table lists the failure behind each check.

Checks on an AI-generated pull request and the failure each one catches
CheckWhat it catches
Your own plan firstAccepting the agent's approach only because it was the first one you saw
What and whyA change that solves a different problem from the one in the ticket
Each changed fileEdits outside the task, such as a config change or a renamed export
One input, one edge caseCode that works on the example and fails on empty or missing data
One failure pathErrors that are caught and ignored, or retried without a limit
New dependenciesA package added for something the codebase or the language already does
Repeated logicA second copy of a helper that already exists, which then drifts from the first
A bug per testTests that pass whatever the code does
What you did not checkAn approval that reads as a full review when it was a partial one

The repeated-logic check matters more with AI-written code. GitClear's analysis of 211 million changed lines found that commits containing a duplicated block of five or more lines rose from 0.70% in 2020 to 6.66% in 2024 (GitClear, 2025). GitClear sells code-analysis software, and the data shows a trend over the years of AI adoption, not proof that AI caused the rise.

Why check your own understanding when the tests pass?

Passing tests show that the code does what the tests check; they do not show whether the reviewer understands the code well enough to debug it later. Two studies suggest that AI help can lower understanding, and that developers misjudge how AI changes their own work.

In an Anthropic trial of 52 developers learning a new Python library, the developers who had an AI chat assistant scored about 15 percentage points lower on a quiz taken minutes after the task than developers who used documentation and web search. The largest gap was on debugging questions, though that breakdown was not part of the pre-registered analysis (Shen & Tamkin, 2026).

In METR's field trial, 16 experienced open-source developers worked on real issues in their own repositories. Afterwards they estimated that AI had made them about 20% faster, while the measurement showed them slower (METR, 2025). A reviewer's sense of having understood a change is not a reliable measure of the understanding, which is why the checklist asks you to answer from memory and then compare.

Can an AI code reviewer check your understanding for you?

An AI reviewer cannot check what you understand. GitHub Copilot code review reviews the pull request, identifies issues and suggests fixes, and none of those comments shows whether the human reviewer understood the change. GitHub's documentation describes the limits of the Copilot review:

“Copilot is not guaranteed to spot all problems or issues in a pull request. Sometimes it will make mistakes. Always validate Copilot's feedback carefully. Supplement Copilot's feedback with a human review.”

GitHub, GitHub Docs, About GitHub Copilot code review

Use an AI reviewer to find issues you would miss, and use the checklist to confirm that you understand both the change and the reviewer's comments on it.

How do you use the coding agent to understand its own code?

Ask the agent to explain the change, then check the explanation against the code instead of accepting it. Ask why it chose this approach over the one you wrote down, what the code does on your edge case, and which lines handle the failure you traced.

The Anthropic trial supports this way of working. Inside the AI group, the developers who asked conceptual questions, asked for explanations with the code, or questioned the code after generating it scored 65-86%. The developers who delegated the task, relied more on the assistant as the session went on, or pasted errors back until they went away scored 24-39% (Shen & Tamkin, 2026). Each of those groups had two to seven people, and the study did not assign anyone to a usage pattern, so treat the split as a pattern, not proof.

Writing your own plan first follows the same logic. In an eye-tracking study of 21 students in a first programming course, the students without difficulties read AI suggestions, kept the ones that matched a plan they already had, and ignored the rest (Prather et al., 2024). The students in that study were novices, not professional reviewers.

How do you still understand the code a week after the review?

Practice recalling the topics of the change after a delay. The checklist tests your understanding on the day of the review; retrieval practice, which means answering questions from memory instead of rereading, is a method with strong evidence for recall a week later.

In the reference experiment, students who read a prose passage once and then practiced recalling it three times remembered 61% of it a week later, and students who read it four times remembered 40% (Roediger & Karpicke, 2006). The evidence on working professionals is thin: there are about six studies, their pooled result has a confidence interval that crosses zero, and there is no randomized trial on programming.

Where Atomic Reps fits

Atomic Reps adds the recall step after the review. When your coding agent finishes a task in Claude Code, GitHub Copilot, Cursor or Codex, it asks you questions on the topics that task worked on, and you answer each one with a letter, from memory.

It does not replace the checklist. The checklist confirms you understand the change before you approve it; Atomic Reps helps you keep the topics behind it.

See how it works in your coding agent

Sources