How to understand AI-generated code
The short answer
You understand AI-generated code when you can explain what it does and why, predict what it does on an input you have not tried, and fix it without asking the agent. Reading the diff once does not get you there. Ask the agent for a plan before it writes anything, read the change in the order the data moves through it, predict each result before you run it, and ask why at every decision you would not have made yourself.
How you use the assistant matters. In an Anthropic trial, developers who asked conceptual questions, asked for explanations with the code, or questioned the code after generating it scored 65-86% on a comprehension quiz. Developers who delegated the task, relied more on the assistant as the session went on, or pasted errors back until they went away scored 24-39% (Shen & Tamkin, 2026). The groups were small, so treat the split as a pattern, not proof. To still understand the code a week later, practice recalling its topics from memory after the task.
What does it mean to understand AI-generated code?
You understand a piece of AI-generated code when you can do four things without the agent: say what the code does, say why it does it that way and not another way, predict what it returns for an input you have not tried, and find the cause when it fails. These four checks are more useful than a feeling of familiarity, because reading code that looks correct gives you the feeling without the ability.
A quick test is to close the agent session and explain the change in your own words, out loud or in the pull request description. The places where you repeat the agent's wording are the places you have not understood yet.
Why is AI-generated code hard to understand?
AI-generated code is hard to understand because the agent made the decisions and you did not. When you write code yourself, you choose the data structure, hit the errors and fix them, and each of those steps teaches you something about the code. When an agent writes the code, you see only the result.
An Anthropic trial measured the cost. 52 developers learned a Python library that was new to them; half had an AI chat assistant and half had documentation and web search. On a quiz taken minutes after the task, the AI group scored about 15 percentage points lower, and the AI group did not finish significantly faster (Shen & Tamkin, 2026). The largest gap was on debugging questions, about 21 points, and the authors attribute the control group's advantage to meeting errors and resolving them alone. That per-topic breakdown was exploratory, not pre-registered.
Your own impression of how well the work is going is not a reliable check. In a METR field trial, 16 experienced open-source maintainers estimated that AI had made them about 20% faster, and the measurement said 19% slower. That trial measured speed, not understanding, but it shows how far a developer's impression can differ from a measurement.
A classroom study of 21 novice programmers, observed with think-aloud recordings and eye tracking, describes the same gap between feeling and ability:
“Students who are already poised to succeed can leverage GenAI to accelerate, while struggling students may be hindered by using GenAI, leaving them with an illusion of competence.”
The classroom study observed students and did not measure a learning outcome, so it describes the gap without measuring its size.
How do you read AI-generated code you did not write?
Read AI-generated code actively: make a prediction or ask a question at each step, instead of scrolling through the diff. These seven habits work with any coding agent:
- Ask for the plan before the code. In Claude Code, plan mode does this; the documentation says "Claude reads files and proposes a plan but makes no edits until you approve." Press Shift+Tab until the status bar shows plan mode on. With another agent, ask for a plan in your first message and tell it not to edit files yet. Compare the plan with what you would have done, and ask about every difference.
- Read the change in the order the data moves, not file by file. Start where the input enters, such as a route, a handler or a command, and follow it to where the result is stored or returned. Claude Code's documentation suggests prompts of this kind, for example "trace the login process from front-end to database".
- Predict before you run. Before each test or manual check, write down what you expect to happen. A wrong prediction shows you exactly which part you did not understand.
- Ask why at each decision you would not have made: why this data structure, why this library call, why this error is caught here and not higher up. If the agent answers that the choice is common practice, ask what breaks if you remove it.
- Change one thing by hand. Remove a check or change a boundary value, and predict the result first. If the tests still pass after you remove a check, find out what the check was for, or whether a test is missing.
- Debug one failure without the agent. Debugging questions showed the largest gap in the Anthropic trial, so fixing one error yourself practices the skill where the AI group trailed most.
- Explain the change in your own words before you merge it. If you cannot write two sentences about why the change is correct, you have not finished reviewing it.
These habits cost time. Use all seven on code you will maintain or debug in production, and fewer on a script you will delete next week.
Which ways of using an AI assistant keep your understanding?
Asking conceptual questions, asking for explanations with the code, and questioning the code after the assistant generates it are the ways of working that scored highest. In the Anthropic trial, the researchers sorted screen recordings of the AI group into six usage patterns and compared quiz scores (Shen & Tamkin, 2026). The group without AI averaged about 65%.
| Usage pattern | Quiz score | Developers |
|---|---|---|
| Generate the code, then question it | 86% | 2 |
| Ask for code with an explanation | 68% | 3 |
| Ask conceptual questions, then write the code yourself | 65% | 7 |
| Delegate the whole task | 39% | 4 |
| Rely on the assistant more as the session goes on | 35% | 4 |
| Paste errors back until they go away | 24% | 4 |
Each usage pattern has two to seven developers, the researchers sorted them after the fact, and nobody was assigned a pattern, so treat the split as a pattern worth copying, not proof. The trial used a chat assistant, not an agent that edits files. The authors read their result as a lower bound for how much developers offload to agentic tools, because those tools need less input from the developer.
Watch for the lowest-scoring pattern when you work with an agent. Pasting an error back until it goes away fixes the build and leaves you without an explanation of why it broke.
How do you keep understanding AI-generated code a week later?
Practice recalling what the code taught you. Understanding a change today does not mean you remember it next week, and reading habits during the task do not bring the material back later.
Retrieval practice means answering questions from memory, then checking your answers. In the reference experiment, students who read a passage once and practiced recalling it three times remembered 61% of it a week later; students who read it four times remembered 40% (Roediger & Karpicke, 2006). A meta-analysis of 222 classroom studies, covering 48,478 students, found that quizzing raised achievement by about half a standard deviation (Yang et al., 2021).
Those studies tested students, not engineers. The evidence on working professionals is thin: about six studies, a pooled result whose confidence interval crosses zero, and no randomized trial on programming.
For AI-generated code, the practical version is short. After the task, with the code closed, answer a few questions on the topics the agent worked on, and check your answers.
Where Atomic Reps fits
Atomic Reps adds the recall step to the tools you already use. When your coding agent finishes a task in Claude Code, GitHub Copilot, Cursor or Codex, it asks you questions on the topics that task worked on, and you answer each one with a letter, from memory.
It does not replace the reading habits above. The reading habits help you understand a change while it is open; Atomic Reps helps you recall its topics after it is closed.
See how it works in your coding agentSources
- Anthropic, Claude Code documentation, Common workflows
- Shen, J. H. & Tamkin, A., arXiv (Anthropic), 2601.20245 (2026)
- METR, developer productivity RCT (2025)
- Prather, J., Reeves, B., Leinonen, J. et al., ICER '24 (2024)
- Roediger, H. L. & Karpicke, J. D., Psychological Science 17(3) (2006)
- Yang, C., Luo, L., Vadillo, M. A., Yu, R. & Shanks, D. R., Psychological Bulletin 147(4) (2021)
