Stop Reviewing Code. Start Reviewing Software.
AI can generate a pull request spanning dozens of files in minutes. Reviewing that same pull request can take an engineer hours. In some cases, even the engineer submitting the PR does not fully understand every change, especially when most of the implementation was generated by AI.
We have accelerated how quickly we can produce code, but not how quickly we can understand it. Still, we expect another engineer to read the resulting changes and approve them.
I think this needs to change. As AI-generated code becomes the norm, human code review should stop being the default quality gate. Most changes should rely on automated verification and human validation of the resulting behaviour. Human code review should be reserved for changes where the risk and potential impact justify closer inspection.
Code review no longer scales
Before AI, writing the implementation was usually the expensive part. A developer could spend days implementing a feature before asking another engineer to review the resulting changes.
AI changes this balance.
Generating hundreds of lines across multiple files is now relatively cheap. Understanding those same lines is not. A reviewer still needs to understand how components interact, follow execution paths, consider edge cases, and determine whether the implementation makes sense within the existing application.
The amount of code we can generate has increased considerably, while the time available for reviewing it has not.
A PR might still be reviewed and approved, but that does not necessarily mean the reviewer understood the complete change. As generated changes become larger, expecting another engineer to thoroughly inspect every line becomes increasingly unrealistic.
We never understood every line anyway
One concern with AI-generated software is that the codebase becomes a black box. If engineers stop reading the implementation, how can they remain responsible for the software they maintain?
I don’t think understanding every line of code has ever been a requirement for understanding a software system.
When joining an existing project, there will already be code written before you became involved. Even within a familiar codebase, different parts of the system are often maintained by different engineers. We also build on frameworks, libraries, databases, and other abstractions without needing to understand their complete implementation.
Instead, we rely on interfaces, architecture, expected behaviour, and the boundaries between different parts of the system. When there is a reason to investigate further, we inspect the implementation.
AI extends this way of working. The role of the software engineer moves more towards architecture, requirements, and constraints, while AI takes on more of the implementation. The important question is therefore not whether we understand every generated line, but whether we can establish that the resulting software behaves as intended.
Replace inspection with verification
Removing human code review does not mean removing quality control. It means changing how we establish confidence in a change.
AI can generate implementation code together with automated tests. These tests can run continuously alongside static analysis, type checking, integration tests, and other automated checks. If we trust AI to generate production code, I think it is reasonable to also let it generate tests for that code.
This does not mean that passing AI-generated tests proves that the implementation is correct. The same incorrect assumption can end up in both the implementation and its tests. Automated testing should therefore be one layer of the QA process, not the final authority on whether a feature works correctly.
Humans still need to identify the important business flows and critical parts of the application. We need to determine whether requirements have been met, perform functional testing, define meaningful end-to-end tests, and verify that the application remains stable after introducing a change. These checks validate the implementation against expectations that exist outside the generated code itself.
The focus of human review therefore moves from reviewing the code to reviewing the software.
Instead of spending a significant amount of time trying to understand every implementation detail, we can spend that time verifying whether the feature actually does what it is supposed to do.
Review based on risk
I don’t think human code review should disappear completely.
Some changes carry enough risk to justify inspecting the implementation before they are deployed. Schema migrations and fundamental changes to data models are obvious examples because their effects can be difficult to reverse. Security-sensitive changes, destructive operations, or changes to critical parts of a system may deserve the same scrutiny even when they can technically be reverted.
The decision to require human code review should therefore be based on the risk and potential impact of the change.
A routine implementation change that can easily be reverted and has passed automated verification should not necessarily require another engineer to inspect every line. A change that could permanently alter data, introduce a security vulnerability, or significantly affect a critical part of the application deserves additional human inspection.
Human code review then becomes a tool we apply where the risk justifies it, rather than something every pull request needs by default.
Our responsibility shifts to a higher level
AI does not remove human responsibility from software engineering. It changes where that responsibility is most useful.
Software engineers increasingly define the architecture, constraints, requirements, critical business flows, and expected behaviour of an application. AI can take on more of the implementation, while automated systems continuously verify the resulting code. Humans can then focus on validating the software and intervene when the risk of a change requires additional inspection.
Trying to maintain line-by-line human code review while the amount of generated code continues to increase does not scale. At some point, we are producing code faster than other engineers can reasonably understand it.
That does not mean we should slow down AI-assisted development so our existing processes can keep up. Our processes should move with the technology.
We should stop treating the amount of code another engineer has read as our measure of confidence. What matters is whether our engineering process can demonstrate that the software meets its requirements, behaves correctly, and can be operated safely.