Code review was built for a slower world
The hard part is the same as it has always been. Gathering requirements and modeling an application that the users expect. But from the moment requirements are even roughly available, Claude Code and a good harness accelerate from code to QAT, and when gaps are identified in UAT, updating the requirements, regenerating code, and returning to UAT can be minutes or hours, not days.
That is the whole argument. The rest of this paper is about why the industry is looking at the wrong stage of that loop, what one project's record shows when the loop runs at that speed, and what still needs a human to read it.
1. The queue moved
Everyone can see the bottleneck. Writing code got cheap and reviewing it did not, so the pull request queue is where the work now waits. The commentary is thick with numbers: pull request volume up 98% with review time up 91%, AI-generated PRs waiting 4.6 times longer to be picked up, 61% of agent-authored PRs receiving no review at all. I hold those figures loosely. Most are telemetry from companies selling review tooling, cited secondhand. But the shape underneath them is right, and Addy Osmani has stated it as plainly as anyone: understanding code costs what it always did.
2. A faster reviewer is the problem restated
The prescription in nearly every one of those pieces is an AI reviewer on the pull request. That is another reader meeting the code cold. It has no record of why the change was made, what it was supposed to do, or what was already checked. It reconstructs the claim from the evidence, which is the expensive part, and the diff carries only the evidence.
The two failure modes everyone warns about, rubber-stamping and over-scrutinizing every line, are the same failure. They are what a reviewer does when handed evidence without a claim. Speeding the reviewer up does not change what it was handed.
3. Nobody reads the assembler, and nobody decided that
When C arrived, plenty of teams kept walking the generated assembly in code review. Nobody announced the day that stopped. It stopped because three things happened. The compiler accrued trust through use, every correct program a vote. The audit stayed cheap and available, so you could look whenever you suspected the mechanism and therefore didn't need to look every time. And the line settled in different places for different people. Kernel, embedded, and compiler writers never stopped reading assembly, because their domain is the mechanism. Application programmers stopped almost entirely.
The useful frame is that the line sits wherever the audit stops paying, and it moves with tooling, not with declarations. A team that walked assembly in 1978 kept doing it while the walk kept catching things. When it went a year without changing an outcome, they quietly stopped.
An agent is not a compiler. A compiler is deterministic and an agent is not, so trust cannot accrue in the translator the same way. But the C teams did not stop reading assembly because the compiler had been proven. They stopped because they could not keep up with it and nothing broke when they didn't. That is the next section's point, in 1978 clothing.
4. The author now outruns every human-paced instrument
The discourse reads the bottleneck as a capacity problem: not enough reviewers for the volume. It is a cadence problem. When the author produces faster than a human can read, reading is not a weak instrument. It is an instrument that no longer gets a turn.
Here is what that looks like in one repository, with one developer. In July 2026 the Sharpee project landed 391 commits, 86 architecture decision records, and sixteen npm releases. A new authoring language, Chord, went from an accepted design to a working compiler running the standard first test story with an 81-of-81 gate in about ten hours. One decision record's six children removed a deprecated call from 382 sites in 21 hours. A nine-phase subsystem was built in seven and a half hours.
On a single day in December 2025, all 43 standard-library actions were refactored three separate times: to one pattern at four in the morning, to a second by five, and to a third by late afternoon, with a written assessment of every action in between. The next morning, running the test story found that the put action had never moved anything. It called the container's validation method and returned. The fix was one line.
It is tempting to read that as four review passes that failed. It is not. Three refactors in thirteen hours is three passes in a window where no human-paced process could have run at all. At hand-written pace those refactors are weeks apart, and someone plays the story between them in the ordinary course of work. The bug survived because the slow loop never got a turn.
That is the honest statement of what this record shows about reading. Human reading was not measured and found wanting. It was sampled at near zero. Every review pass in the record is a model reading model-written code. What the record can say is that at this cadence, the only instruments that fired were the ones that run at machine speed, builds, tests, and audit scripts, and the ones a consumer forced: a story, a live player, a reader following a manual.
5. The before times: log it, prioritize it for later, build around it
I worked at a company in 2010 that was building a customer-facing platform in Silverlight. We had an arbitrary deadline that was clear to everyone as being unachievable. Even so, management argued the deadline was a hard one. We explained there were whole sections of the application that would either not be available or likely fail. We estimated over 150 hard defects that required at least two months of effort. They didn't care. They told us to release the application as-is.
The app was released, we worked on defects, and the customer loudly complained. I left the company before any of it was resolved, but I suspect I was the "fall guy".
Nothing about that was ignorance. The debt was known, counted, and sized before release. The failure was that repair cost real calendar time, and calendar time lost the argument to the date. That is the mechanism of the before times: known debt is overruled by the schedule, externalized to the customer, and then personalized onto whoever sized it. Review had to be strict in that world because whatever got through was permanent. Repair was slow, it competed with new work, and it lost. So teams logged the debt, prioritized it for later, and later never came. They built around it instead, and the workaround became load-bearing, and the next workaround was built on the first.
Two months of effort is the number that collapses. The Sharpee record is unusually clear on what repair looks like now. Roughly a third of everything ever written in the repository is gone: 1,733 of 4,826 source files ever created are absent from the current tree, and somewhere between 27% and 33% of all lines ever written went into files that no longer exist. A text-rendering package was deleted in one commit, 62 files and 4,208 lines, under a decision designed and executed in about a day that explicitly reversed a decision recorded a week earlier. Writing an architecture document surfaced eight issues in one sitting, and the cleanup deleted two entire layers and drove 1,035 unsafe casts to zero. A transcript grammar was cut out of the authoring tool, 168 files and 12,369 lines, and shipped as a major release the same day.
Cheap repair did not mean churn everywhere. The domain core barely moved: 228 of the 290 live standard-library source files and 125 of the 190 world-model files predate 2026, and one core trait file has been touched five times in fourteen months. What got built and deleted was the layer you cannot know without building, and what stayed was the layer you can.
The consequence for the gate is the point. In a world where the fix lands the same day the consumer finds the defect, a defect caught after merge is not a process failure. It is the process. The thing the 2010 team could never afford, letting the customer find it and fixing it immediately, becomes the normal path, and the "later" queue disappears because the fix is faster than the ticket. Remove the two months and the meeting where management overruled us never happens, because there is nothing to overrule.
6. Method/1 had the right artifacts and the wrong economics
The methodologies of the eighties and nineties had the artifact chain right. A functional specification said what the system did. A detailed design specification said how. Code was transcription of the design. The industry abandoned that because transcription was the expensive step and because the specification drifted the moment the code diverged from it; a specification was never executed, so nothing told you when it had gone stale. Agile's answer was to make the code the specification. That was a concession to cost, not a discovery about design.
Generative tooling inverts the cost. Transcription is now the cheap step. The specification can stay live because the code is re-derived from it on demand. The one thing Method/1 could never do, keep the specification and the code in lockstep, is now the easy part. The loop was linear because iteration was expensive. Agile iterated at sprint cadence, in weeks. This loop iterates in hours with Method/1's artifacts intact.
What the human still does is worth listing, because every item is a specification act. Design the plan. State the invariants. State what each unit does, under what conditions, why it matters, and what it rejects. Decide what constrains future work. Judge the acceptance demonstration. In DevArch that middle item is a Behavior Statement, written in conversation before any test exists, and each of its lines becomes a test the suite is then graded against. It is a detailed design specification that executes by derivation. It is necessary. The next sections show it is not sufficient.
6a. The hard part gets better tooling at the edges
Requirements and modeling are still the work. But two edges of the loop that used to be meetings are not meetings any more. Where I work now, product owners, the development manager, and developers trade requirements as markdown on the same git remote the code lives on, and questions and answers between them travel through a knowledge-worker harness over that same remote. A gap found in acceptance testing becomes a change to a markdown file. A developer's question becomes a thread that reaches the person who can answer it. The regeneration starts from the updated text. This is an account of what changed in the room, not a measurement, and I can't offer figures for it.
What it is, structurally, is Method/1's artifact chain with the handoffs removed and the decisions kept. The product owner still owns the requirement, the manager still owns priority, the developer still owns the build. None of them wait on a meeting to move, and none of them transcribe each other. Requirements, questions, answers, and code share one remote and one history, so "what did the requirement say when this was built" is a log query rather than archaeology. The 2010 team had 150 defects in one system and the requirements they violated in another, owned by different people. That gap is where the two months lived.
7. What the record shows, at the right altitude
Across three and a half years and roughly 1,250 session summaries (since extended by 244 more, through October 8), the Sharpee retrospective verified a set of defects against git. I classified each by the channel that found it: a consumer executing the real path, a mechanical check, or a human or model reading. Of eighteen behavioral defects, the kind where code ran and did the wrong thing, thirteen were found by a consumer, two by a mechanical check, and two by reading. Of five structural defects, none were found by a consumer. Of six false completion claims, none were found by a consumer.
Reading in that tally is almost entirely model reading, for the reason given in section 4. The right altitude for the result is not "consumers beat reading." It is that consumers were the only instrument that kept up, and what they found was fixable the same day.
Four cases carry the argument.
The put action, told above.
A save-and-restore service that, from July 2025 to April 2026, silently dropped the score ledger, every capability, every state value, and every relationship on every save. The original code carried the comment "For now, this restores entity traits and locations." Three tests covering it were green the whole time, because they asserted that the save hook was called and the completion event was emitted, never on what was saved. In April a 3,164-test audit read that file, wrote the gap down in its own words, "no test for save operation providing actual save data back," graded the file excellent, and said keep. A static grader reported 177 green and zero yellow. Three consumers saved and restored daily for nine months. It was found by reading, while scoping unrelated work. It was fixed in one session with a 314-line round-trip test. In 2010 that is a priority-two defect that ships in a year.
The grader itself. The project built a script to implement its own red-yellow-green test-quality rule. Its yellow check applied to one directory, and everything else fell through to a line that literally read "GREEN (default)." The project never found this. The retrospective's author read the script four months later. The instrument that measured test quality was itself unread.
The live player. A phase named "Engine Subprocess and Turn Execution" was recorded done in the plan tracker while the engine entry point was an 84-line echo stub and every sandbox test drove a 92-line fixture of that stub. The carve-out was disclosed in the commit message and the session summary's own status read incomplete, so it was known and not seen rather than hidden. It surfaced when a live player opened the site and saw "Waiting for the story to begin." The rule written that day, that the system under test cannot be the thing you wrote to stand in for the system under test, is now a line in DevArch's ruleset.
The column in that table I first labeled "what had already looked and missed it" reads differently at this altitude. In a world where the fix is same-day, those instruments were not failing at their job. Their job had changed.
8. The unit of verification is the demonstration
Pick the instrument that scales with the author and whose findings are fixable at the speed they are found. That is the demonstration, and in the thesis's terms it is UAT. A gap found there goes back to requirements, code is regenerated, and the demonstration runs again in hours. The before-times loop had review in it because the loop was too slow to afford a second pass. This loop affords as many passes as it takes.
Review does not disappear. It moves up a level, to the plan, the behavior statement, the decision record, and the demonstration, which are the things that still compound. For a non-deterministic translator, trust cannot accrue in the translator. It accrues in the ledger: recorded demonstrations and recorded instruments, per session, with the session id. A compiler's trust was folklore. This is a record.
There is one place reading is irreplaceable, and it is not the code. It is the instruments and the decisions. The grader that defaulted to green, and a test harness that silently skipped unasserted commands and was fixed three times in six months, say that instruments are where a human read pays. Decisions compound because cheap repair of code does not make a wrong boundary cheap to move.
What this does to the pull request is simple to state. The PR is a projection of verification that already happened, not the place it happens first. The artifacts exist: the plan, the behavior statement, the integration reality statement, the graded suite, the session summary. Today none of them reach the forge; the diff goes bare. The 61% of agent PRs with no review is only alarming if the diff is the unit of review. In this model it never was.
9. Where the line is
I don't know where the line is yet, and this paper is not going to pretend to. What it can say is what sets the line and how you would know it moved.
Cadence sets it. Whatever cannot run on every change becomes an audit, so the line moves with tooling, with automated consumers, seed-pinned runs, and a manual followed by a reader in a container, and not with trust in the author. The measure is audit yield: a human reading code and changing the outcome is a hit; the apparatus catching it first makes the read redundant. Yield goes to zero for pure transformation first, side-effect code later, integrations last. Security and concurrency stay audit-heavy for the same reason kernel developers still read assembly.
It cannot be computed backward. Across 2,029 files in the Sharpee corpus, about 1,500 of them session summaries, roughly 2% mention in prose how a defect was found, and no structured record carries it. The denominator exists mechanically from mid-2026, when the harness began logging test runs. The numerator was never written down. It can be computed forward with one field at the point every item is already filed. That is a method, not a product.
10. What this record cannot support
One project, one developer, one record. The record is a confession, which is evidence of honesty and also of a single vantage.
No counterfactual. There is no record of a review that examined the put action and passed it, and no record of what a human-paced team would have caught. The claim about human reading in this paper is economic, not empirical.
Cheap repair has a residue. The same throughline that gives the one-third number records what was built and neither used nor removed: a helper module repudiated in writing the day it was written and still in the tree eight months later with only its own test importing it; 4,601 working lines on a branch never merged; a directory explicitly named a parts bin that shipped 146 files for three months before it was archived; a server built over nine days and 35 commits, deployed, and deleted two minutes after its end-to-end tests landed, for a successor that also stalled. Cheap repair did not make reconciling the record cheap. Status lines stopped being load-bearing.
Deletion was never the tool's default. After an unasked deletion in September 2025, the rule "never delete files without confirmation, not even to get a build working" went into the project's instructions, and every large deletion in the record was authorized case by case. Cheap repair is a capability the human directs.
Instruments fail silently, and the only thing that caught them was a human reading them. This paper argues against reading code as the unit of work and for reading instruments and decisions as the audit. Those are different claims, and the paper stands or falls on keeping them apart.
And causation is confounded. The project had no methodology at all until January 2026. The comparison between an eighteen-day detection gap in 2025 and an 87-minute one in 2026 spans four model generations and one developer's accumulated experience, not instrumentation alone.
Sources: Osmani, "Agentic Code Review"; Google Cloud Office of the CTO, "When AI writes the code, who reviews it"; Codacy, "AI Is Breaking Code Review"; the Sharpee history retrospective (throughlines "The Forcing Functions," "The gap between done and true," "Build, Ship, Delete"; verification record, 41 claims, 20 confirmed, 19 corrected, 2 refuted); defects-by-channel.md.