Skip to content

Back to Articles

5 min read ·

Who signs the diff?

Amazon's fix for its March outages was a senior's signature on AI-assisted changes from junior and mid-level engineers. It was the right fix. But an industry that answers AI-assisted code with senior signatures while it cuts junior hiring is betting that somebody else will train its reviewers.

I write automation for a freight forwarder and customs broker in Gurugram, and I run Claude Code and Codex every working day. They earn their place. I also read every diff an agent writes before it merges.

The tools are good now, and the numbers are starting to say so. Last July, METR published a randomized trial with 16 experienced open-source developers on 246 real tasks. They believed AI made them about 20% faster. It made them 19% slower. In February, METR changed the design and reported that the slowdown has probably turned into a speedup. I believe it.

In the same post, METR says that between 30 and 50% of the developers withheld tasks they did not want to do without AI, and it now calls its own data "an unreliable signal". One of them told METR why: "I avoid issues like AI can finish things in 2 hours, but I have to spend 20 hours." That is a fair trade for someone who spent ten years learning to do the 20 hours. For someone who never did them once, it is a trap.


Amazon's retail site had several high-severity outages in a single week, and one of them reportedly blocked checkout for about six hours on March 5. An internal email named a "trend of incidents" with "high blast radius" and "Gen-AI assisted changes". What followed was a 90-day reset for 335 critical systems, plus the senior sign-off.

Amazon says only one incident involved AI directly, when an engineer acted on "inaccurate advice that an AI agent inferred from an outdated internal wiki". Take that at face value, and the response still tells you what the company believes. Its reported fix was not a better model or a bigger eval suite. It was a person who knows the system well enough to be blamed for it, and a signature is worth exactly what that person understands.

The gap between what a team merges and what it understands is review debt.

Amazon's reset is a repayment plan for review debt, paid in senior hours. Agents run it up faster than any tool before them, because the cost of producing a diff fell to almost nothing and the cost of understanding one did not move.

Google's DORA report last September linked AI adoption to higher throughput and lower delivery stability, and put it in one line: "AI doesn't fix a team; it amplifies what's already there." Amplify a team where three people understand the checkout path, and you get more changes to the checkout path and the same three people.


Last May I wrote that an AI-first memo is a hiring freeze. In August, a Stanford team put numbers on it. Working from ADP payroll data, it found about a 13% relative decline in employment for workers aged 22 to 25 in the most AI-exposed jobs, software developers among them. Experienced workers in the same jobs held steady or grew.

Two years ago I was finishing my master's, close enough to that cohort to take the numbers personally.

Nobody becomes a reviewer by reviewing. You get there by writing code that breaks and reading the stack trace at a bad hour. The agent writes that code now, and the junior who would have broken it was never hired.

The industry wants a licensed driver in every passenger seat, and it is rationing learner's permits.

On March 31, Anthropic published version 2.1.88 of Claude Code to npm with a 59.8 MB source map inside, about 512,000 lines of TypeScript that anyone could read. Bun emits source maps by default, and the package did not exclude the map files. Anthropic said no personal or sensitive data leaked. Nobody wrote the exclusion, and a line that nobody wrote never shows up in a diff. The habit of reading the file list before you publish comes from getting burned once, and the cheapest place to get burned is as somebody's junior.

In January, DHH wrote that his agents work on their own "and I just review the final outcome." Just. He can review the outcome because he spent two decades producing the inputs. He created Rails, and as late as last July he said he typed all of his code by hand.

Cutting the junior class while you demand senior signatures is not efficiency. It is freeloading. You are counting on some other company to spend years turning a graduate into the person who signs your diffs, and then you hire that person away. Either that, or you expect seniors to arrive the way the agents do, over an API.


My own rules are boring. If I cannot explain an agent's diff, I do not merge it. I want tests that pin the current behavior before I accept an agent's change, because a pinned test turns "looks right" into "still does what it did". Those two rules make agent code safe to accept. They teach nobody which behavior is worth pinning. That takes years of being wrong near someone who can tell you so.

The best reply is that agents will soon review agents, pay off their own review debt and make a senior's signature look quaint. Who knows. It did not look quaint to Amazon last month.

Every system you gate on a senior's signature will need a senior in five years. How many juniors did you hire this year?