Skip to content

Back to Blog

5 min read ·

The only code generator I trust

The best code generator ever built has no chat window. It is the compiler, and it earned my trust the slow way, by producing the same bytes from the same source, by answering to a published standard, and by surviving more testing than anything I will ever ship. Any tool that wants to write code for me has to clear that bar, and nothing with a chat window clears it yet.

Two weeks ago, on Alphabet's earnings call, Sundar Pichai said that "more than a quarter of all new code at Google is generated by AI, then reviewed and accepted by engineers." The headlines kept the first half of that sentence. I keep rereading the second half.

Since June I have been writing C++ controllers for Zoom Rooms, against an SDK that picked the language for me. By Pichai's accounting, every instruction the CPU has run for that code was generated by a machine. All of it. Nobody reviewed it and nobody accepted it, because nobody has to. You do not open a pull request for the assembly. No CEO has ever gone on an earnings call to announce that the compiler wrote all of Google's machine code.

That missing review step is the entire argument.


Compilers have been writing code since the 1950s (IBM shipped FORTRAN in 1957), and they earned their place with properties too boring for a slide. Compile the same source twice with the same compiler and flags, and you get the same bytes, as long as nothing like __TIME__ or the build directory leaks into the output. Linux distributions spent years hunting those leaks, so that anyone can rebuild a package and check it byte for byte against the download.

The compiler answers to a spec. The C++ standard defines what a + b means for two ints, so the compiler is not guessing what I meant. The standard even says in writing where that definition stops. If the sum overflows, the behavior is undefined, and the compiler owes me no warning. Break the grammar or the type rules, though, and it refuses, with a file name and a line number.

Every program anyone compiles is one more test case. Blaming the compiler is the oldest excuse in programming. It is almost never guilty.

Nothing inside it is magic. The lexer cuts characters into tokens, and the parser checks the tokens against the grammar and builds a tree. Then semantic analysis walks the tree with the rude questions. Was sum ever declared? Why does a function that takes two arguments get three? The optimizer rewrites what survives into something cheaper that means the same thing, and the code generator picks real instructions and registers for a real CPU. Most of the error messages I have cursed at came from one of those stages. The rest came from the linker. Around all of it sit Make (or CMake writing the Makefiles for you), which rebuilds only what changed, and clang-tidy, which reads your code without running it.

Here is the optimizer at work, with GCC for x86-64 at -O2 -masm=intel:

int main() { int a = 5; int b = 10; int sum = a + b; return sum; }

main:
        mov     eax, 15
        ret

The variables are gone. So is the addition. The compiler saw that a and b never change, worked out 5 + 10 at compile time, dropped the variables that nothing reads, and left one instruction that puts 15 in the return register. I have seen this called inlining, and I have called it that myself. It is not. There is no function call here to inline. It is constant propagation and constant folding. The standard permits it under the as-if rule, which lets the compiler rewrite anything as long as the program's observable behavior stays the same.


I will give the new generators this much. More than a quarter of the new code at a company as careful as Google does not get accepted by accident. On the same day as Pichai's call, GitHub announced that Copilot would offer Claude 3.5 Sonnet and o1-preview next to GPT-4o.

But ask the people who use them. In this year's Stack Overflow survey, 76% of developers said they use or plan to use AI tools, and only 43% said they trust their accuracy. Picture a compiler that 43% of its users trusted. You would not call it a compiler. You would call it a first draft. You would review it before you accepted it, which by Pichai's own account is what Google's engineers do.

The bolder pitch calls the prompt the next programming language and the model its compiler, and says that reading the output will soon feel as quaint as reading assembly. That pitch is compiler cosplay. A compiler maps one specified language onto another and refuses when it cannot. A model turns an unspecified wish into plausible text. When it has to guess, it usually does not say so. Ask it for the same function twice and you can get two different functions, like asking the same uncle twice for directions to the wedding. There is no standard for a prompt, and no as-if rule.

Compiler cosplay is not harmless marketing. It tells people like me, six months out of a master's program, that it is fine to stop learning the layers under the prompt. It has that exactly backwards. When generated code is wrong, the parts of the pipeline that catch it the same way every time are the deterministic ones: the compiler, its warnings, the analyzer, the tests. -Wall -Wextra is still the cheapest reviewer anyone will ever hire, and it helps only a programmer who understands what it is complaining about.

Until a model makes the promises GCC makes, the second half of Pichai's sentence is the honest half.

Learn what your compiler does with your code, or stay a tourist in your own program.