Review paper · B.Tech CSE, semester 5
phases of compiler design
The classical six-phase model, measured against a compiler that reports its own phase structure as data.
Harshit Khemani · Kush Ahuja · Mohit Kumar Mishra · Kushagra Agrawalzero.khe.money
Why this paper
does the textbook still describe a real compiler?
Every syllabus teaches six phases in a straight line, front to back. We took one modern compiler, ran it, and asked it to describe itself.
6
phases the textbook names
2
cross-cutting activities beside them
8
phases Zero 0.3.4 actually reports
The two lists don’t line up — and where they diverge is the finding, not an inconvenience.
Zero’s own phase list · zero check --json → compilerPhases
Background
the model we were taught
Six phases in sequence, with a symbol table and an error handler running alongside all of them.
phase 1
Lexical analysis
phase 2
Syntax analysis
phase 3
Semantic analysis
phase 4
Intermediate code generation
phase 5
Code optimization
phase 6
Target code generation
Cross-cutting
Symbol table management — a structure the front end rebuilds every compile.
Cross-cutting
Error detection and reporting — diagnostics rendered as prose, for a person.
Aho, Lam, Sethi & Ullman · Cooper & Torczon
The subject
zero compiles a graph, not a file
Text enters once, through an ingestion gate, and becomes a persistent typed graph. Every compile after that reads stored facts instead of re-deriving them.
108
declarations
109
symbols
176
types
974
source-map rows
That’s the symbol table — not rebuilt per compile, but stored, versioned and queryable. It’s the largest program in our corpus, and the graph is the artifact.
zero query --json --full on p08_lexer
Methodology
we ran the compiler and wrote down what it said
One harness drives every command, on every program, for every target — and writes one JSON file.
8
corpus programs, 6 to 304 lines
10
deliberately broken programs
8
advertised build targets
64
builds attempted
Per program: tokens · parse · check · query · source-map · time · size · mem · test · run, plus one build per target. Prose, tables and charts all read that same file, so the text and the evidence can’t drift apart.
tools/capture.mjs → web/data/capture.json13th Gen Intel(R) Core(TM) i5-13450HX, 16 cores · host win32-x64.exe
The corpus
eight programs, smallest to largest
| Program | Lines | Tokens | Graph nodes | Relative size | Builds |
|---|---|---|---|---|---|
| p01_hello | 6 | 35 | 13 | 8/8 | |
| p02_arith | 31 | 181 | 125 | 8/8 | |
| p03_control | 54 | 237 | 179 | 8/8 | |
| p04_shapes | 85 | 371 | 231 | 5/8 | |
| p05_errors | 54 | 228 | 150 | 3/8 | |
| p06_memory | 101 | 665 | 413 | 5/8 | |
| p07_generics | 45 | 223 | 131 | 5/8 | |
| p08_lexer | 304 | 1,471 | 974 | 5/8 |
zero tokens --json · zero query --json · 680 lines, 2,216 nodes, 2,201 edges in total
Finding 1
the phases didn’t disappear. they moved.
Lexical, syntactic and semantic analysis run once, at an ingestion gate that admits a program into the graph. The steady-state compile path is lowering, codegen and link — nothing else.
Runs once · at the gate
Characters are scanned here and never again. Package compilation reads a binary graph store, so the scanner is absent from the steady-state path.
Runs every compile
Note that resolve is reported before parse. That isn’t our typo — it’s what a compiler does when the names are already bound.
zero check --json → compilerPhases, in the order reported
Finding 2
lowering is the whole clock
Every reported phase except lowering returns 0 ms. The front end isn’t fast — it isn’t running.
100%
of measurable phase time is lowering
53 ms of 53 ms, summed over all 8 programs.
zero time --json, cold cache · summed across the corpus · 0 ms is the compiler’s own reported figure, not a rounding artifact of ours
Finding 3
the two ends accept different languages
17 of 64 builds failed with BLD004 — after zero check had already passed them.
One dot per build. 8 programs × 8 targets.
47 / 64
builds produced a binary
17
refused after passing the check
| Why the backend refused | Builds |
|---|---|
| BLD004 · record | 9 |
| BLD004 · IR_VALUE_CHECK | 4 |
| BLD004 · unsupported instruction | 3 |
| BLD004 · IR_VALUE_RESCUE | 1 |
zero build --target <t> after a clean zero check · match could not be lowered on any target on this build
Finding 3, continued
same graph, different emitter
Backend completeness belongs to the target, not the program. Four targets built everything; three built fewer than half.
One program, two fates
p05_errors cross-compiles to an 16,632-byte binary for Linux — and is refused outright on the Windows host it was compiled on.
zero build --target <t> --emit exe · zero size --json
Finding 4
one command, two verdicts that disagree
zero check --json publishes a top-level verdict and a nested one. On some programs they contradict each other, and the schema presents the wrong one first.
What the top level says
"ok": true, "diagnostics": []
A consumer reading the field the schema presents as the verdict stops here.
What the same payload also says
"targetReadiness": { "buildable": false, "stage": "lower" }
…carrying a BLD004 that names the exact construct it can’t lower.
2 of our 10 error cases hit this. The compiler knew. It just answered the question twice.
zero check --json on e09_lowering · schemaVersion 1
Finding 5
what it costs to be readable by a program
The same question — is this program correct? — answered for a person and for a machine, on a six-line program.
141 bytes
the whole program
4 bytes
zero check — the prose answer
17,896 bytes
zero check --json — the same answer
That’s 4,474× more bytes to say the same thing — and on our largest program the JSON reaches 189,741 bytes. The trade isn’t tokens for tokens. It’s tokens for actionability: a span, a typed repair, and a field a program can branch on without parsing English.
zero check vs zero check --json on p01_hello · bytes measured, not estimated
The error corpus
ten broken programs, three gates
| Case | Intended phase | Refused at | Code |
|---|---|---|---|
| e01_lexical | lexical | import | PAR100 |
| e02_syntax | syntax | import | PAR100 |
| e03_name | resolve | import | NAM003 |
| e04_type | check | import | TYP002 |
| e05_mutability | check | import | TYP009 |
| e06_effect | check | import | ERR003 |
| e07_memory | check | import | MEM003 |
| e08_target | target | check | TAR002 |
| e09_lowering | lower | build | BLD004 |
| e10_match | — | build | BLD004 |
7
stopped at the gate
1
stopped at check
2
got as far as build
0 of 10 left anything in the store. The gate holds.
zero import → zero check → zero build, in that order
The analytical core
six phases, mapped onto what zero reports
| Classical phase | Zero reports it as | Where it runs | How to see it |
|---|---|---|---|
| 1. Lexical analysis | parse | ingestion | zero tokens --json |
| 2. Syntax analysis | parse | ingestion | zero parse --json |
| 3. Semantic analysis | resolveinterfacecheck | both | zero check --json |
| 4. Intermediate code generation | lower | compile-path | zero size --json (loweredIrBytes) |
| 5. Code optimization | lowercodegen | compile-path | zero build --profile release-small | tiny |
| 6. Target code generation | codegenobjectlink | compile-path | zero build --emit exe && zero size --json |
Optimization is the one phase Zero never reports separately — it’s selected by build profile, so its cost is folded into lower and codegen.
The site
phases 1 and 2 run in your browser
Zero 0.3.4 has no WebAssembly target — its own report marks browserCompiler as removed — so the real compiler can’t run client-side. We reimplemented the scanner in TypeScript instead, then checked it against the compiler token by token.
2,666
tokens compared, field by field
8/8
programs matching exactly
8/8
function sets recovered
Kind, text, line, column, offset and length — all six fields, on every token. So phases 1–2 run live on whatever you type. Phases 3–6 are replayed from the capture, and the terminal says so out loud the moment you edit the buffer.
tools/lexer-fidelity.mjs vs zero tokens --json · agreement on this corpus, not a proof of equivalence
Why it matters
the diagnostics are written for a program
The textbook’s second cross-cutting activity is error reporting, and it assumes a human reader. Zero’s assumes a caller that has to decide what to do next.
Classical
A sentence. To act on it, you parse English and hope the wording is stable across releases.
Zero
Every field is branchable. That’s what the extra 17,892 bytes buy.
The six phases were designed when the compiler’s reader was a person at a terminal. Zero is designed for a compiler whose reader is another program.
zero explain --json · zero fix --plan --json
Threats to validity
what this evidence can’t tell you
Everything here describes one build of one compiler on one machine. We’d rather say that than let a reader over-read the numbers.
One version
Zero 0.3.4, build 5b3a90a, captured 2026-08-21. It’s a preview-stage compiler; the BLD004 gaps are unfinished work, not a design position.
One machine
13th Gen Intel(R) Core(TM) i5-13450HX, 16 cores. Timings are medians of repeated runs, but they’re still this host’s timings, and cross-compilation reports itself as not ready.
A small corpus
8 programs, 680 lines. Enough to show the shape, far too little to characterise a language.
Millisecond resolution
0 ms means “under the reporting floor”, not “free”. The claim is that the work moved, not that it vanished.
Every raw figure is in web/data/capture.json — check ours against yours
Conclusion
a vocabulary, not an order
The six phases still name the right work. They no longer describe the sequence a compiler runs them in.
What survives
Lexing, parsing, typing, lowering and emission are all still there, still separable, still worth naming separately.
What changed
They’re split across an ingestion gate and a compile path. The symbol table became a database. Diagnostics became records.
What that costs
Two ends that accept different languages, 17 refused builds, and a verdict field that can be wrong.
Teach the six phases as a decomposition of the work. Then show a compiler that reports its own — because that’s where the model earns its keep, and where it stops.
Full argument, tables and threats to validity in the paper
Thank you
read the whole thing
The paper, the dataset, the corpus and a working playground — zero.khe.money
Harshit Khemani · Kush Ahuja · Mohit Kumar Mishra · Kushagra Agrawal · submitted to Ms. Ankita Sharma