Hack Assembler Internals
← Hack ISA · Common teaching contract · 中文 · Execution and tests →
This guide follows isa/hack/tools/assembler.py and artifact.py. It explains how source becomes profile-specific Hack machine code, then crosses a strict persisted-artifact boundary. Commands default to 16-bit hack16; --profile hack32 selects 32-bit words. Shared directive and manifest rules are defined by the common teaching contract. assembler.py remains a pure library; assembler_cli.py is the thin CLI boundary that handles arguments and publication.
The assembler's responsibility
Input can contain three kinds of information:
Parsed source has two destinations:
records: realMachineWordvalues that enter ROM;AssemblyMetadata: HALT addresses, assertions, effective runtime settings with origins, source identity, and.description.
Workflow reads the description through the same parser for program discovery. artifact.py serializes it into the .hack manifest, while directives still emit no ROM word. This keeps teaching metadata observable without pretending that it is Hack code.
Why a hand-written parser is enough
Hack assembly is a line-oriented language with no nested expressions, precedence, or multiline syntax. A parser generator would add little value here. Regular expressions recognize each line shape, while Python dataclasses form a typed intermediate representation.
Simple does not mean permissive. The parser rejects:
- malformed labels and symbols;
- unknown
comp,dest, andjumpforms; - duplicate
.descriptionor.max_stepsdirectives; - unknown dot directives;
- out-of-range A operands, RAM targets, and assertion values;
For an educational assembler, explicit rejection is better than guessing the author's intent.
Phase 1: preserve source identity
Every meaningful line first gets a SourceLine:
It retains both the line number and original text. If one pseudoinstruction later expands into four machine instructions, all four still point to the same SourceLine. The writer can therefore emit:
This is a useful compiler principle: source locations belong in the IR; do not try to reconstruct them only when reporting an error.
Phase 2: typed statements
Parser.parse() produces a Statement union instead of integers:
The parser dispatch order is deliberate:
- remove text after
//and trim whitespace; - recognize
.description,.assert, and.max_steps; - recognize labels;
- recognize A instructions beginning with
@; - use the first word to recognize Hack+;
- parse everything else as a C instruction.
Dot directives come first so an invalid .assert or .description gets a directive-specific error rather than an ambiguous “bad C instruction.” Handling @ separately similarly distinguishes an invalid operand from a non-A instruction.
source_description(text) scans the same typed statements for workflow discovery. expand() omits DescriptionDirective from code lowering, and assemble_text() copies its text into AssemblyMetadata.description. artifact.py then serializes that value as manifest description; it never enters ROM.
A-instruction validation
An A operand is either:
- a decimal value in the selected A-immediate range:
0..32767forhack16,0..2147483647forhack32; - a Hack symbol matching
[A-Za-z_.$:][A-Za-z0-9_.$:]*.
The parser retains symbolic operands as strings because label addresses are not known yet and variable allocation must not happen early.
C-instruction normalization
Whitespace is allowed around fields and then removed. This source:
becomes:
The COMP, DEST, and JUMP tables validate each field. They are both encoding tables and the accepted language definition; unknown combinations never reach assembly passes.
Phase 3: lower Hack+
expand() walks the typed statements:
- standard instructions and labels enter the code stream;
- directives are collected as side metadata;
- each
PseudoInstructionbecomes standard A/C instructions.
For example:
lowers to:
Each lowered instruction carries PseudoExpansion(instruction, index, count), while source still points at the original JEQ line. The four results therefore retain both their canonical spelling and positions [1/4] through [4/4].
Why lowering must precede label collection
Consider:
SET occupies four ROM words, so AFTER=4. If pass one counted source lines before expanding SET, it would incorrectly assign AFTER=1.
The valid order is therefore:
HALT is also lowered here into a unique private label:
Its private-label address is additionally collected in halt_addresses for the executor.
See the ISA guide for the complete Hack+ lowering table.
Phase 4: two-pass assembly
Pass one: build the ROM symbol table
assemble_text() starts with predefined symbols:
R0..R15 map directly to RAM addresses 0..15; they are predefined assembler symbols, not architectural registers. SP/LCL/ARG/THIS/THAT are aliases for the first five of those same addresses.
It then walks the lowered code:
- a
Labelrecords the current ROM address; - each A/C instruction enters the instruction list and increments that address;
- labels occupy no machine word;
- duplicate labels and programs beyond 32768 words fail immediately.
Every branch label has a stable address when pass one finishes.
Pass two: resolve A symbols and encode
Pass two visits only real A/C instructions.
For an A instruction:
- use a decimal operand directly;
- look up a known symbol;
- allocate an unknown variable consecutively from RAM address
16; - place the value after a leading zero in the selected profile: 15 immediate bits for
hack16, 31 forhack32.
Later uses of the same variable reuse its first allocation.
A C instruction uses the same canonical fields in both profiles:
For example, D=M+1;JGT gives:
Each result is wrapped as:
For ordinary A/C instructions, expansion is None. For Hack+ output, it is a PseudoExpansion containing the canonical instruction and its [i/n] position, so machine code, original source, and lowering provenance remain connected.
Metadata is a side channel, not code
AssemblyMetadata contains:
Directives do not advance ROM. Removing all .assert lines must leave every standard machine word unchanged; the test suite explicitly fixes this property.
What assertion assembly does
The assembler only:
- uses the public parser to enforce bit-exact equality and explicit
signed(...)/unsigned(...)ordered comparisons; - validates Hack targets and widths, then canonicalizes representation;
- retains the source line.
It never reads final registers or evaluates assertions in Python. The executor turns metadata into Sail assertions inside the generated driver.
That separation is intentional: Python translates; Sail executes and judges architectural state.
write_hack(): an auditable artifact
artifact.py writes two record classes.
Opening structured manifest block
write_hack() delegates to the shared renderer. At summary and full, it writes one indented canonical S-expression as a contiguous block at the start of the file, with every line prefixed by //% . At none, it writes the same form as one compact prefixed line. The schema is verylogic.annotated-image; the public envelope contains runtime max_steps, assertions, completion, and ISA metadata. Hack direct assembly omits empty frontend provenance. Hack-specific validation fixes isa=hack, profile=hack16|hack32, source kind asm, profile word width, memory dimensions, and lowered-self-loop completion; standard is never a valid profile. Runtime values carry cli, source, or default origins.
Source-mapped machine words
This default example begins with a normal 16-bit hack16 word. A hack32 line begins with 32 binary digits and is not an ordinary nand2tetris .hack word. This project's profile-aware executor reconstructs the complete contract from the manifest and annotations.
Teaching detail levels
The optional comments argument controls explanatory content without changing machine words or semantic metadata. The manifest is always present and records the selected comment level:
summary is the default. It shows ordinary A/C source as well as pseudoinstruction lowering, with inline assembly comments at the far right. The opening //% manifest block is written at every level. Run pixi run just hack assemble multiply, or override max_steps with pixi run just hack assemble multiply summary --max-steps 10000. apply_runtime_overrides() resolves CLI over source/default before write_hack(), so both each effective value and its origin cross the artifact boundary.
Why load_hack() exists
The executor does not retain the in-memory AssemblyResult. It calls write_hack() and then reloads the file through load_hack().
This creates an explicit boundary:
Without reload, a writer bug could omit metadata while execution still “worked” from the old in-memory object. Reload ensures execution depends only on the artifact actually delivered to disk.
Strict Python artifact-loader rules
The loader here is isa/hack/tools/artifact.py::load_hack(). It verifies:
- exactly one contiguous manifest block appears at the start of the file, uses the comment-level-specific formatting, and has the exact common version-1 key set;
- ISA/profile, source identity, description, comment level, runtime value/origin, assertions, completion, and
isa_metadatahave exact types and legal values; - equality assertions use
bitsmode with no wrapper, while ordered assertions use explicitsignedorunsignedmode; - each machine word has exactly the selected profile width (16 or 32 binary digits), and comment presence matches the manifest level;
- each completion address points at an
@address; 0;JMPself-loop; - ROM has at most 32768 words.
load_hack() accepts only the strict annotated format and rejects a file without an opening manifest block or with a profile mismatch. Raw nand2tetris interoperability requires a separate explicit frontend because a raw image cannot truthfully supply descriptions, assertions, watchdog origin, or completion metadata. hack32 additionally has 32-bit words and is not ordinary nand2tetris .hack compatibility.
Error design
Source problems use AssemblyError(line, message), which the CLI reports through argparse. Because locations entered the IR at parse time, errors involving lowered instructions still identify the user's original line.
Artifact problems use ValueError messages containing file path and artifact line number. These are different error domains:
AssemblyError: invalid source program;ValueErrorfromload_hack: damaged or untrusted intermediate artifact.
Collapsing both into “assembly failed” would hide which boundary broke.
Call map
Recommended function reading order:
Parser.parseParser._parse_instructionexpandassemble_textwrite_hackload_hack
How tests cover the assembler
test_assembler.py fixes contracts at every phase rather than checking only a few example programs:
Integration runs of programs/*.asm then verify that parsed and encoded programs produce the expected architectural state on the Sail ISA.
Suggested exercises
Exercise 1: perform two passes by hand
For this source, list lowered ROM addresses, the symbol table, and machine words:
Notice that variable sum is allocated at 16, while LOOP uses a ROM address from the expanded program.
Exercise 2: design a pseudoinstruction
Write its standard Hack expansion first, then answer:
- which registers does it clobber?
- in which phase must it lower?
- does it need a private label?
- how will it preserve
SourceLine? - should tests fix words, source mappings, or metadata?
Only then modify PSEUDO_NAMES, parsing patterns, and expand().
Exercise 3: inspect a round trip
Open isa/hack/.build/hack16/asm/multiply/multiply.hack (or .build/hack32/asm/multiply/multiply.hack after assembling with --profile hack32)
- each source pseudoinstruction;
- its standard expansion;
- ROM address;
- profile-width machine word;
- file-level metadata.
Continue with Execution and test workflow to follow this .hack artifact into a running program.