The Hack Instruction Set Architecture
← Tutorial · Hack overview · 中文 · Assembler internals →
This document focuses on the Hack ISA itself: the state visible to software, machine-instruction encodings, and the state transition caused by each instruction. The detailed bit diagrams use the default hack16 profile; profile differences are called out explicitly. See the Hack package reference for commands, assertions, and annotated .hack metadata.
What an ISA is
An instruction set architecture (ISA) is the contract between software and a processor implementation. It defines:
- registers and memory visible to programs;
- how machine words decode;
- what every instruction reads, computes, and writes;
- how the program counter advances or branches.
An ISA does not specify how many NAND gates form an adder, how long signals take to settle, or whether an implementation uses a pipeline or cache. Those are circuit or microarchitecture concerns. Two processors implement the same ISA when they produce the same architectural result from the same machine state and instruction.
The layers used by this project are distinct:
Only A or C instructions for the selected profile ultimately reach the modeled CPU. Commands default to hack16; --profile hack32 selects the extension.
Architecture overview
Both profiles use separate instruction and data storage: programs are fetched from ROM while data is read and written in RAM, commonly described as a Harvard architecture. They share a 15-bit PC/address space but differ in word width.
Architectural state
Assembly M is not a separate register. It denotes memory at the current A address:
The standard assembler symbols R0..R15 are aliases for RAM[0]..RAM[15]; they are not additional architectural registers. For example, @R3 loads address 3 into A, and a following D=M reads RAM[3].
In hack16, A is 16 bits and an A instruction carries a 15-bit immediate. In hack32, A is 32 bits and an A instruction carries a 31-bit immediate. In both profiles, RAM addresses and jump targets use only A[14:0]; hack32's upper A data bits still participate in the 32-bit ALU.
Hack platform memory map
The complete nand2tetris Hack platform normally interprets data addresses as:
The Sail model treats every address from 0 through 32767 as plain RAM. Instruction addressing rules are modeled; screen refresh and keyboard input are outside its scope.
Machine-instruction formats
Each profile has only A and C instruction forms. The following diagrams show the canonical 16-bit hack16 layout; hack32 wraps the same C fields in a wider encoding.
A instruction
Standard assembly syntax:
Semantics:
For example, @42 encodes as:
The operand may be a decimal value or a symbol. Symbol resolution is an assembler operation. hack16 encodes 0 plus 15 immediate bits; hack32 encodes 0 plus 31 immediate bits.
C instruction
Standard assembly syntax:
The fields select:
a + comp: the ALU computation;dest: any combination ofA,D, andMto receive the result;jump: whetherPCreceives the address from the old value ofA.
dest and jump are optional; comp is required. In hack32, the 16-bit C layout shown above is the low half of the word and the high half is 0xFFFF, so the complete encoding is 1111111111111111 111accccccdddjjj.
comp: ALU operations
The fixed ALU input x is D. When a=0, y=A; when a=1, y=M=RAM[A].
This table lists the canonical mnemonics accepted by the assembler. — means that no canonical assembly form names that combination. The gate-control ALU itself is total for all 64 values of cccccc; a legal C word may therefore decode and execute a control value that the assembler never emits.
Bitwise and arithmetic results are truncated to the selected word width, so overflow wraps modulo 2^16 in hack16 and modulo 2^32 in hack32. Hack has no separate flags register. Jump conditions inspect whether the current ALU result is zero and whether its most significant bit is one.
dest: write-back targets
From high to low, the ddd bits are write enables for A, D, and M:
Multiple destinations are architecturally simultaneous. In particular, for AM=... or AMD=..., A receives the new result while M still writes RAM at the address from old_A.
jump: conditional control flow
A jump interprets the ALU result out as a two's-complement value of the selected profile width:
When the condition holds, the target is the low 15 bits of A from the start of the instruction—not the ALU result or a newly written value of A.
Executing one instruction
Sail's execute function can be summarized as the following architectural pseudocode.
A instruction
C instruction
Saving old_A is the crucial detail. For example:
If the condition holds, this one instruction:
- computes
D+1; - writes the result to
A; - writes the result to RAM addressed by old
A; - jumps to the ROM address from old
A.
That is why shared model/core.sail begins the C-instruction path by preserving the old value of A.
From standard assembly to machine code
The assembler uses two passes:
- Parse source and expand Hack+ into standard A/C instructions.
- In pass one, record the ROM address of every
(LABEL); labels occupy no ROM word. - In pass two, resolve A-instruction symbols and encode every machine instruction.
- Allocate previously unknown variable symbols consecutively from RAM address
16.
Predefined symbols include:
R0..R15;SP=0,LCL=1,ARG=2,THIS=3, andTHAT=4;SCREEN=16384andKBD=24576.
For hack16, an A instruction is 0 followed by its 15-bit value, and a C instruction is concatenated as:
For hack32, an A instruction is 0 followed by its 31-bit value. A C instruction is 0xFFFF followed by the same canonical 16-bit C encoding.
For example:
uses comp=D, dest=M, and no jump:
How Hack+ lowers to real instructions
Hack+ expansion happens before label resolution and machine-code encoding. Every expanded line uses canonical nand2tetris A/C assembly syntax; final machine words use the selected profile's encoding.
Subtleties worth making explicit:
SET R0, R1stores the symbolR1's address value1inR0; useMOV R0, R1to copyRAM[R1].- Binary memory operations use destination-first order:
SUB R0, R1meansRAM[R0] -= RAM[R1]. MOV, binary memory operations, and conditional pseudoinstructions overwriteD.JNZandJNEare equivalent;JNEmatches the standard Hack jump mnemonic.HALTis not a Hack instruction. The assembler creates a unique__HACKPLUS_HALT_nlabel and a two-instruction self-loop, then records that ROM address as execution metadata. This project's executor stops when it reaches the address; an implementation of the selected profile would remain in the loop.
Complete lowering example
Source:
Conceptual standard Hack assembly after expansion:
Only after this expansion does pass one compute the ROM addresses of DONE and the private HALT label; pass two then emits words for the selected profile. Expanded instructions therefore consume real ROM addresses, and all label addresses refer to the expanded program.
Mapping this ISA to Sail
The current model separates shared semantics from profile layout:
The C-instruction fields use named bit-vector types rather than anonymous bit/bits(n) positions. This documents their roles throughout decoded instructions, alu, and should_jump without changing the raw encoding domain.
decode_hack checks legality before using the profile mapping. In hack16, words with bit 15 clear are A instructions and only the 111 prefix is C; 100, 101, and 110 are illegal. In hack32, words with bit 31 clear are A instructions and C requires high 16 bits 0xFFFF plus a low-half 111 prefix; every other word with bit 31 set is illegal. An illegal word throws HackIllegalInstruction(word) before execute begins, preserving the raw word without committing state.
Once the C prefix is legal, all six-bit ALU controls, both a values, and every dest/jump mask are in the machine encoding domain. Shared alu defines all 64 control results at the selected word width. The assembler deliberately emits only the canonical mnemonic subset shown above; assembler vocabulary must not be confused with machine-code legality.
Compatibility boundary
hack16 is the canonical nand2tetris baseline. hack32 is not an ordinary nand2tetris .hack format: its words are 32 bits, A immediates are 31 bits, and C words carry the additional 0xFFFF high half. Matching assembly mnemonics do not make the binaries interchangeable; use an explicit hack32 profile and a profile-aware manifest loader.
Model boundaries
Keep the complete Hack platform separate from this executable model when interpreting results:
- ROM is declared by the selected profile composition; the generated driver initializes raw profile-width words, while
fetch_hackandhack_stepown instruction fetch and dispatch. - RAM has no Screen or Keyboard device side effects.
HALT, Hack+,.assert, and.max_stepsbelong to the tool layer, not the Hack ISA.- The model specifies architectural transitions, not gate delays, clock-edge details, or nand2tetris HDL.
- Execution uses Sail's C backend and makes no claim of a completed formal equivalence proof.
Study the model in code
- Read the A/C encodings and execution pseudocode in this document.
Open one file under
model/profiles/, follow its include intomodel/core.sail, then return to the profile's mapping clauses; useencdec(instruction)directly for encoding and tracefetch_hack→decode_hack→encdec(word)→execute→hack_stepfor execution. - Compare standard assembly and Hack+ in
programs/basic_alu.asm. - Run
pixi run just hack assemble basic_aluand inspectisa/hack/.build/hack16/asm/basic_alu/basic_alu.hack; add--profile hack32to inspect.build/hack32/asm/basic_alu/basic_alu.hack. - Read
tests/sail/hack16/conformance.sailandtests/sail/hack32/conformance.sailto see profile rules turned directly into tests.
Specification sources
- nand2tetris Chapter 4 / Project 04: Hack machine language;
- nand2tetris Project 05: Hack CPU, Memory, and Computer;
- nand2tetris Project 06: Hack assembler;
- Sail Language Reference: Sail language and backends.