Hack Execution and Test Workflow
← Assembler internals · Common teaching contract · 中文
This guide follows executor.py and workflow.py to explain how Hack machine code becomes a native executable through Sail's C backend, and how unit, ISA conformance, and assembly integration tests form one regression loop.
Responsibility boundaries
This split keeps shared semantics reusable while each projects/<profile>.sail_project fixes one complete ISA model. Running another Hack program requires another generated ROM loader and driver, not a change to fetch, decode, or CPU semantics.
Why profile projects have no main()
The ISA owns the architectural ROM and answers how one step fetches a raw word, decodes it, and changes architectural state. A runnable program must additionally define:
- which raw machine words to load into ROM;
- when execution is complete;
- which final assertions to check;
- which state to print.
Those program-specific answers do not belong in the ISA. executor.write_driver() generates a small Sail main() for each .hack artifact; it loads the image into model-owned ROM, then supplies run control around hack_step().
This resembles a real system boundary: a CPU model defines storage, fetch, decode, and instruction behavior, while a loader or simulator supplies the program image and run-control policy.
Full executor.run() sequence
Steps 6 and 7 are the persisted trust boundary: effective CLI overrides are serialized into .hack before driver generation. Step 12 is the publication boundary: final names change only after staged execution succeeds. See Why load_hack() exists and the common publication contract.
Which settings belong on which surface
Thus CLI coverage is intentional rather than mechanical: only runtime settings use CLI > source > default. CLI cannot silently replace source assertions or descriptions.
Build artifacts
For default hack16 multiply, the per-program output prefix is shown below. --profile hack32 replaces the profile directory with hack32; asm identifies the direct assembly frontend:
The executor can generate:
A complete run creates these files in a temporary staging subdirectory inside .build/<profile>/asm/<program>/, beside the final artifact files. It compiles and executes the staged closure there, then publishes the whole per-program closure only after execution and assertions succeed. Backup-and-replace publication restores previous destinations if installation fails; a failed staged build or run leaves the previously published closure intact.
LoadedHack is the real executor input
After reload, the object contains only:
There is no parser AST, symbol table, or hidden Python assembly state. Driver behavior can come only from:
- machine words in
.hack; - structured metadata at the top of
.hack; - optional per-word teaching comments reloaded from
.hack; - semantics from the selected profile project.
Here, the loader means the host-side Python artifact loader isa/hack/tools/artifact.py::load_hack(), not generated load_program() or the Sail model's fetch_hack(). The artifact loader and driver do not consume every manifest field in the same way:
How the driver loads model-owned ROM
The selected profile composition owns the complete architectural instruction memory and shared raw fetch operation:
fetch_hack(pc) returns the raw profile-width word—16 bits for hack16, 32 bits for hack32; it does not decode it. The model's hack_step() composes the architectural path:
write_driver() no longer emits instruction_at, execute_at, or an address-to-decode match. It emits a load_program() function containing raw assignments instead:
The shown load lines are the default hack16 form; hack32 lines contain 32-bit literals. Python only reloads words and writes these ROM load lines. It never decodes A- or C-instructions; legality, decode, and execution remain in Sail. A legal hack_step() returns unit; an illegal encoding throws HackIllegalInstruction(word) before execution starts. The generated driver calls hack_step() directly and lets an unexpected exception fail the run.
The same none / summary / full level used for .hack controls driver comments. At the default summary, each ROM[index] = word line carries normalized assembly source, Hack+ words also carry [i/n] source => canonical, and inline assembly comments stay at the far right. full adds exact original source text and expands the explanations for assertion sources and final output. These comments come from LoadedHack.word_comments, proving they survived the file boundary.
load_program() writes only the loaded image. Generated main() calls execution_should_continue(PC, steps) before each attempt, so it never calls hack_step() beyond the program's loaded ROM extent.
How the driver stops and validates completion
The generated driver centralizes the subtle run condition in one small function:
The loop itself stays direct:
loaded_rom_words is generated from the reloaded artifact's word count—for basic_alu it is 53, covering ROM addresses 0..52. It is not max_steps and does not limit how many times a loop may revisit those addresses. Each lowered_halt_n value comes from completion.addresses in the reloaded opening //% manifest block; the bit pattern has no special meaning in the Hack ISA.
The condition becomes false in three cases, which the driver distinguishes after the loop:
PCreaches a HALT address recorded in.hackmetadata: valid completion;PCleaves the loaded ROM image: explicit execution error;stepsreachesdriver_max_steps: valid only for an explicitly requested bounded snapshot, otherwise watchdog failure.
Why the HALT loop does not execute here
Hack+ HALT lowers to:
Metadata records the private-label address. The driver sees completion when PC reaches that address and exits before executing the loop. The two loop words remain in .hack, so the artifact is still legal for its selected profile and would stay in the self-loop on a matching implementation.
This is deliberately driver policy. Neither profile has an encoded HALT instruction: Hack+ HALT is an assembler convenience, ROM end comes from the loaded artifact, and neither is an architectural outcome from the selected project. If a future profile gains a real HALT opcode, decoding and returning that outcome would belong in its model.
The 32768-word ROM edge case
PC is 15 bits and cannot exceed 32767. For a completely full 32768-word image, PC >= loaded_rom_words can never become true: there is no representable “one past the ROM” address. A recorded HALT can complete such a program normally; otherwise the explicit or default step budget still bounds the host run and fails as a watchdog unless the source explicitly requests a bounded snapshot.
Generated run loop
The generated main() loads raw words, executes model steps under completion and watchdog control, then evaluates source assertions:
The driver owns completion checks, watchdog enforcement, source assertions, and final output. Fetch, decode, execute, and illegal-word rejection remain one model-owned hack_step(); generated main() supplies only run-control policy.
Two meanings of max_steps
The effective limit has this precedence:
A CLI value replaces source .max_steps in the generated .hack metadata before strict reload. The generator then emits driver_max_steps from reloaded metadata, or uses the default for normal executor runs; execution_should_continue requires steps < driver_max_steps. This guarantees a finite host run even when a program loops forever without reaching lowered HALT metadata.
Watchdog mode
If a program has HALT, reaching the limit requires that PC has already reached one of its generated lowered_halt_n values:
For a program with no HALT and no explicitly requested bounded snapshot, budget exhaustion instead raises maximum step limit reached without lowered HALT or an explicit bounded snapshot. In both cases the meaning is: “the limit prevents an infinite host run; exhausting it is not successful completion.”
Deliberately bounded snapshot
If a program has:
- no HALT;
- at least one
.assert; - an explicitly configured source
.max_stepsor CLI--max-stepsvalue;
then the limit means “run N retired instructions and inspect state.” If execution remains within the loaded image and does not fault, the driver executes exactly N steps and then checks the assertions without requiring lowered-HALT completion.
This is useful for loop snapshots, but terminating programs should prefer HALT because it expresses completion more directly.
Turning source assertions into Sail
The assembler canonicalizes .assert into Assertion; the executor emits the actual Sail expression.
Target mapping
Comparison modes
Equality and inequality always compare architectural bits and reject any signed(...) or unsigned(...) wrapper. Ordered operators require one of those wrappers explicitly. The common parser enforces that distinction before Hack validates A, D, PC, R0..R15, or RAM[index].
Failure messages retain the assembly source line:
A runtime failure can therefore lead back to .asm, rather than only to generated driver code.
Sail-to-native compilation
compile_and_run() performs three stages.
1. Ensure project-local Sail
This selects and validates the pinned official binary under .pixi/sail/; it does not fall back to arbitrary system Sail.
2. Generate C with Sail
Conceptually, the executor gives Sail the selected isa/hack/projects/<profile>.sail_project source closure plus <output-prefix>.driver.sail. Sail type-checks that profile and driver together, then emits <output>.c and <output>.h.
3. Compile, link, and run
Platform compiler:
- Windows:
x86_64-w64-mingw32-gcc; - Linux:
gcc.
Link inputs include:
- Sail-generated C;
- Sail runtime components
rts,elf,sail,sail_config,sail_failure, andcJSON; support/sail_windows_compat.cand its forced-include header;- GMP through
-lgmp.
The Pixi/Conda prefix locates GMP headers and libraries. After a successful link, the executor starts the native program immediately. Every subprocess uses check=True, so a nonzero exit propagates as failure.
What generated main() prints
After execution and successful assertions, the driver:
- prints
ASSERT PASSwhen assertions exist, otherwiseRUN COMPLETE; - prints
A,D, andPC; - prints
R0..R7.
These registers are a human-readable summary, not the test oracle. Sail assert and process exit status determine success; workflow does not scrape console output to guess whether a program passed.
workflow.py as the control plane
Workflow answers which repository example a command names and which tool to invoke. It does not execute ISA semantics.
Source-based program discovery
discover_programs() scans direct program sources matching:
Files are sorted by filename, and the filename stem becomes the command name. For example, multiply.asm is selected by just hack run multiply.
Every bundled source must contain exactly one nonempty description directive:
The assembler parser validates .description, source_description() exposes it to workflow discovery, and assemble_text() stores it in AssemblyMetadata. artifact.py serializes the same text as manifest description; it emits no machine word.
discover_programs() additionally verifies that each file resolves inside the Hack package and actually exists. Direct glob("*.asm") intentionally ignores nested directories and non-assembly files.
Action dispatch
isa/hack/justfile exposes these actions directly as just hack ... recipes. The root test and clean-all recipes aggregate the module commands.
What just hack test verifies
pixi run just hack check checks both projects; pixi run just hack check --profile hack32 is the narrow gate. The complete test command always runs both profiles.
Layer 1: assembler tests
test_assembler.py covers parsing, lowering, encoding, metadata round trips, and malformed-input rejection. It runs quickly without compiling a native program for every assertion.
Layer 2: executor component tests
test_executor.py calls write_driver() and inspects generated Sail to fix:
- raw
ROM[index] = wordloading; - checked
hack_step()dispatch and explicit loaded-image boundary failure; - bounded-loop behavior;
- bit-exact, signed, unsigned, and PC assertions;
- reloaded metadata as the source of truth.
Some tests monkeypatch native compilation to isolate Python boundary behavior.
Layer 3: discovered-program end-to-end tests
Workflow invokes every discovered program with:
An example without .assert fails immediately, preventing bundled sources that merely run without verifying anything.
Layer 4: direct Sail ISA conformance
tests/sail/hack16/conformance.sail and tests/sail/hack32/conformance.sail directly check profile known words and legality plus shared alu(), jump, destination, old-A, PC wraparound, and fault non-commit behavior. This layer tests Sail ISA functions independently of the generated-driver integration path.
Repository-wide tests
First runs global tool tests such as tests/test_install_sail.py, then invokes just hack test. Future ISA modules should be aggregated at the root without teaching the root workflow Hack executor internals.
Debugging a failed run
Follow artifact boundaries from front to back:
- Source error: inspect assembler
line Ndiagnostics. - Encoding error: run
just hack assemble NAMEand inspect the source annotations beside.hackwords. - Metadata error: inspect the opening
//%manifest block in.hack. - Driver error: open
.driver.sailand inspectload_program()ROM lines, checked-step run control, and assertions. - Sail type error: read Sail
-coutput directly. - C build error: verify compiler, runtime include path, GMP, and compatibility layer.
- Architectural assertion failure: use the
.asmline in the message and final register summary. - Batch workflow failure: run one
just hack run NAMEbefore debugging source discovery or batch orchestration.
Do not start by adding prints to workflow.py. It usually propagates a lower-layer error; the strongest evidence is more likely at the .hack, .driver.sail, or generated-C boundary.
Function reading order
Executor
runwrite_driver_assertion_expressioncompile_and_runmain
Workflow
discover_programsselected_program/source_pathassemble/runtestmain
Suggested exercises
Exercise 1: read a generated driver
Open isa/hack/.build/hack16/asm/multiply/multiply.driver.sail (or the corresponding .build/hack32/asm/multiply/ artifact) and locate:
- the
ROM[0] = wordload line; - the HALT address;
- the generated run loop and checked
hack_step()call; - the Sail assertion corresponding to
R2 == 42;
Exercise 2: write a bounded program
Create a self-loop without HALT, add .max_steps 10 and a final .assert. Predict the generated loop and limit checks, then compare with the real .driver.sail.
Exercise 3: distinguish test layers
Deliberately introduce, one at a time:
- an invalid assembly mnemonic;
- an incorrect
.assertexpected value; - a Sail type error in a generated driver;
- a missing or duplicate
.descriptionin one bundled source.
Observe which layer catches each error. A clear toolchain should fail as close as possible to the root cause.
Return to the Hack overview or open the Hack package reference.