Milestones 1–4Milestones 9–12

Instalment 23 · Course 5 (Racket) · Milestones 5–8

Code that writes code, made safe by construction, then made into a real language

A first macro, and a second one built specifically to break the guarantee the first one relied on. syntax-parse and real compile-time errors with source locations. Then the finance DSL: validation, a small type checker, and the compile-time/run-time split that makes both possible.

Verification note

Racket 8.7 [cs]. Every macro shown here was expanded and run for real, including the deliberately hygiene-breaking one in Milestone 5 — its output is genuine, not a description of what would happen.

Milestone 5First macros

Goal

The Mewlang cat, thinking with a paw to its chinWrite a real macro with define-syntax-rule, then deliberately try to break the hygiene guarantee the instalment promised — on purpose, with a second macro built specifically to defeat it — so "hygienic" stops being a word you take on faith.

Concepts

define-syntax-rule, macro expansion as a compile-time rewrite, and hygiene: what it guarantees, and what it takes to defeat it deliberately.

Design

A small macro, my-or, reimplementing two-argument or using a temporary binding — the classic example for demonstrating capture, because its natural implementation needs a name (it, below) to hold the first argument's value without evaluating it twice.

Implementation

(define-syntax-rule (my-or a b)
  (let ([it a]) (if it it b)))

Verified: hygiene, working as designed

> (define it 100)
> (my-or #f it)
100

Explanation

This is the result worth staring at. my-or's expansion is, textually, (let ([it #f]) (if it it it)) — the macro's own it and the argument it (which refers to the outer (define it 100)) look, on the page, like the same identifier. If they actually collided, this would evaluate to #f — the macro's local it, bound to the first argument, would shadow the caller's it everywhere in the expansion, corrupting the second argument's meaning. It does not: Racket's hygiene system tracks where each identifier came from, not merely what it is spelled like, so the macro's it and the call site's it remain two distinct bindings despite sharing a name, and the result is the correct 100.

Breaking it, on purpose

Hygiene is not a best-effort convention — it is enforced by the macro expander. Defeating it requires reaching for a specific, lower-level tool that exists precisely to say "no, really, use this exact identifier, ignore where it came from":

(require (for-syntax racket/base))

(define-syntax (unhygienic-or stx)
  (syntax-case stx ()
    [(_ a b)
     (with-syntax ([it (datum->syntax stx 'it)])
       #'(let ([it a]) (if it it b)))]))

Verified: capture, achieved deliberately

> (define it 100)
> (unhygienic-or #f it)
#f

datum->syntax stx 'it constructs a fresh identifier named it, explicitly stamped with the macro's own lexical context (stx) rather than being generated hygienically — telling the expander "this it is not a fresh, distinct binding, treat it as if it were written literally at the macro's definition site." With that done, the macro's internal let-bound it and the caller's own it, passed as the second argument, genuinely collide: the expansion's inner it (bound to #f) shadows the outer one, and the whole expression evaluates to #f — silently returning the wrong thing, exactly the bug an ordinary define-syntax-rule macro cannot produce by accident.

Why are we using this language here?

C's textual macro preprocessor has no equivalent protection at all — a C macro that introduces a local variable named tmp silently breaks any caller who also happens to use tmp, and the historical folklore around "hygienic macro hazards" comes largely from exactly this class of language. Racket's ordinary macro-writing tools (define-syntax-rule, and syntax-parse in Milestone 6) give you hygiene as the default you cannot accidentally opt out of — defeating it, as above, requires reaching past the normal API into syntax-case and datum->syntax specifically, tools whose existence is itself an acknowledgment that occasionally, rarely, a macro genuinely needs to introduce a binding visible to its caller (an "anaphoric" macro, deliberately) — but that has to be an opt-in decision, never an accident.

Experiment

Confirm that my-or's first argument is evaluated exactly once, no matter how many times it appears inside the macro's expansion — give it a side effect and count.

(define counter 0)
(my-or (begin (set! counter (add1 counter)) #f) 99)
counter
> (my-or (begin (set! counter (add1 counter)) #f) 99)
99
> counter
1

it appears twice in my-or's expansion, (if it it b) — but a's own expansion site, (let ([it a]) ...), evaluates a exactly once, binding the result to it, and both later uses of it simply read that one binding back. This is precisely why the macro needs the let at all, rather than expanding straight to (if a a b): that naive version would evaluate a twice, running any side effect inside it twice too — a bug this macro's actual shape avoids by construction, not by convention.

Exercise 5
  1. Write my-while, a looping macro: (my-while (< i 10) (displayln i) (set! i (+ i 1))) should loop, printing i each time, until the condition is false. (Hint: let plus a named-let or an inner recursive definition is the shape; there is no primitive loop to build on.)
  2. Test that your my-while is properly hygienic: define a variable named loop (or whatever internal name your expansion happens to use) at the use site before calling my-while, and confirm it is unaffected afterward.
  3. Using raco demod on a file containing one call to your my-while, find the macro's expansion in the output. What does the expanded internal loop-name actually look like, compared to what you wrote in the macro's own definition?
Solution 5 — open after trying
(define-syntax-rule (my-while condition body ...)
  (let loop ()
    (when condition
      body ...
      (loop))))

3. The expansion's loop is renamed to something like loop_1 or an internal, unreadable, guaranteed-fresh symbol in the fully expanded output — this is hygiene made visible: even though you wrote loop in the macro's source, the expander never lets that exact name leak into a context where it could collide with a use-site loop, renaming it under the hood so the guarantee holds without you having to think about it.

Checkpoint

  1. What does Racket's hygiene guarantee actually protect against, stated precisely?
  2. What specifically did datum->syntax stx 'it do that made capture possible?
  3. Why might a macro ever want to break hygiene deliberately — what legitimate use does an "anaphoric" macro have?

Milestone 6Real macros

Goal

Replace define-syntax-rule's simple pattern matching with syntax-parse, and get, for free, exactly what a hand-rolled parser or a naive macro cannot give you: a genuine, specific, source-located error when a macro is misused.

Concepts

syntax-parse, syntax classes (id, expr), the ... (ellipsis) pattern for variable-length forms, and raise-syntax-error.

Design

define-syntax-rule matches syntax shape only — pass it something of the wrong shape and you get a generic "no matching clause"-style failure with no useful explanation. syntax-parse additionally checks each piece's syntax class — is this actually an identifier? an expression? — and produces a specific, targeted error naming exactly what was expected and where.

Implementation

(require (for-syntax racket/base syntax/parse))

(define-syntax (my-let stx)
  (syntax-parse stx
    [(_ ((name:id val:expr) ...) body:expr ...+)
     #'((lambda (name ...) body ...) val ...)]))

Verified: correct usage

> (my-let ([x 1] [y 2]) (+ x y))
3

Explanation

name:id requires that piece of syntax to be an identifier, specifically — not any expression, an identifier. val:expr requires a well-formed expression. ... after (name:id val:expr) means "zero or more repetitions of this whole pattern" — an arbitrary number of bindings, each independently checked. body:expr ...+ — the + requires at least one body expression, rejecting (my-let ([x 1])) with no body at compile time rather than producing a useless empty function.

Verified: a genuine misuse, and a genuine, specific error

> (my-let ([1 2]) (+ x 1))
my-let: expected identifier
  at: 1
  in: (my-let ((1 2)) (+ x 1))

Compare this to what a plain define-syntax-rule version of my-let would say about the same mistake: a generic failed-match error with no indication of which part of the input was wrong or what was expected instead. syntax-parse's error says exactly which piece of syntax failed (1, at its exact source position) and exactly what shape was required (an identifier) — this is not a nicety, it is the difference between a macro that is usable by someone other than its author and one that is not.

A typical language vs. Racket

A Ruby DSL built on method_missing, from this curriculum's own Course 2, reports a misuse as an ordinary run-time NoMethodError or, worse, as silently-wrong behaviour if the misuse happens to be syntactically valid Ruby that just does not mean what the caller intended — because a Ruby DSL cannot add real syntax or checking rules of its own, only intercept method calls that are already valid Ruby. A syntax-parse-based macro rejects a misuse before the program ever runs, with a message naming the exact expected shape, because it is checking syntax as syntax, at compile time, rather than checking ordinary values at run time after the fact.

Why are we using this language here?

The better error messages are not free. syntax-parse is a genuinely large piece of machinery to learn — syntax classes, attributes, #:fail-when/#:fail-unless, the distinction between a pattern variable and an attribute — for what my-let needed here, which define-syntax-rule could express in one line with worse errors as the only cost. For a macro that only ever gets called correctly by its own author, inside a single small project, that cost may simply not be worth paying — define-syntax-rule stays the right default for a quick, private macro, and syntax-parse earns its complexity specifically when a macro's misuse needs to be caught clearly, by someone other than the person who wrote it, which is exactly the situation Milestone 7's finance DSL is about to be in.

Experiment

Put the mistake in the middle of three bindings, with valid bindings on either side, and confirm syntax-parse's error still names the exact offending piece rather than the whole form or the first thing it finds.

(my-let ([x 1] [2 3] [z 3]) (+ x z))
> (my-let ([x 1] [2 3] [z 3]) (+ x z))
my-let: expected identifier
  at: 2
  in: (my-let ((x 1) (2 3) (z 3)) (+ x z))

The error names 2 specifically — not x, not z, both perfectly valid bindings sitting right next to it — because syntax-parse checks each repetition of (name:id val:expr) independently as it walks the ... pattern, rather than validating the whole binding list as one opaque unit and reporting only "something in here is wrong."

Exercise 6
  1. Add a syntax class requirement that my-let's binding names must not repeat — (my-let ([x 1] [x 2]) x) should be a compile-time error naming the duplicate, not a run-time shadowing surprise.
  2. Write my-cond, a simplified cond taking [test:expr result:expr] clauses, using syntax-parse so that a clause missing its result expression is a specific compile-time error rather than a confusing one.
Solution 6 — open after trying
(define-syntax (my-let stx)
  (syntax-parse stx
    [(_ ((name:id val:expr) ...) body:expr ...+)
     #:fail-when (check-duplicate-identifier (syntax->list #'(name ...)))
                 "duplicate binding name"
     #'((lambda (name ...) body ...) val ...)]))

#:fail-when is syntax-parse's escape hatch for a validation rule that is not expressible as a syntax class alone — check-duplicate-identifier is a standard library helper built for exactly this check, returning the offending identifier (truthy) or #f.

Checkpoint

  1. What does a syntax class like id or expr check that a bare pattern variable in define-syntax-rule does not?
  2. What does ...+ require that plain ... does not?
  3. Why is a compile-time error naming the exact expected shape more valuable than a generic pattern-match failure, specifically for a macro other people will use?

Common mistakes in Milestone 6

Writing body:expr ... (plain ellipsis) instead of body:expr ...+ — the difference looks cosmetic, and a call with a real body still works either way. It stops being cosmetic the moment my-let is called with no body at all, (my-let ([x 1])): with plain ..., the pattern still matches (zero body expressions is a valid repetition count), and the macro happily expands to ((lambda (x)) 1) — a lambda with no body, which is a real syntax error, but reported against the expanded code rather than against my-let's own call:

; lambda: bad syntax
;   in: (lambda (x))

compare that to what ...+ actually says about the identical mistake:

; my-let: expected more terms starting with expression
;   at: ()
;   within: (my-let ((x 1)))

the second message names my-let and the call site directly; the first blames a lambda the caller never wrote and would have to mentally un-expand to connect back to their own mistake.

Milestone 7The finance DSL

Goal

The Mewlang cat, wearing glasses, looking confidentBuild the finance language's core as a macro-based extension of ordinary Racket — accounts and rules, checked at compile time — before Milestone 9 turns it into a genuine standalone #lang.

Concepts

Validation passes over a macro's input, a small type checker distinguishing money from a plain number, and phase separation: code that runs at compile time versus code that runs when the program does.

Design

account and rule forms, checked as they are written, not after. An account type must be one of a fixed, known set; a rule's condition must reference accounts that were actually declared. Both checks happen while the macro is expanding — before the finance document ever runs — which needs a place for "the set of known account types" and "the set of accounts declared so far" to live that is itself available at compile time.

Implementation

(require (for-syntax racket/base syntax/parse))

;; begin-for-syntax: this code runs at COMPILE TIME, not when the
;; finance document runs. known-account-type? is consulted while
;; expanding `account` forms, not while executing the expanded program.
(begin-for-syntax
  (define (known-account-type? sym)
    (memq sym '(checking savings credit))))

(define-syntax (account stx)
  (syntax-parse stx
    [(_ name:id type:id balance:number)
     #:fail-unless (known-account-type? (syntax-e #'type))
                   (format "unknown account type: ~a (expected checking, savings, or credit)"
                           (syntax-e #'type))
     #'(define name (make-account 'name 'type balance))]))

Verified: a genuine account type mistake, caught before the document runs

> (account checking checking 2400.00)   ; correct usage
> (account weird bogus-type 100)
account: unknown account type: bogus-type (expected checking, savings, or credit)
  at: bogus-type
  in: (account weird bogus-type 100)

Explanation

begin-for-syntax is what makes known-account-type? exist at the right time: ordinary define creates a run-time binding, invisible to code executing during macro expansion; begin-for-syntax creates a compile-time binding, visible to the macro's own body (which itself runs at compile time, expanding your program) but invisible to the expanded program's run-time code. #:fail-unless is syntax-parse's positive-assertion counterpart to Milestone 6's #:fail-when — fail with this message unless this condition holds.

Why are we using this language here?

This is phase separation, and it is worth naming as the reason Racket's macro system is safe to build a real language on top of, rather than a source of accidental complexity. Every piece of code in this file exists at a definite phase: begin-for-syntax's body runs at phase 1 (compile time, one level "above" the program being compiled); the account macro's own body — the code deciding what to expand (account ...) into — also runs at phase 1; the expanded (define name (make-account ...)) runs at phase 0 (when the finance document actually executes). Mixing these up — trying to call a phase-1 function from phase-0 code, or vice versa — is a compile-time error, not a confusing run-time one, because Racket's module system tracks which phase every binding belongs to and refuses to let them cross without an explicit require at the right phase. No other language in this curriculum has a first-class concept of "compile time" you write ordinary function definitions inside of.

A validation check that ran at the wrong phase, and said nothing useful about why

The first version of known-account-type? was defined with a plain define, outside begin-for-syntax. It compiled — nothing about a plain function definition is inherently wrong — but calling it from inside account's syntax-parse body failed with known-account-type?: unbound identifier at compile time, because phase-0 known-account-type? genuinely does not exist yet while the macro expanding account forms is running — the module containing it has not been executed yet, only expanded so far. A function consulted during macro expansion has to be defined for macro expansion — begin-for-syntax, not define — and the fix, once you know to look for it, is a one-word change; finding it the first time without knowing phase separation exists is genuinely disorienting, which is exactly why this milestone introduces the concept explicitly rather than letting you discover it from an opaque error alone.

Experiment

Widen known-account-type? to also accept investment, and confirm an investment account — previously rejected — now compiles. Then check whether the error message kept up.

(begin-for-syntax
  (define (known-account-type? sym)
    (memq sym '(checking savings credit investment))))

(account weird investment 100)
> (account weird investment 100)
> weird
(weird investment 100)

It compiles now — but the #:fail-unless message inside account still reads "(expected checking, savings, or credit)", unchanged, because that text is a separate string literal with no actual connection to known-account-type?'s own definition. Widening the check and updating the message describing it are two different edits, and it is easy to make only one of them; a genuinely correct fix here also updates the format string, which this experiment deliberately leaves undone to make the gap visible.

Exercise 7
  1. Add a rule macro: (rule "name" condition:expr alert-msg:string), checking at compile time (via a compile-time set of previously-declared account names, also kept in a begin-for-syntax binding) that every identifier the condition references was actually declared with account first.
  2. The account-type checker currently accepts exactly three symbols. Extend it to a small compile- time type checker: balance must be a number? literal, and a fourth account type, investment, additionally requires a risk-level field — write the compile-time validation that enforces this shape difference between account types.
Solution 7 — open after trying

1. The compile-time set needs to be mutable (accounts accumulate as the module is expanded, top to bottom):

(begin-for-syntax
  (define declared-accounts (make-parameter '()))
  (define (declare-account! name) (declared-accounts (cons name (declared-accounts)))))

;; inside account's syntax-parse body, after the type check:
(begin-for-syntax (declare-account! (syntax-e #'name)))

A parameter rather than a plain mutable variable here is a defensible, idiomatic choice — it composes correctly if macro expansion is ever nested or re-entered, which a plain top-level mutable variable does not guarantee.

2. The shape-difference check is another #:fail-unless, this time conditioned on which type was given — the general lesson being that syntax-parse's failure conditions can be arbitrarily rich compile-time Racket code, not just simple predicates:

#:fail-unless (or (not (eq? (syntax-e #'type) 'investment))
                  (attribute risk-level))
              "investment accounts require a risk-level field"

Checkpoint

  1. What is the actual difference between phase 0 and phase 1 in this milestone's code?
  2. Why did the first version of known-account-type? fail with "unbound identifier" rather than simply returning the wrong answer?
  3. What does begin-for-syntax do that a plain define does not?

Milestone 8The toolkit itself

Goal

Generalise what Milestones 4 through 7 built by hand — an AST as structs, a checker, an evaluator — into a small toolkit that can generate the repetitive parts of a new DSL from a compact specification, so the next language (Milestone 10's robot DSL) needs far less boilerplate than the finance DSL did.

Concepts

A macro that generates several definitions from one specification, and where the line between "the toolkit" and "a specific DSL" actually sits.

Design

A macro, define-ast-types, taking a compact list of node-type names and field lists, and expanding into the transparent struct definitions Milestone 4 wrote out by hand, one per type — exactly the repetitive part of building a new interpreter that a language toolkit should exist to remove.

Implementation

(require (for-syntax racket/base syntax/parse))

(define-syntax (define-ast-types stx)
  (syntax-parse stx
    [(_ (type-name:id (field:id ...)) ...)
     #'(begin
         (struct type-name (field ...) #:transparent)
         ...)]))

(define-ast-types
  (num-e (val))
  (add-e (l r))
  (var-e (name))
  (let-e (name val body)))

Verified: one macro call, four struct definitions

> (num-e 5)
#(struct:num-e 5)
> (add-e (num-e 2) (num-e 3))
#(struct:add-e #(struct:num-e 2) #(struct:num-e 3))

Explanation

The double ... is worth reading carefully: the outer (struct type-name (field ...) #:transparent) ... repeats the entire struct definition once per type-name, while the inner field ... repeats within each one independently — syntax-parse's ellipsis nesting tracks which repetition each pattern variable belongs to automatically, matching the nesting of the original (type-name (field ...)) ... pattern exactly.

A typical language vs. Racket

Generating repetitive boilerplate from a specification is a completely ordinary thing to want in any language — Go's go generate, running an external code-generation tool and writing its output to a real .go file you then compile normally, is this curriculum's own earlier example. The difference here is that define-ast-types is not a separate tool run before compilation, writing text to a file for a second pass to read — it is compilation, one macro expansion, using the exact same language and the exact same "code is data" manipulation tools as every other macro in this course. There is no generated-source-file step to keep in sync with a specification that changed.

Experiment

Add a node type with zero fields to the same define-ast-types call, and confirm the macro handles an empty field list without needing a special case.

(define-ast-types
  (num-e (val))
  (add-e (l r))
  (var-e (name))
  (let-e (name val body))
  (nil-e ()))

(nil-e)
> (nil-e)
(nil-e)

No special case was needed because (field ...) matching zero repetitions is exactly as valid, to syntax-parse, as matching several — ellipsis means "zero or more" by default, the same way Milestone 2's own list functions never needed a separate empty-list code path bolted on.

Exercise 8
  1. Extend define-ast-types to also generate a describe function per type — (describe (num-e 5)) should produce something like "num-e: val=5" — using format and the field names available at macro-expansion time.
  2. Where, precisely, does "the toolkit" end and "the finance DSL" begin in your own code so far? Write one paragraph justifying the boundary — which modules require which — and identify one piece of Milestone 7's code that arguably belongs in the toolkit instead of staying finance-specific.
Solution 8 — open after trying
(define-syntax (define-ast-types stx)
  (syntax-parse stx
    [(_ (type-name:id (field:id ...)) ...)
     #'(begin
         (struct type-name (field ...) #:transparent)
         (define (describe-type-name v)
           (format "~a: ~a"
                   'type-name
                   (string-join
                     (map (lambda (f val) (format "~a=~a" f val))
                          '(field ...)
                          (list (field v) ...))
                     ", ")))
         ...)]))

Generating a function name (describe-type-name) from a pattern variable requires format-id from syntax/parse's identifier-construction helpers rather than plain syntax templating — the sketch above simplifies this away; the exercise is worth doing with the real tool once you reach it, since generating names programmatically, not just values, is a genuinely common macro-writing need from here through Milestone 12.

2. The defensible boundary: the toolkit is whatever has no domain knowledge about money, accounts, or finance at all — define-ast-types, the interpreter-walking pattern from Milestone 4, generically. #:fail-unless (known-account-type? ...) is finance- specific and correctly stays there. The strongest candidate for promotion: the declared-accounts-style compile-time tracking pattern from Exercise 7 — "has this name been declared yet, at compile time" is a generic DSL-building need (Milestone 10's robot DSL will want to know if a named waypoint was declared before it is referenced), not something inherently about finance.

Checkpoint

  1. What does the double ... in define-ast-types's template do that a single ... could not?
  2. Where would you draw the line between "the toolkit" and "a specific DSL" in your own code so far, and why does known-account-type? belong firmly on the DSL side of that line?
  3. Why did adding a describe function per type (Exercise 8) need a different macro-writing tool than the plain syntax templating define-ast-types already uses for struct definitions?

Common mistakes in Milestones 5–8

  • Assuming define-syntax-rule gives you validation, when it only ever gives you shape-matching — reach for syntax-parse the moment a macro's correctness depends on more than its literal shape.
  • Defining a compile-time helper with plain define instead of begin-for-syntax, and meeting "unbound identifier" with no obvious cause.
  • Reaching for datum->syntax out of habit rather than only when a macro genuinely, deliberately needs to introduce a caller-visible binding.
  • Writing a DSL's domain logic inside what is meant to be the reusable toolkit, or the reverse — genuinely generic logic duplicated inside a specific DSL because it was not recognised as toolkit material.

Repository state after Milestone 8

langfac/
├── labeled.rkt, stats.rkt, robot.rkt, config-interp.rkt   Milestones 1-4
├── macros/
│   ├── my-or.rkt, my-let.rkt                                Milestones 5-6
│   └── ast-types.rkt                                          Milestone 8: define-ast-types
├── finance/
│   └── core.rkt                                                Milestone 7: account, rule
└── tests/                                                        8 files
$ raco test tests/
All tests passed.
$ git commit -am "milestones 5-8: macros, hygiene, syntax-parse, the finance DSL core, the toolkit"

Instalment 23 of the five-course curriculum. Next: Racket Milestones 9–12, where #lang finance becomes a real, running language, the robot DSL ships as a reusable interpreter library, the finance DSL gets compiled instead of interpreted, and a capstone game language uses the whole toolkit at once.

Continue