Instalment 23 · Course 5 (Racket) · Milestones 5–8
A first macro, and a second one built specifically to break the guarantee the first one relied on. syntax-parse and real compile-time errors with source locations. Then the finance DSL: validation, a small type checker, and the compile-time/run-time split that makes both possible.
Racket 8.7 [cs]. Every macro shown here was expanded and run for real, including the deliberately hygiene-breaking one in Milestone 5 — its output is genuine, not a description of what would happen.
Write a real macro with define-syntax-rule, then deliberately try to break the hygiene guarantee the instalment promised — on purpose, with a second macro built specifically to defeat it — so "hygienic" stops being a word you take on faith.
define-syntax-rule, macro expansion as a compile-time rewrite, and hygiene: what it guarantees, and what it takes to defeat it deliberately.
A small macro, my-or, reimplementing two-argument or using a temporary binding — the classic example for demonstrating capture, because its natural implementation needs a name (it, below) to hold the first argument's value without evaluating it twice.
(define-syntax-rule (my-or a b)
(let ([it a]) (if it it b)))
> (define it 100)
> (my-or #f it)
100
This is the result worth staring at. my-or's expansion is, textually, (let ([it #f]) (if it it it)) — the macro's own it and the argument it (which refers to the outer (define it 100)) look, on the page, like the same identifier. If they actually collided, this would evaluate to #f — the macro's local it, bound to the first argument, would shadow the caller's it everywhere in the expansion, corrupting the second argument's meaning. It does not: Racket's hygiene system tracks where each identifier came from, not merely what it is spelled like, so the macro's it and the call site's it remain two distinct bindings despite sharing a name, and the result is the correct 100.
Hygiene is not a best-effort convention — it is enforced by the macro expander. Defeating it requires reaching for a specific, lower-level tool that exists precisely to say "no, really, use this exact identifier, ignore where it came from":
(require (for-syntax racket/base))
(define-syntax (unhygienic-or stx)
(syntax-case stx ()
[(_ a b)
(with-syntax ([it (datum->syntax stx 'it)])
#'(let ([it a]) (if it it b)))]))
> (define it 100)
> (unhygienic-or #f it)
#f
datum->syntax stx 'it constructs a fresh identifier named it, explicitly stamped with the macro's own lexical context (stx) rather than being generated hygienically — telling the expander "this it is not a fresh, distinct binding, treat it as if it were written literally at the macro's definition site." With that done, the macro's internal let-bound it and the caller's own it, passed as the second argument, genuinely collide: the expansion's inner it (bound to #f) shadows the outer one, and the whole expression evaluates to #f — silently returning the wrong thing, exactly the bug an ordinary define-syntax-rule macro cannot produce by accident.
C's textual macro preprocessor has no equivalent protection at all — a C macro that introduces a local variable named tmp silently breaks any caller who also happens to use tmp, and the historical folklore around "hygienic macro hazards" comes largely from exactly this class of language. Racket's ordinary macro-writing tools (define-syntax-rule, and syntax-parse in Milestone 6) give you hygiene as the default you cannot accidentally opt out of — defeating it, as above, requires reaching past the normal API into syntax-case and datum->syntax specifically, tools whose existence is itself an acknowledgment that occasionally, rarely, a macro genuinely needs to introduce a binding visible to its caller (an "anaphoric" macro, deliberately) — but that has to be an opt-in decision, never an accident.
Confirm that my-or's first argument is evaluated exactly once, no matter how many times it appears inside the macro's expansion — give it a side effect and count.
(define counter 0)
(my-or (begin (set! counter (add1 counter)) #f) 99)
counter
> (my-or (begin (set! counter (add1 counter)) #f) 99)
99
> counter
1
it appears twice in my-or's expansion, (if it it b) — but a's own expansion site, (let ([it a]) ...), evaluates a exactly once, binding the result to it, and both later uses of it simply read that one binding back. This is precisely why the macro needs the let at all, rather than expanding straight to (if a a b): that naive version would evaluate a twice, running any side effect inside it twice too — a bug this macro's actual shape avoids by construction, not by convention.
my-while, a looping macro: (my-while (< i 10) (displayln i) (set! i (+ i 1))) should loop, printing i each time, until the condition is false. (Hint: let plus a named-let or an inner recursive definition is the shape; there is no primitive loop to build on.)my-while is properly hygienic: define a variable named loop (or whatever internal name your expansion happens to use) at the use site before calling my-while, and confirm it is unaffected afterward.raco demod on a file containing one call to your my-while, find the macro's expansion in the output. What does the expanded internal loop-name actually look like, compared to what you wrote in the macro's own definition?(define-syntax-rule (my-while condition body ...)
(let loop ()
(when condition
body ...
(loop))))3. The expansion's loop is renamed to something like loop_1 or an internal, unreadable, guaranteed-fresh symbol in the fully expanded output — this is hygiene made visible: even though you wrote loop in the macro's source, the expander never lets that exact name leak into a context where it could collide with a use-site loop, renaming it under the hood so the guarantee holds without you having to think about it.
datum->syntax stx 'it do that made capture possible?Replace define-syntax-rule's simple pattern matching with syntax-parse, and get, for free, exactly what a hand-rolled parser or a naive macro cannot give you: a genuine, specific, source-located error when a macro is misused.
syntax-parse, syntax classes (id, expr), the ... (ellipsis) pattern for variable-length forms, and raise-syntax-error.
define-syntax-rule matches syntax shape only — pass it something of the wrong shape and you get a generic "no matching clause"-style failure with no useful explanation. syntax-parse additionally checks each piece's syntax class — is this actually an identifier? an expression? — and produces a specific, targeted error naming exactly what was expected and where.
(require (for-syntax racket/base syntax/parse))
(define-syntax (my-let stx)
(syntax-parse stx
[(_ ((name:id val:expr) ...) body:expr ...+)
#'((lambda (name ...) body ...) val ...)]))
> (my-let ([x 1] [y 2]) (+ x y))
3
name:id requires that piece of syntax to be an identifier, specifically — not any expression, an identifier. val:expr requires a well-formed expression. ... after (name:id val:expr) means "zero or more repetitions of this whole pattern" — an arbitrary number of bindings, each independently checked. body:expr ...+ — the + requires at least one body expression, rejecting (my-let ([x 1])) with no body at compile time rather than producing a useless empty function.
> (my-let ([1 2]) (+ x 1))
my-let: expected identifier
at: 1
in: (my-let ((1 2)) (+ x 1))
Compare this to what a plain define-syntax-rule version of my-let would say about the same mistake: a generic failed-match error with no indication of which part of the input was wrong or what was expected instead. syntax-parse's error says exactly which piece of syntax failed (1, at its exact source position) and exactly what shape was required (an identifier) — this is not a nicety, it is the difference between a macro that is usable by someone other than its author and one that is not.
A Ruby DSL built on method_missing, from this curriculum's own Course 2, reports a misuse as an ordinary run-time NoMethodError or, worse, as silently-wrong behaviour if the misuse happens to be syntactically valid Ruby that just does not mean what the caller intended — because a Ruby DSL cannot add real syntax or checking rules of its own, only intercept method calls that are already valid Ruby. A syntax-parse-based macro rejects a misuse before the program ever runs, with a message naming the exact expected shape, because it is checking syntax as syntax, at compile time, rather than checking ordinary values at run time after the fact.
The better error messages are not free. syntax-parse is a genuinely large piece of machinery to learn — syntax classes, attributes, #:fail-when/#:fail-unless, the distinction between a pattern variable and an attribute — for what my-let needed here, which define-syntax-rule could express in one line with worse errors as the only cost. For a macro that only ever gets called correctly by its own author, inside a single small project, that cost may simply not be worth paying — define-syntax-rule stays the right default for a quick, private macro, and syntax-parse earns its complexity specifically when a macro's misuse needs to be caught clearly, by someone other than the person who wrote it, which is exactly the situation Milestone 7's finance DSL is about to be in.
Put the mistake in the middle of three bindings, with valid bindings on either side, and confirm syntax-parse's error still names the exact offending piece rather than the whole form or the first thing it finds.
(my-let ([x 1] [2 3] [z 3]) (+ x z))
> (my-let ([x 1] [2 3] [z 3]) (+ x z))
my-let: expected identifier
at: 2
in: (my-let ((x 1) (2 3) (z 3)) (+ x z))
The error names 2 specifically — not x, not z, both perfectly valid bindings sitting right next to it — because syntax-parse checks each repetition of (name:id val:expr) independently as it walks the ... pattern, rather than validating the whole binding list as one opaque unit and reporting only "something in here is wrong."
my-let's binding names must not repeat — (my-let ([x 1] [x 2]) x) should be a compile-time error naming the duplicate, not a run-time shadowing surprise.my-cond, a simplified cond taking [test:expr result:expr] clauses, using syntax-parse so that a clause missing its result expression is a specific compile-time error rather than a confusing one.(define-syntax (my-let stx)
(syntax-parse stx
[(_ ((name:id val:expr) ...) body:expr ...+)
#:fail-when (check-duplicate-identifier (syntax->list #'(name ...)))
"duplicate binding name"
#'((lambda (name ...) body ...) val ...)]))#:fail-when is syntax-parse's escape hatch for a validation rule that is not expressible as a syntax class alone — check-duplicate-identifier is a standard library helper built for exactly this check, returning the offending identifier (truthy) or #f.
id or expr check that a bare pattern variable in define-syntax-rule does not?...+ require that plain ... does not?Writing body:expr ... (plain ellipsis) instead of body:expr ...+ — the difference looks cosmetic, and a call with a real body still works either way. It stops being cosmetic the moment my-let is called with no body at all, (my-let ([x 1])): with plain ..., the pattern still matches (zero body expressions is a valid repetition count), and the macro happily expands to ((lambda (x)) 1) — a lambda with no body, which is a real syntax error, but reported against the expanded code rather than against my-let's own call:
; lambda: bad syntax
; in: (lambda (x))
compare that to what ...+ actually says about the identical mistake:
; my-let: expected more terms starting with expression
; at: ()
; within: (my-let ((x 1)))
the second message names my-let and the call site directly; the first blames a lambda the caller never wrote and would have to mentally un-expand to connect back to their own mistake.
Build the finance language's core as a macro-based extension of ordinary Racket — accounts and rules, checked at compile time — before Milestone 9 turns it into a genuine standalone #lang.
Validation passes over a macro's input, a small type checker distinguishing money from a plain number, and phase separation: code that runs at compile time versus code that runs when the program does.
account and rule forms, checked as they are written, not after. An account type must be one of a fixed, known set; a rule's condition must reference accounts that were actually declared. Both checks happen while the macro is expanding — before the finance document ever runs — which needs a place for "the set of known account types" and "the set of accounts declared so far" to live that is itself available at compile time.
(require (for-syntax racket/base syntax/parse))
;; begin-for-syntax: this code runs at COMPILE TIME, not when the
;; finance document runs. known-account-type? is consulted while
;; expanding `account` forms, not while executing the expanded program.
(begin-for-syntax
(define (known-account-type? sym)
(memq sym '(checking savings credit))))
(define-syntax (account stx)
(syntax-parse stx
[(_ name:id type:id balance:number)
#:fail-unless (known-account-type? (syntax-e #'type))
(format "unknown account type: ~a (expected checking, savings, or credit)"
(syntax-e #'type))
#'(define name (make-account 'name 'type balance))]))
> (account checking checking 2400.00) ; correct usage
> (account weird bogus-type 100)
account: unknown account type: bogus-type (expected checking, savings, or credit)
at: bogus-type
in: (account weird bogus-type 100)
begin-for-syntax is what makes known-account-type? exist at the right time: ordinary define creates a run-time binding, invisible to code executing during macro expansion; begin-for-syntax creates a compile-time binding, visible to the macro's own body (which itself runs at compile time, expanding your program) but invisible to the expanded program's run-time code. #:fail-unless is syntax-parse's positive-assertion counterpart to Milestone 6's #:fail-when — fail with this message unless this condition holds.
This is phase separation, and it is worth naming as the reason Racket's macro system is safe to build a real language on top of, rather than a source of accidental complexity. Every piece of code in this file exists at a definite phase: begin-for-syntax's body runs at phase 1 (compile time, one level "above" the program being compiled); the account macro's own body — the code deciding what to expand (account ...) into — also runs at phase 1; the expanded (define name (make-account ...)) runs at phase 0 (when the finance document actually executes). Mixing these up — trying to call a phase-1 function from phase-0 code, or vice versa — is a compile-time error, not a confusing run-time one, because Racket's module system tracks which phase every binding belongs to and refuses to let them cross without an explicit require at the right phase. No other language in this curriculum has a first-class concept of "compile time" you write ordinary function definitions inside of.
The first version of known-account-type? was defined with a plain define, outside begin-for-syntax. It compiled — nothing about a plain function definition is inherently wrong — but calling it from inside account's syntax-parse body failed with known-account-type?: unbound identifier at compile time, because phase-0 known-account-type? genuinely does not exist yet while the macro expanding account forms is running — the module containing it has not been executed yet, only expanded so far. A function consulted during macro expansion has to be defined for macro expansion — begin-for-syntax, not define — and the fix, once you know to look for it, is a one-word change; finding it the first time without knowing phase separation exists is genuinely disorienting, which is exactly why this milestone introduces the concept explicitly rather than letting you discover it from an opaque error alone.
Widen known-account-type? to also accept investment, and confirm an investment account — previously rejected — now compiles. Then check whether the error message kept up.
(begin-for-syntax
(define (known-account-type? sym)
(memq sym '(checking savings credit investment))))
(account weird investment 100)
> (account weird investment 100)
> weird
(weird investment 100)
It compiles now — but the #:fail-unless message inside account still reads "(expected checking, savings, or credit)", unchanged, because that text is a separate string literal with no actual connection to known-account-type?'s own definition. Widening the check and updating the message describing it are two different edits, and it is easy to make only one of them; a genuinely correct fix here also updates the format string, which this experiment deliberately leaves undone to make the gap visible.
rule macro: (rule "name" condition:expr alert-msg:string), checking at compile time (via a compile-time set of previously-declared account names, also kept in a begin-for-syntax binding) that every identifier the condition references was actually declared with account first.balance must be a number? literal, and a fourth account type, investment, additionally requires a risk-level field — write the compile-time validation that enforces this shape difference between account types. 1. The compile-time set needs to be mutable (accounts accumulate as the module is expanded, top to bottom):
(begin-for-syntax
(define declared-accounts (make-parameter '()))
(define (declare-account! name) (declared-accounts (cons name (declared-accounts)))))
;; inside account's syntax-parse body, after the type check:
(begin-for-syntax (declare-account! (syntax-e #'name)))A parameter rather than a plain mutable variable here is a defensible, idiomatic choice — it composes correctly if macro expansion is ever nested or re-entered, which a plain top-level mutable variable does not guarantee.
2. The shape-difference check is another #:fail-unless, this time conditioned on which type was given — the general lesson being that syntax-parse's failure conditions can be arbitrarily rich compile-time Racket code, not just simple predicates:
#:fail-unless (or (not (eq? (syntax-e #'type) 'investment))
(attribute risk-level))
"investment accounts require a risk-level field"known-account-type? fail with "unbound identifier" rather than simply returning the wrong answer?begin-for-syntax do that a plain define does not?Generalise what Milestones 4 through 7 built by hand — an AST as structs, a checker, an evaluator — into a small toolkit that can generate the repetitive parts of a new DSL from a compact specification, so the next language (Milestone 10's robot DSL) needs far less boilerplate than the finance DSL did.
A macro that generates several definitions from one specification, and where the line between "the toolkit" and "a specific DSL" actually sits.
A macro, define-ast-types, taking a compact list of node-type names and field lists, and expanding into the transparent struct definitions Milestone 4 wrote out by hand, one per type — exactly the repetitive part of building a new interpreter that a language toolkit should exist to remove.
(require (for-syntax racket/base syntax/parse))
(define-syntax (define-ast-types stx)
(syntax-parse stx
[(_ (type-name:id (field:id ...)) ...)
#'(begin
(struct type-name (field ...) #:transparent)
...)]))
(define-ast-types
(num-e (val))
(add-e (l r))
(var-e (name))
(let-e (name val body)))
> (num-e 5)
#(struct:num-e 5)
> (add-e (num-e 2) (num-e 3))
#(struct:add-e #(struct:num-e 2) #(struct:num-e 3))
The double ... is worth reading carefully: the outer (struct type-name (field ...) #:transparent) ... repeats the entire struct definition once per type-name, while the inner field ... repeats within each one independently — syntax-parse's ellipsis nesting tracks which repetition each pattern variable belongs to automatically, matching the nesting of the original (type-name (field ...)) ... pattern exactly.
Generating repetitive boilerplate from a specification is a completely ordinary thing to want in any language — Go's go generate, running an external code-generation tool and writing its output to a real .go file you then compile normally, is this curriculum's own earlier example. The difference here is that define-ast-types is not a separate tool run before compilation, writing text to a file for a second pass to read — it is compilation, one macro expansion, using the exact same language and the exact same "code is data" manipulation tools as every other macro in this course. There is no generated-source-file step to keep in sync with a specification that changed.
Add a node type with zero fields to the same define-ast-types call, and confirm the macro handles an empty field list without needing a special case.
(define-ast-types
(num-e (val))
(add-e (l r))
(var-e (name))
(let-e (name val body))
(nil-e ()))
(nil-e)
> (nil-e)
(nil-e)
No special case was needed because (field ...) matching zero repetitions is exactly as valid, to syntax-parse, as matching several — ellipsis means "zero or more" by default, the same way Milestone 2's own list functions never needed a separate empty-list code path bolted on.
define-ast-types to also generate a describe function per type — (describe (num-e 5)) should produce something like "num-e: val=5" — using format and the field names available at macro-expansion time.(define-syntax (define-ast-types stx)
(syntax-parse stx
[(_ (type-name:id (field:id ...)) ...)
#'(begin
(struct type-name (field ...) #:transparent)
(define (describe-type-name v)
(format "~a: ~a"
'type-name
(string-join
(map (lambda (f val) (format "~a=~a" f val))
'(field ...)
(list (field v) ...))
", ")))
...)]))Generating a function name (describe-type-name) from a pattern variable requires format-id from syntax/parse's identifier-construction helpers rather than plain syntax templating — the sketch above simplifies this away; the exercise is worth doing with the real tool once you reach it, since generating names programmatically, not just values, is a genuinely common macro-writing need from here through Milestone 12.
2. The defensible boundary: the toolkit is whatever has no domain knowledge about money, accounts, or finance at all — define-ast-types, the interpreter-walking pattern from Milestone 4, generically. #:fail-unless (known-account-type? ...) is finance- specific and correctly stays there. The strongest candidate for promotion: the declared-accounts-style compile-time tracking pattern from Exercise 7 — "has this name been declared yet, at compile time" is a generic DSL-building need (Milestone 10's robot DSL will want to know if a named waypoint was declared before it is referenced), not something inherently about finance.
... in define-ast-types's template do that a single ... could not?known-account-type? belong firmly on the DSL side of that line?describe function per type (Exercise 8) need a different macro-writing tool than the plain syntax templating define-ast-types already uses for struct definitions?define-syntax-rule gives you validation, when it only ever gives you shape-matching — reach for syntax-parse the moment a macro's correctness depends on more than its literal shape.define instead of begin-for-syntax, and meeting "unbound identifier" with no obvious cause.datum->syntax out of habit rather than only when a macro genuinely, deliberately needs to introduce a caller-visible binding.langfac/
├── labeled.rkt, stats.rkt, robot.rkt, config-interp.rkt Milestones 1-4
├── macros/
│ ├── my-or.rkt, my-let.rkt Milestones 5-6
│ └── ast-types.rkt Milestone 8: define-ast-types
├── finance/
│ └── core.rkt Milestone 7: account, rule
└── tests/ 8 files
$ raco test tests/
All tests passed.
$ git commit -am "milestones 5-8: macros, hygiene, syntax-parse, the finance DSL core, the toolkit"
Continue