InstalmentMilestones 5–8

Instalment 7 · Course 2 (Ruby) · Milestones 1–4

A gem, then a block, then a language, then a data structure

We build the DSL surface twice: once with an explicit builder so you can see what it costs, then with instance_eval so you can see what it buys. Then we throw away the idea that building should run anything.

Verification note

All code was run on Ruby 3.2.3 and every output block below is copied from that run. The final suite is 20 tests and 50 assertions, all passing. The failure demonstrations in Milestone 3 are real program output, not illustrations.

Milestone 1The gem, a Step, and a Registry

Goal

The Mewlang cat, typing on a laptopA working gem skeleton with three ideas in it: a Step that describes work, a Registry that knows how to perform it, and an error hierarchy that a user can rescue. No DSL yet.

Concepts

Classes and attr_reader, freezing for immutability, value equality with ==/eql?/hash, duck typing with respond_to?, Hash#fetch with a block, error class hierarchies, and minitest.

Design

The central separation, which the whole course rests on: a Step says what, the Registry says how. A step named :fetch with options {from: "arxiv"} is a description. Whether that means an HTTP call, a fixture file, or a no-op in a dry run is the registry's business. Once those are separate, testing a pipeline stops requiring a network, and users can add verbs without touching your code.

The second decision is that a Step is frozen on construction. Ruby lets you mutate almost anything, and a description that changes between validation and execution is a bug you cannot reproduce. Freezing is how you opt out.

Implementation

lib/automation/errors.rb

# frozen_string_literal: true

module Automation
  # Every error this gem raises inherits from Error, so a user can write
  # `rescue Automation::Error` and catch everything we throw without
  # catching everything in the world.
  class Error < StandardError; end

  # Raised when a pipeline references a step nobody registered.
  class UnknownStep < Error
    attr_reader :name, :known

    def initialize(name, known, location: nil)
      @name = name
      @known = known
      message = +"unknown step #{name.inspect}"
      message << " at #{location}" if location
      message << "; known steps: #{known.sort.join(', ')}" unless known.empty?
      message << "; no steps are registered" if known.empty?
      super(message)
    end
  end

  # Raised when a step implementation fails. Keeps the original as #cause.
  class StepFailed < Error
    attr_reader :step

    def initialize(step, cause)
      @step = step
      super("step #{step.name} failed: #{cause.message} (#{cause.class})")
    end
  end
end

Explanation

lib/automation/step.rb

module Automation
  # Step is one instruction in a pipeline: a name, positional arguments and
  # options. It describes work; it does not perform it.
  class Step
    attr_reader :name, :args, :options

    def initialize(name, *args, **options)
      @name = name.to_sym
      @args = args.freeze
      @options = options.freeze
      freeze
    end

    def to_s
      parts = args.map(&:inspect) + options.map { |k, v| "#{k}: #{v.inspect}" }
      "#{name}(#{parts.join(', ')})"
    end

    def ==(other)
      other.is_a?(Step) && name == other.name &&
        args == other.args && options == other.options
    end
    alias eql? ==

    def hash = [self.class, name, args, options].hash
  end
end

Explanation

lib/automation/registry.rb

module Automation
  # Registry maps step names to implementations. An implementation is
  # anything that responds to #call: a lambda, a method object, or an
  # instance of a class you wrote.
  class Registry
    def initialize
      @steps = {}
    end

    def register(name, callable = nil, &block)
      impl = callable || block
      raise ArgumentError, "register(#{name.inspect}) needs a callable or a block" if impl.nil?
      unless impl.respond_to?(:call)
        raise ArgumentError, "step #{name.inspect} must respond to #call, got #{impl.class}"
      end

      @steps[name.to_sym] = impl
      self
    end

    def fetch(name, location: nil)
      @steps.fetch(name.to_sym) { raise UnknownStep.new(name, known, location: location) }
    end

    def registered?(name) = @steps.key?(name.to_sym)
    def known = @steps.keys
  end
end

Explanation

The tests

class StepTest < Minitest::Test
  def setup
    @step = Automation::Step.new("fetch", "papers", from: "arxiv")
  end

  def test_name_is_always_a_symbol
    assert_equal :fetch, @step.name
    assert_equal :fetch, Automation::Step.new(:fetch).name
  end

  def test_is_frozen_all_the_way_down
    assert @step.frozen?
    assert @step.options.frozen?
    assert_raises(FrozenError) { @step.options[:from] = "elsewhere" }
  end

  def test_value_equality_and_hashing
    twin = Automation::Step.new(:fetch, "papers", from: "arxiv")
    assert_equal twin, @step
    assert_equal 1, [@step, twin].uniq.size
    refute_equal Automation::Step.new(:fetch, "papers", from: "other"), @step
  end

  def test_to_s_reads_like_the_dsl
    assert_equal 'fetch("papers", from: "arxiv")', @step.to_s
  end
end
$ ruby -Ilib -Itest test/test_step.rb
4 runs, 9 assertions, 0 failures, 0 errors, 0 skips

Note [@step, twin].uniq.size: that assertion is the one that fails if you define == and forget hash. Tests that pin the consequences of a protocol, rather than the protocol itself, catch more.

Exercise 1

Make Registry enumerable and inspectable:

  • Add each yielding [name, implementation] pairs, and include Enumerable. Then confirm that registry.map, registry.count, registry.sort_by and registry.group_by all work without writing them.
  • Add Registry#describe(name) returning a human-readable signature of a registered step, derived from the implementation itself, so describe(:fetch) prints which options it requires and which it accepts.

Hint for the second part: every callable in Ruby answers parameters. Try ->(a, b:, c: 1) {}.parameters in irb before writing any code.

Solution 1 — open after trying
class Registry
  include Enumerable

  def each(&block)
    return to_enum(:each) unless block_given?

    @steps.each(&block)
    self
  end

  def describe(name)
    impl = fetch(name)
    required = impl.parameters.select { |type, _| type == :keyreq }.map(&:last)
    optional = impl.parameters.select { |type, _| type == :key }.map(&:last)
    rest     = impl.parameters.any? { |type, _| type == :keyrest }

    parts = []
    parts << "requires: #{required.join(', ')}" unless required.empty?
    parts << "accepts: #{optional.join(', ')}"  unless optional.empty?
    parts << "accepts any options"              if rest
    "#{name}(#{parts.join('; ')})"
  end
end
registry.register(:fetch) { |_input, what, from:, since: "1d"| }
registry.describe(:fetch)
# => "fetch(requires: from; accepts: since)"

registry.count                      # => 1
registry.map { |name, _| name }     # => [:fetch]
registry.sort_by { |name, _| name } # => [[:fetch, #<Proc...>]]

Two things worth keeping from this.

Enumerable is free power. Define each, include the module, and about sixty methods appear. This is Ruby's answer to interfaces: instead of implementing a big contract, you implement one method and a module implements the contract in terms of it. The Go equivalent would be writing Map, Filter, GroupBy and the rest yourself for every type.

parameters is the reflection that Milestone 4 turns into a validator. It reports an array of [kind, name] pairs, where the kinds are :req, :opt, :rest, :key, :keyreq, :keyrest and :block. A language that can ask a function what arguments it wants can check a call before making it, which is how we will get compile-time-ish errors out of a dynamic language.

Experiment

Remove freeze from Step#initialize and run the suite: one test fails with a clear message. Now remove only the .freeze on @options and run again: the object is frozen but its Hash is not, so @step.options[:from] = "x" succeeds. That is shallow freezing, and it is the reason Data in Milestone 4 is worth having, since it freezes what it holds.

Common mistakes in Milestone 1
  • The Mewlang cat, giving an annoyed side-eye from aboveDefining == without hash and eql?. Everything looks fine until uniq, group_by or a Hash key behaves strangely.
  • Forgetting super in a custom exception's initialize. The message silently becomes the class name.
  • Mutating a frozen string literal in a file with the magic comment. Use +"..." or build with interpolation.
  • Requiring a base class instead of respond_to?(:call). It makes your gem hostile to lambdas and to anyone else's objects.
  • Using [] where you need fetch. A nil implementation produces NoMethodError: undefined method 'call' for nil three frames away from the cause.

Checkpoint

  1. Why does Step freeze both itself and its options Hash?
  2. What breaks if you define == but not hash?
  3. Why does Registry#register check respond_to?(:call) rather than the class?
  4. What does Hash#fetch with a block do that Hash#[] does not?
  5. Why must every error in a gem inherit from one base class?
Why are we using this language here?

The Mewlang cat, facing forwardNothing in Milestone 1 needs Ruby specifically. Registry#register accepting "anything that responds to #call" is duck typing, and it is genuinely convenient: a lambda, a Method object and a plain object all work with zero adapter code. But the same design is one interface declaration away in Go, and a Python Protocol gets you the same check with static tooling behind it. What this milestone actually shows off is smaller and more concrete: Data.define-adjacent value semantics done by hand (==, eql?, hash, freeze) are four separate decisions in Ruby that a case class or a Go struct with a generated comparator would bundle for you automatically.

The honest cost showed up immediately: forgetting hash while defining == breaks uniq and Hash lookups with no warning at write time, only at use time, and only if a test happens to exercise it. Milestone 4's Data.define makes this whole category of mistake impossible by generating all three together — which is itself an admission that hand-rolled value equality in Ruby is a trap worth avoiding once a better tool exists.

Milestone 2Blocks: the DSL with its receiver showing

Goal

Automation.pipeline("research") do |p| ... end builds a runnable pipeline. The builder is an explicit block parameter, which is deliberately one step short of the DSL we want, because the difference between this and Milestone 3 is the entire lesson.

Concepts

yield and block parameters, reduce as an interpreter, exception wrapping with automatic cause, and the difference between a build-time and a run-time error.

Design

The Mewlang cat, thinking with a paw to its chinThree objects, each with one job:

  Automation.pipeline(name) { |p| ... }
        │
        ├─ creates a Builder, hands it to the block
        │
        ├─ Builder#step collects Step descriptions
        │
        └─ Builder#to_pipeline produces a frozen Pipeline

  Pipeline#run  →  reduce over the steps, looking each one up

The interpreter is one line, and it is worth seeing before it gets dressed up: a pipeline is a fold. Each step takes the previous step's output and returns the next input. That single decision (values flow through, nothing is shared) is what makes steps composable and testable, and it is the same reasoning that made Go's Action a pointer-free value.

Implementation

module Automation
  class Pipeline
    attr_reader :name, :steps

    def initialize(name, steps, registry: Automation.registry)
      @name = name.to_sym
      @steps = steps.freeze
      @registry = registry
      freeze
    end

    def run(input = nil)
      @steps.reduce(input) { |acc, step| perform(step, acc) }
    end

    private

    def perform(step, input)
      impl = @registry.fetch(step.name)
      impl.call(input, *step.args, **step.options)
    rescue StandardError => e
      raise e if e.is_a?(Error)
      raise StepFailed.new(step, e)
    end
  end

  class Builder
    def initialize(name)
      @name = name
      @steps = []
    end

    def step(name, *args, **options)
      @steps << Step.new(name, *args, **options)
      self
    end

    def to_pipeline(registry:) = Pipeline.new(@name, @steps, registry: registry)
  end

  def self.pipeline(name, registry: self.registry)
    builder = Builder.new(name)
    yield builder
    builder.to_pipeline(registry: registry)
  end
end

Explanation

Running it

Automation.register(:fetch) do |_input, source, from:|
  puts "  fetching #{source} from #{from}"
  ["Attention Is All You Need (AI)", "Cats sleep a lot (biology)", "Scaling laws (AI)"]
end

Automation.register(:filter)    { |items, topic:| items.select { |i| i.include?(topic) } }
Automation.register(:summarize) { |items, max_words:| items.map { |i| i.split.first(max_words).join(" ") + "..." } }

research = Automation.pipeline("research") do |p|
  p.step :fetch, "papers", from: "arxiv"
  p.step :filter, topic: "AI"
  p.step :summarize, max_words: 3
end

puts research
p research.steps.map(&:to_s)
p research.run
research (3 steps)
["fetch(\"papers\", from: \"arxiv\")", "filter(topic: \"AI\")", "summarize(max_words: 3)"]
  fetching papers from arxiv
["Attention Is All...", "Scaling laws (AI)..."]

Two failure paths, also real output:

--- an unregistered step fails at run time, not build time ---
built fine: broken (1 steps)
run failed: unknown step :summarise; known steps: fetch, filter, summarize

--- a failing step is wrapped, with the cause preserved ---
Automation::StepFailed: step explode failed: divided by 0 (ZeroDivisionError)
cause: ZeroDivisionError

Read the first one carefully, because it is the flaw that drives Milestone 4. A pipeline containing a typo builds successfully and fails only when that step executes. If :summarise were the fourth step of an hour-long job, you would find out in an hour. Go would have caught this at compile time. We will catch it with a validator, which is Ruby's equivalent: not free, but available.

Ordinary configuration vs an executable builder
YAML                              Ruby builder
────                              ────────────
steps:                            Automation.pipeline("research") do |p|
  - name: fetch                     p.step :fetch, "papers", from: "arxiv"
    from: arxiv                     p.step :filter, topic: "AI"
  - name: filter                  end
    topic: AI

static, safe to parse            executable, arbitrary code
needs a schema to validate       can validate with reflection
no loops, no variables           3.times { |i| p.step :fetch, page: i }
no editor support for values     completion, refactoring, debugger
errors: "line 7: bad key"        errors: a Ruby backtrace

The builder already wins on power and loses on safety. Neither difference has anything to do with how it looks; those come next.

Exercise 2

Give Pipeline composition. Implement Pipeline#| (the pipe operator) so that fetching | processing returns a new pipeline whose steps are the concatenation of both, with a name derived from the two. Then implement Pipeline#+ as an alias, and Pipeline#repeat(n) returning a pipeline whose steps are repeated n times.

Constraints: the originals must not change; both pipelines must use the same registry or you should raise a clear error; and (a | b) | c must equal a | (b | c) in steps. Write tests for all three.

Solution 2 — open after trying
class Pipeline
  def |(other)
    unless other.is_a?(Pipeline)
      raise ArgumentError, "cannot compose a Pipeline with #{other.class}"
    end
    if registry_of(other) != @registry
      raise Error, "cannot compose pipelines using different registries"
    end

    Pipeline.new(:"#{name}_#{other.name}", steps + other.steps, registry: @registry)
  end
  alias + |

  def repeat(times)
    raise ArgumentError, "times must be positive" unless times.positive?

    Pipeline.new(:"#{name}_x#{times}", steps * times, registry: @registry)
  end

  protected

  # protected, not private: readable by other Pipelines, not by outsiders.
  def registry_of(other) = other.instance_variable_get(:@registry)
end
combined = fetching | processing
combined.steps.map(&:name)   # => [:fetch, :filter, :summarize]
fetching.steps.map(&:name)   # => [:fetch]   (unchanged)

Three Ruby-specific points hide in those fifteen lines.

Operators are methods. def |(other) defines what a | b means. The list of overloadable operators is long (+ - * / % ** == <=> [] []= << & | ^ ! =~), and the discipline is to define one only when the meaning is obvious to a reader who has not seen your code. "Pipe two pipelines together" qualifies; almost nothing else in this project would.

alias + | makes both spellings the same method. Aliases are resolved at definition time, so redefining | later does not change +, which surprises people.

protected is the rarely-used third visibility. A protected method can be called with an explicit receiver, but only from inside the same class. That is exactly what comparing two pipelines needs, and it is the one situation where protected is the right answer rather than a confusion. (Using instance_variable_get at all is a little impolite; a cleaner design exposes attr_reader :registry and accepts that users can see it.)

steps * times works because Array#* with an integer repeats the array, which is a nice example of Ruby's core library being unusually generous.

Experiment

Register a step that returns nil and put it in the middle of a pipeline. Then watch the next step receive nil and fail with something unhelpful like undefined method 'select' for nil. The fold is unforgiving: every step must return something the next one can use. This is the argument for the Context object that Milestone 5 introduces, where the value flowing through is a structured thing with a payload and metadata rather than a bare return value.

Common mistakes in Milestone 2
  • Forgetting the block parameter: Automation.pipeline("x") { step :fetch } raises NoMethodError because step is not defined on the top-level object. That error is exactly what Milestone 3 removes.
  • reduce without returning the accumulator. A step whose last expression is puts returns nil and poisons the rest of the fold.
  • Rescuing StandardError and swallowing your own errors, producing double-wrapped messages.
  • Calling impl.call(input, step.options) without the double splat, which passes the Hash as a positional argument. Since Ruby 3.0 these are strictly different and the error message (wrong number of arguments) does not say why.
  • Freezing the pipeline but not the steps array, so a caller can append to it afterwards.

Checkpoint

  1. What does yield builder do, and what happens to the block's return value?
  2. Explain impl.call(input, *step.args, **step.options) in terms of how Step collected them.
  3. Why does perform re-raise errors that are already Automation::Error?
  4. Where does e.cause come from, and what is the Go equivalent?
  5. Why is a typo in a step name not caught until run time?
Why are we using this language here?

yield builder is the whole trick in this milestone, and it is smaller than it looks: a block that receives an object and calls methods on it is one line of Ruby, and the same builder pattern exists in Go (a function taking a *Builder), in Java (a fluent builder class) and everywhere else. Nothing here required a dynamic language yet — Milestone 3 is where that changes.

What this milestone is honest about instead is the price of choosing "return a result" over "raise an exception". A Pipeline#run that returns bare values, with no result wrapper, makes a failed step indistinguishable from a step that legitimately returned nil. We pay for that simplicity later, in Milestone 5, with a whole RunResult type built specifically to stop conflating the two. A language feature this milestone leans on without remarking on it: rescue StandardError => e inside a bare method body works because a method definition is an implicit begin block — convenient here, and one more thing a reader coming from a language with explicit try blocks has to learn once and then never think about again.

Milestone 3instance_eval: what it buys, and what it costs

Goal

Remove the receiver, so the block reads as a language:

Automation.define_pipeline("research") do
  fetch "papers", from: "arxiv"
  filter topic: "AI"
  summarize max_words: 3
end

Then find the three ways this breaks, watch each one break for real, and fix them.

Concepts

instance_eval, self and the implicit receiver, method_missing with respond_to_missing?, BasicObject, and delegating to the caller via block.binding.receiver.

Design

The mechanism is one line: builder.instance_eval(&block) runs the block with self set to the builder, so fetch "papers" is a message to the builder. The builder has no fetch method, so method_missing catches it and records a step.

That is the whole of the good news. The bad news has three parts, and every Ruby DSL author meets all three:

  1. Your caller's methods vanish. self is now the builder, so a helper method on the object where the block was written is no longer reachable.
  2. Object's own methods shadow your verbs. A builder is an Object, so it already responds to format, print, select, method, hash, display, test and about fifty others. method_missing never fires for those, so a step with a colliding name silently does something else.
  3. Typos become steps. If method_missing accepts everything, a misspelling is recorded as a perfectly valid step that fails much later.

Let us watch all three happen.

The naive version, and its failures

class NaiveDSL
  def initialize(name)
    @name = name
    @steps = []
  end

  def method_missing(verb, *args, **options, &_block)
    @steps << Step.new(verb, *args, **options)
    self
  end

  def respond_to_missing?(_verb, _include_private = false) = true

  def steps = @steps
end

Failure 1: the caller's helper is captured as a step

class Report
  def initialize(topic) = @topic = topic
  def default_topic = @topic          # a helper on the OUTER object

  def naive
    dsl = Automation::NaiveDSL.new("r")
    dsl.instance_eval do
      filter topic: default_topic     # self is the builder now
    end
    dsl.steps
  end
end

p Report.new("AI").naive.map(&:to_s)
["default_topic()", "filter(topic: #<Automation::NaiveDSL:0x00007f754b7183c0 @name=\"r\",
 @steps=[#<Automation::Step:0x00007f754b718118 @name=:default_topic, ...>]>)"]

Look at what happened. default_topic was not a NoMethodError; it was captured by method_missing and became a step called default_topic. Then, because method_missing returns self, its return value was the builder, which got passed as the topic: option of the next step. The user asked for one step and got two, one of which contains a builder as data.

The Mewlang cat, giving an unimpressed side-eyeThis is the worst kind of bug: no exception, no warning, a plausible-looking result, and a failure that surfaces somewhere else entirely. It is the price of method_missing accepting everything.

Failure 2: a step name that collides with an Object method

Automation.register(:format) { |items| items.map(&:upcase) }

naive = Automation::NaiveDSL.new("collide")
naive.instance_eval { format "%s", "x" }    # Kernel#format wins, silently
p naive.steps.map(&:to_s)
[]

No step at all. Kernel#format exists on every object, so the method was found and method_missing never ran. The user's format step vanished without a trace. Every name on Object is a landmine: p, print, puts, select, test, method, display, hash, send, class, freeze, trust, then, tap.

The version we keep

  # DSLBuilder is the version we keep. Two changes from NaiveDSL:
  #
  #   1. It inherits from BasicObject, which has almost no methods, so a
  #      step named `format` or `print` is not silently swallowed by
  #      Object's own methods.
  #   2. It remembers the object the block was written in, and forwards
  #      anything that is not a registered step back to it, so helper
  #      methods from the caller still work inside the block.
  class DSLBuilder < BasicObject
    def initialize(name, outer, registry)
      @name = name
      @outer = outer
      @registry = registry
      @steps = []
    end

    def method_missing(verb, *args, **options, &block)
      if !@registry.registered?(verb) && @outer.respond_to?(verb, true)
        return @outer.__send__(verb, *args, **options, &block)
      end

      @steps << ::Automation::Step.new(verb, *args, **options)
      self
    end

    def respond_to_missing?(verb, include_private = false)
      @registry.registered?(verb) || @outer.respond_to?(verb, include_private)
    end

    # Deliberately ugly names: anything readable might collide with a step.
    def __steps__ = @steps
  end

  def self.define_pipeline(name, registry: self.registry, &block)
    raise ArgumentError, "define_pipeline needs a block" unless block

    outer = block.binding.receiver
    builder = DSLBuilder.new(name, outer, registry)
    builder.instance_eval(&block)
    Pipeline.new(name, builder.__steps__, registry: registry)
  end

Explanation

Both fixes, working

--- 1. bare verbs, no receiver ---
["fetch(\"papers\", from: \"arxiv\")", "filter(topic: \"arxiv\")", "summarize(max_words: 2)"]
["papers from"]

--- 2. local variables still work (the block is a closure) ---
["fetch(\"papers\", from: \"arxiv:cs.AI\")", "summarize(max_words: 2)"]

--- 3. the caller's own methods: broken, then fixed by delegation ---
["default_topic()", "filter(topic: #<Automation::NaiveDSL...>)"]      # naive
["filter(topic: \"AI\")"]                                              # delegating

--- 4. a step whose name collides with an Object method ---
[]                                                                    # naive
["format()"]                                                          # BasicObject

Section 2 of that output is worth a moment: instance_eval changes self, but the block is still a closure, so local variables from the enclosing scope (source, limit) remain visible. Methods and locals behave differently, and knowing which is which is most of understanding Ruby scope.

And because the block is ordinary Ruby, control flow comes free:

pipeline = Automation.define("x") do
  puts "  (puts works inside the block: forwarded to the outer self)"
  3.times { |i| summarize index: i }
end

p pipeline.step_names
p pipeline.steps.map { |s| s.options[:index] }
  (puts works inside the block: forwarded to the outer self)
[:summarize, :summarize, :summarize]
[0, 1, 2]

Three steps generated by a loop. That is the whole argument for an executable DSL over a data format, in four lines: no for_each: key to invent, no template language, no escaping rules. Your users already know how to write a loop.

The three ways to write a Ruby DSL
Explicit builder
do |p| p.step ... end
instance_eval
do fetch ... end
Hybrid
do |p| ... end with both
Reads like a languageNoYesPartly
Caller's methods workYes, alwaysOnly with delegationYes
Name collisionsImpossibleReal; needs BasicObjectImpossible
TyposNoMethodError on the builderBecome steps unless checkedNoMethodError
Editor supportSomeNoneSome
Used byMany librariesRSpec, Rake, Sinatra, GemfileRails routing, some gems

The hybrid deserves a mention: instance_exec(builder, &block) sets self and passes the builder, so users who want the explicit form can have it and everyone else can omit it. If your DSL is for a library other people extend, this is often the kindest choice. We take the pure instance_eval route because the course is about seeing the technique clearly.

Is block.binding.receiver too clever?

A fair objection. It reaches into the caller's environment without being asked, which is exactly the sort of thing that makes Ruby codebases hard to reason about. The conservative alternative is to require the caller to hand you the context: Automation.define("x", context: self) do ... end. It is uglier and it is honest.

My judgement is that the binding trick earns its keep here because the alternative is the failure in Section 3, which is silent and severe. But notice what it is doing: your DSL now behaves differently depending on where the block was written, which is genuinely harder to explain than "steps are methods on a builder". Every metaprogramming decision in this course has this shape, and the discipline is to ask "what would I have to explain to a new maintainer?" before reaching for the clever thing.

Exercise 3

Add nesting. Support a group verb that takes a name and a block, so a pipeline can be organised:

Automation.define_pipeline("research") do
  fetch "papers", from: "arxiv"

  group "cleaning" do
    deduplicate
    normalize case: :lower
  end

  summarize max_words: 200
end

Requirements: groups may nest to any depth; pipeline.steps should still be able to produce a flat list for execution; the group name must be recorded so error messages can say research/cleaning/normalize; and delegation to the caller must still work inside a nested block. Write a test with two levels of nesting.

Hint: what object should the inner block be evaluated against, and how does it get the same outer?

Solution 3 — open after trying
class DSLBuilder < BasicObject
  def initialize(name, outer, registry, path = [])
    @name = name
    @outer = outer
    @registry = registry
    @path = path
    @steps = []
  end

  def group(group_name, &block)
    # A nested builder: same outer, same registry, deeper path.
    nested = DSLBuilder.new(group_name, @outer, @registry, @path + [group_name.to_s])
    nested.instance_eval(&block)
    @steps.concat(nested.__steps__)
    self
  end

  def method_missing(verb, *args, **options, &block)
    if !@registry.registered?(verb) && @outer.respond_to?(verb, true)
      return @outer.__send__(verb, *args, **options, &block)
    end

    @steps << ::Automation::Step.new(
      verb, *args, **options.merge(__path__: (@path + [verb.to_s]).join("/"))
    )
    self
  end
end

The essential move is that group creates another builder of the same class and evaluates the inner block against it, then merges the results. Nesting an interpreter is almost always "make another one of yourself with different context", and you will do the identical thing in Racket in Course 5 and in the Perl parser in Course 3.

Two design choices worth arguing about.

Flattening in group versus keeping a tree. This solution flattens immediately, which keeps run unchanged and loses structure. Keeping a tree (a GroupNode whose children are steps or groups) preserves the shape for visualisation and lets you later say "retry this whole group", at the cost of a runner that must walk recursively. Milestone 4's AST is the right place to make that choice, and for a real tool I would keep the tree.

Smuggling __path__ into the options Hash is a hack: it pollutes the data the step receives, and a step with **opts will see it. The clean version puts the path on the node itself, which is what the AST does in the next milestone. Recognising that a piece of metadata does not belong in the payload is the thought that produces Milestone 4.

Experiment

Make DSLBuilder inherit from Object instead of BasicObject and register a step called :p or :test. Then try to use it. Then change it back and note the :: prefixes you suddenly need. Finally, run BasicObject.instance_methods.sort and Object.instance_methods.size in irb and compare: 8 methods against 58 or so. That difference is the entire cost and benefit.

Common mistakes in Milestone 3
  • method_missing without respond_to_missing?. Your object lies about itself, and any library that checks respond_to? before calling will skip you.
  • method_missing that never calls super in a context where typos matter. In a DSL builder, accepting everything is a deliberate choice that must be paired with validation (Milestone 4), not an accident.
  • Forgetting :: on constants inside a BasicObject. The error, uninitialized constant Automation::DSLBuilder::Step, is a confusing way of saying "constant lookup does not work the way you assumed".
  • Assuming instance_eval hides local variables. It does not; blocks are closures. Only self changes.
  • Using instance_eval where a plain block parameter would do. If your DSL has three verbs and is used once per project, the explicit builder is better code. Reach for the trick when the file will be read a hundred times.
  • Naming plumbing methods readably. steps, name and run on a DSL builder are names your users will want for their own verbs.

Checkpoint

  1. What exactly does instance_eval change, and what does it leave alone?
  2. Why did default_topic become a step in the naive version, and why is that worse than an exception?
  3. Why did the format step disappear entirely, and what fixes it?
  4. What is a Binding, and what does block.binding.receiver give you?
  5. Why does the registry check come first in method_missing?
  6. Give one situation where the explicit builder is the better design.

Milestone 4Stop executing. Build a data structure.

Goal

The DSL produces an immutable AST that records where every step was written. A validator checks the whole pipeline against the registry before anything runs, reporting problems with file and line. A pipeline can be converted to plain data and back, which gives us a safe mode for untrusted input.

Concepts

Data.define for AST nodes, caller_locations for source tracking, reflection on parameters to check arguments, immutable transformation, and serialisation as a security boundary.

Design

Milestone 2's Pipeline mixes three responsibilities: it describes the steps, it holds the registry, and it runs. Separating them gives us four things we cannot otherwise have:

CapabilityNeeds
Dry run: show what would happenA description that can be walked without executing
Validation with file:lineSource locations captured at build time
Self-modification (Milestone 11)A value you can transform into a new value
A safe mode for untrusted pipelinesA representation that is pure data, with no code in it

So the AST becomes the centre of the system, and everything else is a function over it:

   block ──► ASTBuilder ──► PipelineNode ──┬──► Validator  ──► problems
                            (frozen Data)  ├──► Runner     ──► results
                                           ├──► to_h       ──► plain data
                                           └──► transforms ──► a NEW PipelineNode
   data  ──► from_h ────────►┘  (no code executed, ever)

Implementation

lib/automation/ast.rb

module Automation
  # The AST. Every node is a Data object: immutable, value-compared, and
  # carrying the source location it came from so errors can point at the
  # user's file rather than at ours.
  module AST
    StepNode = Data.define(:name, :args, :options, :location) do
      def to_s
        parts = args.map(&:inspect) + options.map { |k, v| "#{k}: #{v.inspect}" }
        "#{name}(#{parts.join(', ')})"
      end

      def to_h = { name: name, args: args, options: options, location: location }
    end

    HandlerNode = Data.define(:kind, :callable, :location)

    PipelineNode = Data.define(:name, :steps, :handlers, :location) do
      def step_names = steps.map(&:name)
      def find(name) = steps.find { |s| s.name == name.to_sym }
      def handler(kind) = handlers.find { |h| h.kind == kind }

      def to_h
        { name: name, location: location,
          steps: steps.map(&:to_h), handlers: handlers.map(&:kind) }
      end

      # Returns a NEW pipeline: transformations never mutate.
      def with_steps(new_steps) = with(steps: new_steps.freeze)

      def insert_before(name, node)
        index = steps.index { |s| s.name == name.to_sym }
        raise Error, "no step named #{name.inspect} in #{self.name}" if index.nil?

        with_steps(steps.dup.insert(index, node))
      end
    end
  end
end

Explanation

Capturing source locations

    def method_missing(verb, *args, **options, &block)
      if !@registry.registered?(verb) && @outer.respond_to?(verb, true)
        return @outer.__send__(verb, *args, **options, &block)
      end

      @steps << ::Automation::AST::StepNode.new(
        name: verb, args: args.freeze, options: options.freeze, location: __where__
      )
      self
    end

    private

    # Two frames up: method_missing -> the user's line.
    def __where__
      frame = ::Kernel.caller_locations(2, 1).first
      "#{::File.basename(frame.path)}:#{frame.lineno}"
    end

caller_locations(start, length) returns frames from the call stack as objects with path, lineno and label. Counting the frames is fiddly and worth doing by experiment: frame 1 is method_missing itself, frame 2 is the line in the user's block that triggered it. Get it wrong and every step reports the same location inside your own library, which looks plausible and is useless.

This is the single highest-value feature in the milestone. A DSL that can say demo4.rb:29: unknown step :summarise feels like a compiler. A DSL that says NoMethodError in automation/runner.rb:31 feels like a leaky abstraction. The difference is six lines.

The validator, using reflection as a type check

  class Validator
    def initialize(registry)
      @registry = registry
    end

    def problems(pipeline)
      pipeline.steps.flat_map { |step| problems_for(step) }
    end

    def validate!(pipeline)
      found = problems(pipeline)
      raise InvalidPipeline, found unless found.empty?

      pipeline
    end

    private

    def problems_for(step)
      unless @registry.registered?(step.name)
        return ["#{step.location}: unknown step #{step.name.inspect} " \
                "(known: #{@registry.known.sort.join(', ')})"]
      end

      params = @registry.fetch(step.name).parameters
      missing_keywords(step, params) + unknown_keywords(step, params)
    end

    def missing_keywords(step, params)
      required = params.select { |type, _| type == :keyreq }.map(&:last)
      (required - step.options.keys).map do |key|
        "#{step.location}: step #{step.name} is missing required option #{key.inspect}"
      end
    end

    def unknown_keywords(step, params)
      return [] if params.any? { |type, _| type == :keyrest } # accepts **rest

      allowed = params.select { |type, _| %i[key keyreq].include?(type) }.map(&:last)
      (step.options.keys - allowed).map do |key|
        hint = allowed.empty? ? "it takes no options" : "it accepts: #{allowed.join(', ')}"
        "#{step.location}: step #{step.name} got unknown option #{key.inspect}; #{hint}"
      end
    end
  end

This is the part that answers Part 0's honest criticism. Ruby cannot check your program before it runs, but a program can check itself, because every callable can be asked what arguments it wants. parameters returns pairs like [[:opt, :items], [:keyreq, :topic], [:key, :since]], and from that you can derive: which options are required, which are permitted, and whether the step accepts anything at all via **rest.

Details that make it usable rather than merely correct:

What it looks like

--- the pipeline is data ---
research (4 steps, 0 handlers)
[:fetch, :filter, :summarize, :save_to]
{:name=>:research,
 :location=>"demo4.rb:11",
 :steps=>
  [{:name=>:fetch, :args=>["papers"], :options=>{:from=>"arxiv"}, :location=>"demo4.rb:12"},
   {:name=>:filter, :args=>[], :options=>{:topic=>"arxiv"}, :location=>"demo4.rb:13"},
   {:name=>:summarize, :args=>[], :options=>{:max_words=>3}, :location=>"demo4.rb:14"},
   {:name=>:save_to, :args=>["knowledge_base"], :options=>{}, :location=>"demo4.rb:15"}],
 :handlers=>[]}

--- nothing has executed yet; now it runs ---
  saved 1 items to knowledge_base
["papers from arxiv"]

--- validation catches mistakes before anything runs ---
pipeline is invalid:
  - demo4.rb:28: step fetch is missing required option :from
  - demo4.rb:29: unknown step :summarise (known: fetch, filter, save_to, summarize)
  - demo4.rb:30: step filter got unknown option :limit; it accepts: topic

--- building executes nothing, even for steps that would explode ---
built a pipeline containing :explode and nothing happened

--- when_failed runs the user's handler ---
  handler: explode failed with boom

--- transformations return new pipelines ---
[:fetch, :filter, :summarize, :save_to]
[:fetch, :filter, :deduplicate, :summarize, :save_to]

--- the data-only path: no code, no eval ---
[:fetch, :summarize]
["papers from"]

Three errors, each with the user's own file and line, reported together, before a single step ran. Compare that with Milestone 2, where the same pipeline would have fetched successfully and then failed on the typo. That is what the AST bought.

The safe mode

  # The safe path: build the same AST from plain data, with no code
  # execution at all. This is what you expose to untrusted input.
  def self.from_h(hash)
    steps = hash.fetch(:steps, []).map do |s|
      AST::StepNode.new(
        name: s.fetch(:name).to_sym,
        args: (s[:args] || []).freeze,
        options: (s[:options] || {}).transform_keys(&:to_sym).freeze,
        location: s[:location] || "(data)"
      )
    end
    AST::PipelineNode.new(
      name: hash.fetch(:name).to_sym, steps: steps.freeze,
      handlers: [].freeze, location: hash[:location] || "(data)"
    ).freeze
  end
An executable DSL is remote code execution

Say it plainly, because it is the most important practical fact about this technique. Automation.define takes a block of Ruby, so a pipeline file can open sockets, read ~/.ssh, or delete things. If pipelines are written by your own team and live in your own repository, that is fine and no different from any other code. If pipelines arrive from a web form, a customer, or a plugin marketplace, you must not eval them.

from_h is the answer: the same AST, built from JSON or YAML, containing no callables. The registry decides what any given step name can do, so an untrusted pipeline can only compose verbs you chose to expose. Ruby's $SAFE is gone (removed in 3.0) and never worked well; there is no sandbox. The boundary has to be "which representations can contain code", and this is where you draw it.

Two further precautions for the untrusted path: cap the number of steps, and do not allow handlers, since a handler is a block.

The tests that matter

  def test_building_executes_nothing
    ran = false
    @registry.register(:side_effect) { ran = true }

    Automation.define("dangerous", registry: @registry) { side_effect }

    refute ran, "building a pipeline must not run any step"
  end

  def test_every_step_records_where_it_was_written
    pipeline = Automation.define("located", registry: @registry) do
      fetch "papers", from: "arxiv"
    end

    assert_match(/test_define\.rb:\d+/, pipeline.find(:fetch).location)
  end

  def test_transformations_return_a_new_pipeline
    pipeline = Automation.define("t", registry: @registry) { summarize }
    node = Automation::AST::StepNode.new(name: :fetch, args: [], options: {}, location: "(test)")

    extended = pipeline.insert_before(:summarize, node)

    assert_equal %i[summarize], pipeline.step_names, "the original must not change"
    assert_equal %i[fetch summarize], extended.step_names
  end

  def test_caller_methods_are_forwarded_not_captured_as_steps
    pipeline = Automation.define("delegating", registry: @registry) do
      summarize max_words: helper_value
    end

    assert_equal 7, pipeline.find(:summarize).options[:max_words]
    assert_equal %i[summarize], pipeline.step_names
  end

  def helper_value = 7
$ ruby -Ilib -Itest test/all.rb
....................

20 runs, 50 assertions, 0 failures, 0 errors, 0 skips

test_building_executes_nothing is the invariant test of this course, the equivalent of Go's food conservation. It is two lines and it pins the architectural decision everything else depends on. If someone later "optimises" the builder by executing a step eagerly, this test is what stops them.

Note also test_caller_methods_are_forwarded_not_captured_as_steps: it is the naive-version failure from Milestone 3, turned into a regression test. Every bug you find by hand should become a test before you fix it, and a DSL's failures are especially worth pinning because they are silent.

Exercise 4

Make to_h and from_h a lossless round trip, and prove it.

  • Round trip: for any pipeline without handlers, Automation.from_h(pipeline.to_h) == pipeline must be true. Find out what currently breaks it (there is at least one thing) and fix it.
  • JSON: add to_json and Automation.from_json. Symbols do not survive JSON, so decide how to handle that and document the decision.
  • Hardening: make from_h reject input that is not safe: more than max_steps steps, step names that are not simple identifiers, and option values that are not strings, numbers, booleans, arrays or hashes of those. Raise a specific error naming the offending step.
  • Prove it with a test that round-trips three different pipelines and a test for each rejection.

Hint for the first part: compare pipeline.to_h with what from_h reconstructs, field by field, in irb. The mismatch is small and instructive.

Solution 4 — open after trying

What breaks the round trip. to_h preserves location (say "demo.rb:12") but from_h is usually given data with no location and substitutes "(data)", so the nodes differ. Data compares every field, including that one. Two defensible fixes: keep the original location when present (round trip preserves provenance), or exclude location from equality by comparing a normalised form. I prefer the first, with an explicit method for the second:

PipelineNode = Data.define(:name, :steps, :handlers, :location) do
  # Equality ignoring provenance, for round-trip tests and diffs.
  def same_shape?(other)
    other.is_a?(PipelineNode) &&
      name == other.name &&
      steps.map { |s| [s.name, s.args, s.options] } ==
        other.steps.map { |s| [s.name, s.args, s.options] }
  end
end

JSON and symbols. JSON has no symbol type, so :fetch becomes "fetch" and comes back as a String. Rather than guessing, convert at the boundary: from_h already calls to_sym on names and transform_keys(&:to_sym) on options. Document that option values stay strings, because converting them would be lossy in the other direction (a step legitimately wanting the string "lower" cannot be distinguished from one wanting :lower). If a step wants a symbol, it should convert it itself.

def self.from_json(text, max_steps: 100)
  from_h(JSON.parse(text, symbolize_names: true), max_steps: max_steps)
end

Hardening. The core of it:

SAFE_NAME = /\A[a-z_][a-z0-9_]*\z/
SAFE_SCALARS = [String, Integer, Float, TrueClass, FalseClass, NilClass].freeze

def self.from_h(hash, max_steps: 100)
  raw = hash.fetch(:steps, [])
  raise UnsafePipeline, "too many steps: #{raw.size} > #{max_steps}" if raw.size > max_steps
  raise UnsafePipeline, "handlers are not allowed in data pipelines" if hash[:handlers]&.any?

  steps = raw.map do |s|
    name = s.fetch(:name).to_s
    unless name.match?(SAFE_NAME)
      raise UnsafePipeline, "step name #{name.inspect} is not a plain identifier"
    end
    (s[:options] || {}).each do |key, value|
      unless safe_value?(value)
        raise UnsafePipeline, "step #{name}: option #{key} has unsupported value #{value.class}"
      end
    end
    # ... build the StepNode
  end
  # ...
end

def self.safe_value?(value)
  case value
  when *SAFE_SCALARS then true
  when ::Array then value.all? { |v| safe_value?(v) }
  when ::Hash  then value.all? { |k, v| safe_value?(k) && safe_value?(v) }
  else false
  end
end

Three things worth noticing about the hardening.

  • Allow-list, never deny-list. safe_value? lists what is permitted and rejects everything else. A deny-list ("reject Procs and Methods") fails the moment someone finds a type you did not think of, and someone always does.
  • The recursion in safe_value? is what stops a nested Hash from smuggling a callable three levels down. Validators that check only the top level are a recurring source of real vulnerabilities.
  • A step limit is not paranoia. Without it, a 10-million-step pipeline is a denial of service that costs the attacker one HTTP request. Every parser of untrusted input needs a size bound, and this is ours.

Finally, note that from_h being safe depends entirely on the registry containing only safe steps. If someone registers :shell, the data path executes shell commands as designed. Safety here is a property of the whole system, not of one method, and saying so in your README is part of the job.

Experiment

Change caller_locations(2, 1) to caller_locations(1, 1) and look at what the locations become: every step now reports a line inside define.rb. Then try 3 and watch them all point at the line that called Automation.define. Getting this right is pure experiment, and it is worth doing once so that you know how to fix it in your own libraries.

Common mistakes in Milestone 4
  • Counting stack frames wrong, so every error points inside your gem. Always test with a regex like /test_define\.rb:\d+/.
  • Validating one problem at a time. Collect them all; flat_map exists for this.
  • Mutating the AST in a transformation. steps << node raises FrozenError if you froze properly, and silently corrupts shared state if you did not.
  • Forgetting that Data compares every field. Including location, which is why the round trip in Exercise 4 fails first time.
  • Treating from_h as safe without bounding it. Unbounded size, arbitrary names and nested values are all attack surface.
  • Putting metadata in the options Hash instead of on the node, which leaks plumbing into the data your steps receive.

Checkpoint

  1. Name four capabilities that only exist because building and running are separate.
  2. What does Data.define give you that the hand-written Step class did not?
  3. How does the validator know which options a step accepts?
  4. Why does caller_locations(2, 1) use 2?
  5. Why does a step declared with **opts skip option checking, and is that a bug?
  6. What makes from_h safe for untrusted input, and what would make it unsafe again?
  7. Why must insert_before return a new pipeline rather than mutating?
Why are we using this language here?

The Mewlang cat, wearing glasses, looking confidentMilestone 3 is the strongest case for Ruby in this curriculum. Four lines of instance_eval plus method_missing turned a builder API into something that reads like a language, and the loop example (3.times { summarize index: i } producing three steps) shows what you get that a data format cannot offer at any price.

Milestone 4 is the honest correction. Everything we built there (source locations, a validator, an allow-list for untrusted input, a test asserting that building does not execute) is work that a compiled language would either give you free or make unnecessary. Racket, in Course 5, will do this at compile time: a typo in a step name becomes an error before the program runs, with the source location handled by the macro system rather than by counting stack frames. That comparison is the reason these two courses are adjacent in my recommended order.

The fair summary: Ruby lets you build the front end of a language in an afternoon, and then asks you to rebuild, by hand and at run time, the parts of a compiler you actually needed.

Repository state after Milestone 4

automation/
├── Gemfile, automation.gemspec, Rakefile, README.md
├── lib/
│   ├── automation.rb            requires, global registry, register/reset
│   └── automation/
│       ├── version.rb
│       ├── errors.rb            Error, UnknownStep, StepFailed, InvalidPipeline
│       ├── step.rb              milestone 1's frozen value object
│       ├── registry.rb          name -> callable, with useful failures
│       ├── pipeline.rb          milestone 2: builder + reduce interpreter
│       ├── dsl.rb               milestone 3: NaiveDSL and DSLBuilder
│       ├── ast.rb               StepNode, HandlerNode, PipelineNode (Data)
│       ├── define.rb            ASTBuilder, Automation.define, from_h
│       ├── validator.rb         problems with file:line
│       └── runner.rb            a minimal interpreter over the AST
└── test/
    ├── test_helper.rb
    ├── test_step.rb             value semantics, freezing
    ├── test_registry.rb         duck typing, useful errors
    ├── test_define.rb           locations, immutability, nothing executes
    └── test_validator.rb        every validation rule
$ ruby -Ilib -Itest test/all.rb
20 runs, 50 assertions, 0 failures, 0 errors, 0 skips
$ git commit -am "milestone 4: an AST, source locations, and a validator"

Note that pipeline.rb and dsl.rb are still there. Keep them: they are the Milestone 2 and 3 designs, they still pass their tests, and a reader of your repository can follow the same progression you did. Deleting the earlier versions is throwing away the argument.

The Mewlang cat, stretching and relaxedInstalment 7 of the five-course curriculum. Next: Ruby Milestones 5–8, where the runner grows a context and middleware, failures get retries and handlers that actually work, plugins arrive via define_method and method_missing, and the steps start doing real work against HTTP, the filesystem and SQLite.

Continue