Milestones 9–12

Instalment 10 · Course 2 (Ruby) · Advanced phase and finish

Measuring the magic, compiling it away, and checking whether any of it stuck

Six advanced topics with real measurements, a final challenge that turns your interpreter into a compiler, forty-one questions, and everything you need to put this on GitHub.

Verification note

Every benchmark, allocation count and code-reading answer below was produced by running the code on Ruby 3.2.3. Timings come from a single-core container, so treat ratios as meaningful and absolute numbers as indicative.

Part AThe advanced phase

A1 · The object model as a debugging tool

The Mewlang cat, in a neutral curious side poseMetaprogramming makes "where did this method come from?" a real question. Ruby answers it directly, and these four calls will save you hours:

Automation::Steps::Filter.instance_method(:matching).owner        # which module defined it
Automation::Steps::Filter.instance_method(:matching).source_location # file and line
Automation::Steps::Filter.ancestors.first(5)                      # the lookup chain
object.singleton_methods                                          # methods on this object alone

For a method generated by define_method inside a class macro, source_location points at the macro, not at the option :matching line that triggered it. That is worth knowing before it confuses you: to make generated methods self-documenting you have to record the call site yourself, exactly as Milestone 4 did for steps.

The lookup chain, demonstrated:

module A; def who = "A"; end
module B; def who = "B"; end
class C
  include A
  prepend B
  def who = "C"
end
p C.new.who, C.ancestors.first(4)
"B"
[B, C, A, Object]

prepend goes before the class, include goes after it. So a prepended module wins over the class's own method (which is why it works for instrumentation), and an included module loses to it (which is why include provides defaults rather than overrides). Memorising the order is less useful than remembering that ancestors will always tell you.

ToolAdds methods toPositionUse for
include Minstancesafter the classshared behaviour the class may override
prepend Minstancesbefore the classwrapping with super: instrumentation, logging
extend Mthe object itselfits singleton classclass methods, per-object behaviour
using Mlexical scoperefinementmonkey-patching without global effect

A2 · Measure before you metaprogram

Milestone 7 replaced method_missing with generated methods and I claimed it was faster. Here is the number, 300,000 calls each:

                       user     system      total        real
define_method      0.019719   0.000818   0.020537 (  0.020542)
method_missing     0.040257   0.000662   0.040919 (  0.040936)

Twice as slow, which is the shape you would predict: method_missing runs only after a full method lookup has failed, so you pay for the search and then for the dispatch. Two honest observations about that result.

It is only 20 nanoseconds per call. For a DSL evaluated once at startup this is irrelevant, and choosing generated methods was justified by typo detection and introspection, not speed. If someone tells you method_missing is "too slow" for a configuration DSL, ask them for the number.

Benchmark ships with Ruby. benchmark/ips (iterations per second, with statistical comparison) is the better tool and is a gem; the standard Benchmark.bm above is enough to settle a 2× question. Use the real profiler (stackprof, also a gem) before optimising anything in anger.

A3 · Lazy pipelines for data that does not fit

Our runner passes an Array between steps, so a pipeline over 10 million records needs 10 million records in memory. Ruby's answer is Enumerator::Lazy: chain map and select without materialising anything, and pull only what you consume.

def records = (1..2_000_000).lazy.map { |i| { id: i, topic: i.even? ? "AI" : "bio" } }

# eager, over 200,000 records
(1..200_000).map { |i| { id: i, topic: i.even? ? "AI" : "bio" } }
            .select { |r| r[:topic] == "AI" }
            .map { |r| r[:id] }
            .first(5)

# lazy, over 2,000,000 records
records.select { |r| r[:topic] == "AI" }.map { |r| r[:id] }.first(5)
[2, 4, 6, 8, 10]
[2, 4, 6, 8, 10]
eager over 200k: 0.3803s | lazy over 2M: 0.0002s

Same answer; the lazy version was roughly two thousand times faster over ten times as much data, because it built exactly ten records. Chained eager calls allocate a full intermediate array per step; lazy chains fuse into one pass and stop as soon as the consumer is satisfied.

To support this in the gem, a step declares that it streams:

class Filter < Automation::Plugin
  streaming true      # promises: takes an Enumerable, returns an Enumerable,
                      # never calls #size, #sort or #to_a on its input

  def call(records) = records.select { |r| matches?(r) }
end

The runner then wraps the initial payload in .lazy when every step up to the first non-streaming one declares it, and materialises at the first step that does not. The metadata pattern from Milestone 7 makes this a five-line feature: the declaration is already the mechanism.

The catch to document: a lazy payload means failures move. An exception from fetch no longer happens during the fetch step; it happens later, when save_to finally pulls a value, so your beautiful per-step error reporting blames the wrong step. Streaming and precise error attribution are in genuine tension, and the usual resolution is to materialise at well-chosen boundaries rather than everywhere or nowhere.

A4 · Concurrency, honestly

CRuby has a global VM lock, so Ruby threads do not run Ruby code in parallel. They do run in parallel during I/O, because the lock is released around blocking calls:

def slow_io(n) = sleep(0.05) || n

serial   = Benchmark.realtime { 4.times { |i| slow_io(i) } }
parallel = Benchmark.realtime { 4.times.map { |i| Thread.new { slow_io(i) } }.each(&:join) }
serial 0.201s | threaded 0.069s

Three times faster on a single core, because all four sleeps overlapped. Our pipeline steps are HTTP calls and file writes, so this is exactly the case that benefits, and a fan_out step that runs N sub-pipelines concurrently is worth having.

Build it with a bounded pool rather than a thread per item, for the same reason Go's Milestone 11 bounded its workers:

def fan_out(items, concurrency: 8)
  queue = Queue.new
  items.each { |i| queue << i }
  results = Array.new(items.size)

  workers = concurrency.times.map do
    Thread.new do
      while (index = queue.pop(true) rescue nil)
        results[index] = yield(items[index])
      end
    end
  end
  workers.each(&:join)
  results
end

Queue is thread-safe and ships with Ruby; pop(true) is non-blocking and raises ThreadError when empty, which is how the workers know to stop. Writing to distinct indices of a preallocated array is safe here for the same reason it was in Go: separate slots, no resizing.

The Mewlang cat, looking up with curiosityWhat about Ractors?

Ractors give real parallelism for Ruby code by giving each one its own GVL, at the cost of strict isolation: objects crossing a Ractor boundary must be immutable (frozen and deeply so) or copied, and most gems are not Ractor-safe. They remain officially experimental and print a warning on first use.

Our AST would actually pass the isolation test, since it is frozen Data all the way down, which is a pleasing accident of good design. Our steps would not: they close over registries, logger objects and sockets. The honest position is that Ractors are worth an afternoon of curiosity and are not yet worth building a gem's public API around. If you need CPU parallelism in Ruby today, the practical answers are multiple processes or moving the hot work into C, and if those sound unappealing, that is a reason to consider Go for that component.

A5 · Allocation discipline

Ruby's garbage collector is good and allocation is still the main cost in most Ruby hot loops. With GC disabled to keep the measurement stable, 10,000 iterations each:

{:unfrozen_literal => 10005,
 :frozen_literal   => 2,
 :symbol           => 2,
 :data_node        => 30005,
 :hash_node        => 20002}

Note the measurement technique. My first attempt, without GC.disable, reported negative allocations for the string case, because the collector ran mid-measurement and freed more than the loop created. Naive allocation counting is noise unless you disable the GC or use ObjectSpace.trace_object_allocations, and a benchmark that can report a negative count is a benchmark you should not trust in either direction.

A6 · Refinements, types, and docs

Refinements are scoped monkey-patching: a change to a core class that exists only in the file or module that asks for it.

module StepSugar
  refine Symbol do
    def to_step = "step(#{self})"
  end
end

:fetch.to_step              # => NoMethodError

module Scoped
  using StepSugar
  def self.demo = :fetch.to_step
end

Scoped.demo                 # => "step(fetch)"
:fetch.to_step              # => NoMethodError, still

This is the responsible way to add methods to classes you do not own, and it is why using exists. We did not use refinements in the DSL for a reason worth stating: they are lexically scoped, so they do not apply inside a block that was defined elsewhere, which is precisely how our DSL blocks are evaluated. A refinement-based DSL would work in the file that declared it and mysteriously fail everywhere else.

Types for a dynamic gem

Ruby 3 ships RBS, a separate signature language, and steep type-checks against it:

# sig/automation.rbs
module Automation
  class Registry
    def register: (Symbol | String, ?_Callable, ?spec: Hash[Symbol, untyped]?, ?doc: String?)
                  { (untyped, *untyped) -> untyped } -> self
    def fetch: (Symbol | String, ?location: String?) -> _Callable
    def known: () -> Array[Symbol]
  end

  interface _Callable
    def call: (untyped, *untyped, **untyped) -> untyped
  end
end

Signatures live outside the code, so they can be added incrementally and do not clutter a DSL. Be realistic about coverage: the parts of this gem that are ordinary objects (registry, config, store, AST) type-check well; the parts built on method_missing and generated methods cannot be described at all, because their signatures do not exist until run time. A gem like this one is exactly where static typing gives least, which is worth knowing when someone proposes adding it everywhere.

For documentation, YARD reads @param/@return comments and, more importantly here, lets you document methods that do not exist:

# @!method fetch(from:, limit: nil)
#   Fetch a JSON array of records from an HTTP endpoint.
#   @param from [String] URL returning a JSON array
#   @param limit [Integer, nil] keep at most this many records

That @!method directive is how you stop dynamic definitions from being invisible to your users. It is also a smell detector: if you cannot write the directive, your users cannot discover the method.

CI and supply chain

# .github/workflows/ci.yml (the shape, not the whole file)
strategy:
  matrix:
    ruby: ["3.2", "3.3", "3.4"]
steps:
  - uses: ruby/setup-ruby@v1
    with: { ruby-version: "${{ matrix.ruby }}", bundler-cache: true }
  - run: bundle exec rake test
  - run: bundle exec rubocop
  - run: gem install bundler-audit && bundle audit check --update

Test the Ruby versions your gemspec claims to support, or do not claim them. bundler-audit checks your lockfile against known advisories, which for a gem with zero runtime dependencies is a short conversation and remains worth automating.


Part BThe final challenge

Compile the pipeline instead of interpreting it

The Mewlang cat, thinking with a paw to its chinEverything so far interprets: the runner walks the AST at run time, looks each step up in a registry, and calls it. Now write a compiler that turns a pipeline into standalone Ruby source which runs with no interpreter, no registry lookups, and no dependency on this gem.

$ automation compile examples/research.rb --output research_compiled.rb
$ ruby research_compiled.rb          # no `require "automation"` anywhere

Requirements

  1. Automation::Compiler.new(registry).compile(pipeline) returns a String of valid Ruby defining a single method run(input = nil) and calling it.
  2. Step implementations are inlined into the generated source. A step registered as a block or lambda has its source extracted; a step registered as a Plugin subclass has its class definition emitted.
  3. Retries, dry-run support and when_failed handling are generated as ordinary Ruby control flow, not as a middleware chain.
  4. The generated file has no runtime dependency on the gem and passes ruby -c (syntax check) and RuboCop.
  5. Compilation fails loudly, naming the step, when a step's implementation cannot be inlined.

Constraints

Acceptance criteria

  1. Differential testing. For at least five pipelines, interpreted and compiled runs produce identical payloads and identical failure behaviour. Write this as a test that generates, writes, shells out to ruby, and compares.
  2. Speed. The compiled version is measurably faster on a pipeline of ten trivial steps run 10,000 times. Report the ratio and be honest if it is small.
  3. Safety. A test proves that a hostile option value ends up as data, not code.
  4. Failure. Compiling a pipeline whose step was registered as a C-implemented method or an object without extractable source raises a clear error naming the step.

Hints

The Mewlang cat, happy and celebratingThis is the most valuable exercise in the Ruby course, because it is the bridge to Course 5. Racket does this at compile time, with the language's own macro system, and produces better error messages while doing it. Building the clumsy version by hand is what makes the elegant version legible.

Solution — only look after trying

Shape

module Automation
  class Compiler
    class NotCompilable < Error; end

    def initialize(registry) = @registry = registry

    def compile(pipeline)
      steps = pipeline.steps.each_with_index.map { |step, i| compile_step(step, i) }

      <<~RUBY
        # Generated by automation #{VERSION} from #{pipeline.location}
        # Pipeline: #{pipeline.name}. Do not edit; regenerate instead.
        # frozen_string_literal: true

        #{inlined_definitions(pipeline).join("\n\n")}

        #{steps.join("\n\n")}

        def run(input = nil)
        #{run_body(pipeline).gsub(/^/, "  ")}
        end

        puts run.inspect if $PROGRAM_NAME == __FILE__
      RUBY
    end

    private

    def compile_step(step, index)
      <<~RUBY
        def step_#{index}_#{step.name}(input)
          #{call_expression(step)}
        end
      RUBY
    end

    # Values become literals via #inspect, never via interpolation.
    def literal(value)
      case value
      when String, Symbol, Numeric, TrueClass, FalseClass, NilClass then value.inspect
      when Array then "[#{value.map { |v| literal(v) }.join(', ')}]"
      when Hash then "{#{value.map { |k, v| "#{literal(k)} => #{literal(v)}" }.join(', ')}}"
      else raise NotCompilable, "cannot compile a #{value.class} option value"
      end
    end
  end
end

The four decisions that matter

1. inspect is the safe interpolation. For every type the allow-list permits, value.inspect produces a literal that evaluates back to an equal value, with quoting and escaping handled by Ruby itself. Building "from: \"#{value}\"" by hand is how you get injection; the hostile string from the acceptance criteria comes out as an inert, correctly escaped string literal. The else branch is as important: refusing to compile a Proc or a socket is the right answer, not guessing.

2. Plugin classes compile, bare blocks mostly do not. Extracting the text of a block from its source_location means re-parsing the file and finding where the block ends, which is genuinely hard (method_source does it by trying successively longer substrings until ruby -c accepts them, which tells you how hard). A plugin class, by contrast, lives in a known file that you can copy wholesale. The honest resolution is that compilation is a feature of class-based steps, and the compiler raises NotCompilable naming any block-registered step, with a message suggesting the plugin form. Discovering that a feature imposes a constraint on an earlier design decision is the normal shape of this kind of work.

3. Differential testing is the only credible verification.

  def assert_compiles_identically(pipeline, input = nil)
    interpreted = Automation::Runner.new(registry: test_registry).run(pipeline, input)

    Dir.mktmpdir do |dir|
      file = File.join(dir, "compiled.rb")
      File.write(file, Automation::Compiler.new(test_registry).compile(pipeline))

      assert system("ruby", "-c", file, out: File::NULL), "generated code is not valid Ruby"

      compiled = `ruby #{file}`.chomp
      assert_equal interpreted.payload.inspect, compiled
    end
  end

Two implementations, one specification, compared on real inputs. This is the same technique compiler authors use against a reference interpreter, and it is the only way to have confidence in a generator: unit-testing the generated string tests your formatting, not your semantics.

4. The speed result is probably disappointing, and that is the lesson. Expect something between 1.2× and 2× on trivial steps. The interpreter's overhead is a registry lookup and a few lambda calls per step; if your steps do any real work (an HTTP request, a regex over a megabyte), that overhead is noise. Compilation's real payoff here is not speed, it is artefact independence: a single reviewable file you can commit, audit, ship to a machine that does not have your gem, and diff when it changes. Say so in your README rather than claiming a performance win you cannot substantiate.

What this teaches about Course 5

Notice what was painful: getting at the source text of user code, generating syntactically valid output by string concatenation, keeping the generated code readable, and producing errors that point at the user's file. All four are string-level problems, because Ruby gives you no access to its own syntax as data.

Racket's macros operate on syntax objects: structured, already-parsed code that carries its own source locations and binding information. The equivalent compiler there is not a string generator; it is a function from syntax to syntax, it cannot produce invalid output, and error messages point at the user's line for free. When you meet that in Course 5, this challenge is the thing it will be compared against.


Part CKnowledge check

C1 · Twenty conceptual questions

  1. Explain what instance_eval changes and what it leaves alone. Name the four kinds of identifier and how each behaves inside the block.
  2. Why is BasicObject the right superclass for a DSL builder, and what does it cost?
  3. What is a Binding, and what is block.binding.receiver used for here? What is the conservative alternative?
  4. Give three reasons the DSL builds an AST instead of executing as it parses.
  5. How does a dynamic language produce build-time errors with file and line numbers?
  6. When should you prefer define_method over method_missing, and when is method_missing genuinely necessary?
  7. What two methods must always be defined together, and what breaks otherwise?
  8. Explain the class-macro pattern using option as the example. What are the two effects of one call?
  9. Why must inherited call super?
  10. Compare include, prepend, extend and using in one sentence each.
  11. Why is a proc's return dangerous when you store user callbacks, and what do you use instead?
  12. When does Ruby set Exception#cause? Why did the runner not get one?
  13. What is the middleware fold, and why is the list reversed before folding?
  14. Why does the runner return a RunResult rather than raising, and what is the cost of that choice?
  15. What makes from_h safe for untrusted pipelines, and what would make it unsafe again?
  16. Why must policies be idempotent, and how do you detect that they are not?
  17. Why does the project produce successor pipelines instead of mutating running ones?
  18. What does the frozen_string_literal comment actually do, and what did it measure at?
  19. Explain Ruby's GVL in terms of what threads can and cannot speed up.
  20. For a DSL gem, what counts as a breaking change under semantic versioning?
Answers to C1
  1. It changes self for the duration of the block, and nothing else. Locals still resolve (the block is a closure), constants resolve lexically where the block was written, method calls go to the new self (so the caller's methods are unreachable without delegation), and instance variables resolve on the new self, which is why @server silently became nil in three of my tests.
  2. Because it has almost no methods, nearly every verb falls through to method_missing rather than hitting an inherited Object method. The cost is that you lose everything too: constants need ::, puts and raise are unavailable, and plumbing methods need __ugly__ names so they do not collide with user verbs.
  3. A Binding captures an execution context: local variables, self, and the lexical scope. block.binding.receiver recovers the self where the block was written so unknown verbs can be forwarded back to it. The conservative alternative is requiring the caller to pass it: define("x", context: self), uglier and more honest.
  4. Dry runs and validation need to walk the pipeline without side effects; source locations and diffs need it to be data; and transformations need it to be a value you can produce a new version of. The invariant test test_building_executes_nothing pins all three.
  5. By capturing caller_locations at the moment each construct is built, and storing the file and line on the node. Everything downstream (validator, explain, logs, CLI) then reports the user's own file. Getting the frame depth right is experimental, and wrong depths report inside your own library.
  6. Prefer define_method whenever you know the names: the methods really exist, are faster, appear in instance_methods, and anything not caught by them is a typo you can report. method_missing is necessary when the set is genuinely open, and remains useful as the fallback that turns unknown names into good errors.
  7. method_missing and respond_to_missing?. Without the second, respond_to? lies, method() raises NameError, and any library that checks before calling will skip your object.
  8. option :max_words, default: 10 records metadata in a class-level Hash and calls define_method to generate a reader. The metadata drives validation, documentation and CLI help; the reader makes max_words work inside call. One declaration, two uses, which is why attr_accessor, validates and let all look like this.
  9. Because other libraries may define inherited earlier in the ancestor chain, and skipping super silently breaks them. The same applies to every hook: included, method_added, const_missing.
  10. include: instance methods, after the class in lookup. prepend: instance methods, before the class, so super reaches the original. extend: methods on that object alone, via its singleton class. using: activates a refinement in the current lexical scope only.
  11. A return inside a proc returns from the enclosing method, so a user's handler can make your library return from somewhere unexpected; a lambda's return exits only the lambda. A block captured with &block is a proc, so store lambdas when you can, or document the hazard.
  12. Only when an exception is raised inside a rescue block. Our runner constructs StepFailed without raising it at that point, so no chain was recorded and the handler received the wrapper. The fix is an explicit original field.
  13. Each middleware is (step, context, nxt); the fold wraps the innermost behaviour in each layer, producing one callable. The list is reversed so that folding produces the layers in the order written; folding forwards runs them inside-out.
  14. Because a failed run still has partial results worth showing to a CLI, a dashboard or a test, and an exception discards that context. The cost is that an ignored result is an ignored error, which bit me in Milestone 7; mitigate with run! at the top level or a non-zero exit code.
  15. It builds the AST from JSON with no callables anywhere, so an untrusted pipeline can only compose verbs you registered. It becomes unsafe if you register a dangerous verb (:shell), if you allow handlers (which are blocks), or if you skip the size and value allow-lists.
  16. Because they run on every define, so a policy that appends unconditionally grows the pipeline without bound. Detect it by applying policies repeatedly until the result stops changing, with a round limit; value equality on Data makes the fixpoint check one ==.
  17. Because a run should have one fixed definition: logs describe something that still exists, a bad transformation is discarded rather than corrupting a live run, and the runner never has to handle steps appearing mid-flight. Query planners and JIT compilers work this way too.
  18. It freezes every string literal in the file, so identical literals are deduplicated into one object. Measured: 10,005 objects allocated for 10,000 unfrozen literals against 2 for frozen ones.
  19. CRuby's global VM lock means only one thread runs Ruby bytecode at a time, so threads do not speed up CPU work. The lock is released around blocking I/O, so they do help there: four 50 ms sleeps took 201 ms serially and 69 ms threaded.
  20. Anything that changes the meaning of a pipeline file: renaming a verb or an option, changing a default, making an optional option required, or changing what a step returns. None of these is a Ruby method signature change, which is why the usual "public API" framing misleads for DSL gems.

C2 · Ten code-reading questions

Predict the output. All ten were run; answers below.

# 1
module A; def who = "A"; end
module B; def who = "B"; end
class C
  include A
  prepend B
  def who = "C"
end
p C.new.who, C.ancestors.first(4)

# 2
def run_block = [yield(1, 2, 3), block_given?]
p run_block { |a, b| [a, b] }

# 3  (no frozen_string_literal comment in this file)
s = "abc".dup
s << "d"
p s, s.frozen?, "abc".frozen?

# 4
Point = Struct.new(:x, :y)
pt = Point.new(1, 2)
pt.x = 9
D = Data.define(:x, :y)
d = D.new(x: 1, y: 2)
p pt.to_a, d.with(x: 9), d.respond_to?(:x=)

# 5
def f(a, *rest, k: 1, **opts) = [a, rest, k, opts]
p f(1, 2, 3, k: 4, z: 5)
h = { k: 9, z: 8 }
p f(1, **h)
p f(1, h)

# 6
@a = false
@a ||= "set"
@b = nil
@b ||= "set"
p @a, @b

# 7
p([1, 2, 3].reduce([]) { |acc, n| acc << n * 2 })
p([1, 2, 3].reduce(0) { |acc, n| acc + n if n.odd? })

# 8
class Gen
  limit = 3
  define_method(:with_closure) { limit }
  def with_def = defined?(limit) ? limit : "limit is not visible here"
end
p Gen.new.with_closure, Gen.new.with_def

# 9
class Ghost
  def method_missing(n, *) = n.to_s.start_with?("go_") ? :ok : super
end
g = Ghost.new
p g.go_home, g.respond_to?(:go_home)

# 10
class Widget; end
Widget.class_eval { def hello = "instance method" }
Widget.instance_eval { def hi = "class method" }
p Widget.new.hello, Widget.hi
Answers to C2
1.  "B"  /  [B, C, A, Object]
2.  [[1, 2], true]
3.  "abcd"  /  false  /  false
4.  [9, 2]  /  #<data D x=9, y=2>  /  false
5.  [1, [2, 3], 4, {:z=>5}]
    [1, [], 9, {:z=>8}]
    [1, [{:k=>9, :z=>8}], 1, {}]
6.  "set"  /  "set"
7.  [2, 4, 6]  then NoMethodError (nil accumulator)
8.  3  /  "limit is not visible here"
9.  :ok  /  false
10. "instance method"  /  "class method"
  1. prepend puts B before C in the chain, so B wins over the class's own method. include puts A after C, so A never gets a look in.
  2. Blocks are lenient about arity: three yielded values, two parameters, the third is dropped. A lambda would have raised ArgumentError.
  3. Without the magic comment, string literals are unfrozen, so "abc".frozen? is false. Add # frozen_string_literal: true at the top and the third value becomes true. The answer to a question about Ruby strings genuinely depends on a comment.
  4. Struct is mutable and positional; Data is immutable with keyword construction, with for copies, and no writers at all.
  5. The first two lines show splat and double-splat collecting. The third is the Ruby 3 separation: a bare Hash is positional, so it lands in rest and k keeps its default. In Ruby 2 this was auto-converted to keywords, and the change broke a great deal of code.
  6. ||= assigns when the current value is nil or false. Using it to memoise a legitimately false value re-evaluates every time; use defined? or a Hash with key? when false is a real answer.
  7. acc << n * 2 works because << returns the array. The second raises because when n is even the block returns nil, and the next iteration calls nil + n. In reduce, the block's return value is the accumulator, which is why each_with_object is safer when you are building something.
  8. define_method takes a block, and blocks are closures, so limit is captured. def opens a new scope and sees nothing from the class body. This is the main practical reason to reach for define_method.
  9. method_missing handles the call, but respond_to? has no idea, because respond_to_missing? was not defined. The object works and lies about itself.
  10. class_eval evaluates in the context of the class, so def defines an instance method. instance_eval evaluates with the class as self, so def defines a method on the class object, that is, a class method. The names feel backwards to everyone at first.

C3 · Five debugging exercises

  1. Symptom: a pipeline built in a test has one step instead of two, no error, no warning.
    summarize = stub_step(:summarize) { |items| items.first(2) }
    pipeline = build_pipeline("t") do
      fetch from: "x"
      summarize
    end
  2. Symptom: explain reports every step's location as define.rb:113 instead of the user's file.
    def __where__
      frame = ::Kernel.caller_locations(1, 1).first
      "#{::File.basename(frame.path)}:#{frame.lineno}"
    end
  3. Symptom: a when_failed handler prints step fetch failed: step fetch failed: timeout, doubled.
    handler.callable.call(error, step)   # error is the StepFailed wrapper
  4. Symptom: after adding a policy, every define is slower than the last and eventually raises "rewrite produced 501 steps".
    Automation::Policies.register(:audit) do |pl|
      pl.rewrite { append step(:audit_log) }
    end
  5. Symptom: a third-party step works when registered as a lambda and crashes the validator when registered as a custom object.
    params = @registry.fetch(step.name).parameters
Answers to C3
  1. A local variable shadowed the verb. Once Ruby's parser has seen summarize = ..., the bare word summarize inside the block is that variable, not a method call, so no step is added. Rename the local (summarizer) or force a call with summarize(). Diagnose it by printing pipeline.step_names, which is faster than staring at the block.
  2. Wrong stack depth. caller_locations(1, 1) from inside __where__ returns __where__'s caller, which is your own method_missing or __add_step__, not the user. Count the frames experimentally and pin the result with a test asserting the location matches the test file's own name, which is what stops the bug returning.
  3. Passing the wrapper instead of the original. The handler should receive the cause: error.respond_to?(:original) ? error.original : error. The doubled message is the tell: your wrapper's message already contains the inner message, so printing the wrapper inside a message that describes the step says everything twice.
  4. A non-idempotent policy. It appends unconditionally, so every application adds another audit_log, and policies run on every define including on pipelines produced by rewrites. Guard it: pl.find(:audit_log) ? pl : pl.rewrite { ... }. The step limit turning it into an error rather than an out-of-memory crash is the limit doing its job.
  5. A duck-typing assumption. The registry accepts anything with #call, but the validator additionally requires #parameters, which lambdas and Methods have and arbitrary objects do not. Ask for impl.method(:call).parameters instead, and return nil (meaning "cannot check") rather than raising for anything that cannot answer.

C4 · Five implementation exercises

  1. Conditional steps. Add only_if and unless options evaluated at run time against the context vars, so fetch from: url, only_if: :full_refresh is skipped unless the var is set. The result must record the step as :skipped, and explain must show the condition.
  2. Sub-pipelines. Let a pipeline include another by name (include_pipeline "cleanup"), with cycle detection and a clear error naming the cycle.
  3. Parallel fan-out. Implement a fan_out step that runs several sub-pipelines concurrently with a bounded thread pool, merges their payloads, and reports the slowest branch.
  4. Persistence and resume. Checkpoint the payload after each step so an interrupted run can resume from the last completed step. Decide what makes a payload safely resumable and document the limits.
  5. A second front end. Write a YAML loader producing the same AST as the Ruby DSL, with the same validation and source locations (YAML parsers can report line numbers). Prove with a test that the two front ends produce equal pipelines for equivalent input.

C5 · One substantial challenge

Build a language server for your DSL. Implement a minimal LSP server (stdin/stdout, JSON-RPC) that gives an editor: completion for step names with their documentation, hover showing a step's options and defaults, and diagnostics from the validator with correct ranges, for any file using your gem.

Requirements: it must derive everything from the registry rather than a hard-coded list, so a user's plugins appear automatically; diagnostics must update on save; and it must not execute the pipeline file more than once per change. Hints: the hard part is not the protocol (it is a few message types) but deciding when it is safe to load a user's file, given that loading is executing. Think about a subprocess, a timeout, and what to do when the file has a syntax error.

C6 · You should now be able to explain

C7 · You should now be able to implement


Part DShipping it: README, portfolio, interview

D1 · README draft

# automation

A Ruby DSL for describing automation pipelines, with the unusual property
that a pipeline is **data you can inspect, validate, diagram and rewrite**
before anything runs.

No runtime dependencies. Ruby 3.2+.

```ruby
Automation.define("research") do
  retry_on Automation::HttpError, times: 3, backoff: :exponential

  fetch     from: "https://example.com/papers.json", limit: 20
  filter    field: :topic, matching: "AI"
  summarize field: :abstract, max_words: 40
  save_to   collection: "knowledge_base"

  when_failed { |error, step| warn "#{step.name}: #{error.message}" }
end
```

    $ automation explain examples/research.rb
    $ automation run examples/research.rb --dry-run
    $ automation graph examples/research.rb > pipeline.mmd

## Why this exists

Every team eventually writes a YAML file that wants to be a program, then
invents conditionals, loops and variables badly. This starts from the other
end: it is Ruby, so loops and conditionals were always there, and the DSL is
a thin surface over an AST that the tooling can reason about.

## What you get

- **Errors before anything runs**, with your file and line:

      pipeline is invalid:
        - research.rb:7: step fetch is missing required option :from
        - research.rb:8: unknown step :summarise; did you mean summarize?
        - research.rb:9: step filter got unknown option :limit; it accepts: field, matching

- **Dry runs** that provably touch nothing (there is a test asserting the
  fake HTTP server receives zero requests).
- **Retries** with linear or exponential backoff, and failure handlers.
- **Inspection**: `explain`, Mermaid diagrams, and diffs between versions.
- **Rewriting**: pipelines are values, so transformations produce new ones.
- **Policies**: organisation-wide rules applied to every pipeline at define
  time (always log first, cap expensive options, require a retry policy).
- **A safe mode**: load pipelines from JSON, executing no code at all.
- **Testing helpers** for your own pipelines and plugins.

## Adding a step

```ruby
class Translate < Automation::Plugin
  step_name :translate
  doc "Translate a field into another language."

  option :field, default: :abstract
  option :to, required: true

  def call(records)
    records.map { |r| r.merge(field => Translator.call(r[field], to: to)) }
  end
end
Translate.register!
```

The declared options drive validation, `--help`, `explain` output and typo
suggestions. Nothing else needs to know your step exists.

## Trust

`automation run file.rb` **executes Ruby**. Only run pipeline files you
trust, exactly as you would with a Rakefile or a Gemfile.

For untrusted input use the data path, which builds the same AST from JSON
and executes no code:

    $ automation run --safe pipeline.json

The registry is the boundary: an untrusted pipeline can only compose verbs
you chose to register.

## Known limitations

- A local variable with the same name as a step silently shadows the verb
  inside a `define` block. Name spies and helpers distinctly.
- Instance variables do not resolve inside the DSL block (`self` is the
  builder). Bind them to locals first.
- `PStore` is the default store; it rewrites the whole collection per
  transaction and does not scale past a few megabytes. Swap in SQLite
  behind `Automation::Store`'s three methods.
- Steps registered as blocks cannot be compiled to standalone Ruby; use
  plugin classes if you need that.

## Licence

MIT

Three deliberate choices in that README. It leads with the property that is unusual (pipelines are inspectable data) rather than with "a DSL for pipelines", which describes fifty gems. It shows real error output, because that is what convinces an experienced reader that the tool was used rather than published. And the known limitations section lists the two DSL traps I fell into, which costs nothing and is the difference between a user trusting your judgement and discovering the traps alone at midnight.

D2 · GitHub project description

A Ruby DSL where pipelines are inspectable data: validated with file and line before anything runs, diagrammable, rewritable, and loadable from JSON with no code execution. No runtime dependencies.

Topics: ruby, dsl, metaprogramming, ast, pipeline, automation, workflow, gem, plugin-architecture, interpreter.

D3 · Performance considerations

D4 · Security considerations

D5 · What to put in your portfolio

The Mewlang cat, wearing glasses, looking confidentPresent this as a study of what it takes to make a DSL trustworthy, not as "a pipeline runner". The interesting narrative is the sequence of realisations:

  1. A block plus instance_eval makes configuration read like a language, in about four lines.
  2. Doing that naively breaks in three silent ways, each demonstrated with real output: caller methods swallowed as steps, verbs shadowed by Object methods, typos accepted as valid.
  3. The fix that matters is not syntactic: build an AST and execute nothing, which then gives validation with file and line, dry runs, diagrams, diffs and rewriting.
  4. Declared plugin metadata drives validation, documentation and typo suggestions from one declaration.
  5. An executable configuration format is remote code execution, so there is a second, data-only front end sharing the same AST and engine.

Keep a docs/ folder with the explain output, a rendered Mermaid diagram, and the error-message screenshots. In an interview, the error messages are the artefact to lead with: "here is a dynamic language producing compile-time-quality diagnostics, and here is how".

Cut, if you want it tighter: the terminal-free extras (graph is charming but inessential), and the Milestone 2 Pipeline class, unless you keep it explicitly as the "before" half of the comparison and say so.

D6 · Interview questions someone could ask

QuestionWhat a strong answer includes
How does the DSL work?instance_eval changing self, method_missing or generated methods recording steps, and the fact that nothing executes: the block builds an AST.
Why not just use YAML?Loops, conditionals, variables and an editor come free; the cost is that a config file is now a program, which is why there is a JSON path sharing the same AST.
What goes wrong with instance_eval?All four identifier behaviours, plus the two silent failures with the actual symptoms you observed, plus the BasicObject and delegation fixes.
Dynamic languages cannot catch errors early. Discuss.Mostly true, and here is the counter-move: capture caller_locations at build time, reflect on declared or inferred parameters, and report every problem at once with file and line. Then be honest that Racket does this properly at compile time.
When is metaprogramming the wrong tool?When the names are known (use define_method), when the reader cannot follow it, when it breaks introspection, and any time the same result is achievable with a plain object. Cite the editor and tooling cost.
How do you test a DSL?Assert the shape without running; ship spies and an isolated registry so users can test too; contract tests for plugins; an invariant test that building executes nothing.
Tell me about a bug.The local variable that silently ate a verb, which caught you twice, including once after you had documented it. Then the conclusion: that is a design smell in the DSL, not only a mistake.
How would you make it production-ready?Response size limits and SSRF protection on fetch, SQLite instead of PStore, structured logging, provenance for policy changes, JSON output for CI, and an opt-in compile step for auditable artefacts.
Self-modifying pipelines: safe or not?Distinguish mutating a running pipeline from generating a successor; explain why only the second is defensible, and what idempotence and step limits protect against.
Would you use Ruby again for this?Yes for the front end, and name the four features that make it. Then name the six bugs, all of which were dynamic-language failures caught by tests, and say what you would use instead if the priority were verified correctness rather than expressiveness.

D7 · Extensions worth building


Course 2 completeWhat you built, and what it argues

A gem with twelve packages, a CLI, a plugin system, testing helpers for its users, two front ends sharing one AST, and 34 tests. More importantly, an argument you can make with evidence: the syntax is the least interesting part of a DSL. Anyone can make configuration look like a language in an afternoon with instance_eval. What makes it usable is everything that follows from refusing to execute while you parse.

Ruby's answer to the curriculum's central question, stated precisely: problems where the shape of the solution is best expressed as a small language, and where you want to grow that language at run time, from inside the program, without a build step or a parser. Ruby makes the front end nearly free and hands you the responsibility for everything a compiler would have done: checking names, reporting locations, keeping tooling informed, and testing enough to survive the absence of a type checker.

Course 5 takes the same ambition and moves it to compile time. Racket's macros turn a language definition into a function from syntax to syntax, checked before the program runs, with source locations carried automatically. When you get there, the thing to compare it against is Milestone 4's caller_locations frame counting and Part B's string-concatenating compiler. Both work. One is doing by hand what the other has as a primitive.

The Mewlang cat, viewed from behind, walking awayInstalment 10 of the five-course curriculum, and the end of Course 2. Next: Course 3, Perl, and the Text Archaeologist. Parts 0–2 first (what we are building, installation and CPAN, the language crash course), then twelve milestones turning ugly heterogeneous data into a searchable knowledge graph.

Next: Perl instalment