Milestones 5–8Wrap-up

Instalment 9 · Course 2 (Ruby) · Milestones 9–12

Testing a language, drawing it, rewriting it, and putting it in a box

The gem ships testing helpers for its own users, learns to explain and diagram itself, gains the ability to rewrite pipelines into new ones, and becomes a command you can install.

Verification note

Ruby 3.2.3. The suite finishes at 34 tests and 94 assertions, all passing, and every CLI transcript below is real output from exe/automation. Two new bugs are documented, one of which is the Milestone 6 shadowing trap catching me a second time in my own test file.

Milestone 9Testing, including testing other people's pipelines

Goal

The Mewlang cat, yawningShip an Automation::Testing module so that anyone using the gem can test their pipelines and plugins without inventing their own scaffolding. Then use it to test our own.

Concepts

Spies rather than mocks, isolated registries, assertions that speak the domain's language, testing shape without execution, and making slow behaviour (retries, backoff) fast in tests.

Design

A DSL has a testing problem its users hit immediately: to test a pipeline you must either run the real steps (slow, networked, destructive) or stub them, and stubbing requires knowing how your registry works. If you do not answer that question, every user invents a different answer and most of them will be wrong.

So testing support is a feature of the gem, not of the gem's test suite. Three pieces:

PieceJob
RecorderA step implementation that remembers its calls and returns what you told it to
test_registry and stub_stepIsolation: one registry per test, no global state to reset
Assertionsassert_steps, assert_run_failed and friends, so failures read usefully

Note the choice of spy over mock. A mock asserts expectations up front ("you will be called once with these arguments") and fails inside the code under test, producing confusing backtraces. A spy records what happened and lets you assert afterwards, in the test, where the failure message belongs. For a pipeline, where the interesting question is usually "what did the third step actually receive", the spy wins easily.

Implementation

    # Recorder is a step implementation that remembers how it was called
    # and returns whatever you told it to.
    class Recorder
      attr_reader :calls

      def initialize(result = nil, &block)
        @result = result
        @block = block
        @calls = []
      end

      def call(input, *args, **options)
        @calls << { input: input, args: args, options: options }
        @block ? @block.call(input, *args, **options) : @result
      end

      def called? = !@calls.empty?
      def call_count = @calls.size
      def last_input = @calls.last&.fetch(:input)
      def last_options = @calls.last&.fetch(:options)
    end

Twenty lines and no dependency on a mocking library. It works because the registry accepts anything with #call, a decision made back in Milestone 1 that keeps paying.

    def test_registry = @test_registry ||= Registry.new

    def stub_step(name, result = nil, spec: nil, &block)
      recorder = Recorder.new(result, &block)
      test_registry.register(name, recorder, spec: spec)
      recorder
    end

    def build_pipeline(name = "test", strict: true, &block)
      Automation.define(name, registry: test_registry, strict: strict, &block)
    end

    def assert_run_failed(result, step: nil, error: nil)
      refute result.ok?, "expected the run to fail, but it succeeded"
      assert_equal step.to_sym, result.failed_step.name if step
      assert_kind_of error, unwrap(result.error) if error
      result
    end

@test_registry ||= Registry.new gives every test its own registry, created on first use and discarded with the test instance. No global state means no teardown and no order dependence, which is worth more than any amount of careful cleanup code.

What a user's test looks like

class TestingHelpersTest < Minitest::Test
  include Automation::Testing

  def test_stubs_record_how_they_were_called
    fetch = stub_step(:fetch, [{ title: "one" }])
    # NOT `upcase = ...`: a local variable of that name would shadow the
    # verb inside the DSL block and the step would silently disappear.
    upcaser = stub_step(:upcase) { |records| records.map { |r| r.merge(title: r[:title].upcase) } }

    pipeline = build_pipeline("t") do
      fetch from: "somewhere"
      upcase
    end

    result = assert_run_ok(run_pipeline(pipeline))

    assert fetch.called?
    assert_equal({ from: "somewhere" }, fetch.last_options)
    assert_equal [{ title: "ONE" }], result.payload
    assert_equal [{ title: "one" }], upcaser.last_input
  end

  def test_assertions_about_shape_need_no_execution
    stub_step(:fetch)
    stub_step(:save_to)

    pipeline = build_pipeline("t") do
      fetch from: "x"
      save_to collection: "papers"
    end

    assert_steps %i[fetch save_to], pipeline
    assert_step_options({ collection: "papers" }, pipeline, :save_to)
    assert_pipeline_valid pipeline
  end

  def test_retry_behaviour_is_testable_without_waiting
    attempts = 0
    stub_step(:flaky) do
      attempts += 1
      raise IOError, "nope" if attempts < 3

      "ok after #{attempts}"
    end

    pipeline = build_pipeline("t") do
      retry_on IOError, times: 5, backoff: :none
      flaky
    end

    result = assert_run_ok(run_pipeline(pipeline))
    assert_equal "ok after 3", result.payload
    assert_equal 3, result.results.first.attempts
  end
end

backoff: :none is why RetryPolicy has a backoff field instead of a hard-coded schedule. A retry test that sleeps is a test people delete. Making slow behaviour configurable is a testability decision you make when you design the feature, not afterwards.

test_assertions_about_shape_need_no_execution is the one to copy into your own projects: it checks the pipeline is well-formed without running a thing, which is only possible because building produces data.

Common mistakes in Milestone 9

The Mewlang cat, winking playfullyThe shadowing bug, a second time, in my own test

My first version of the first test read:

upcase = stub_step(:upcase) { |records| ... }

pipeline = build_pipeline("t") do
  fetch from: "somewhere"
  upcase                        # the local variable, not the verb
end

The test failed with Expected: [{:title=>"ONE"}], Actual: [{:title=>"one"}]. The pipeline had one step instead of two, because upcase resolved to the local variable holding the Recorder, and evaluating a variable adds no step.

I documented this exact hazard in Milestone 6 and then walked into it again ten minutes later, which tells you how easy it is. Two conclusions. First, the naming convention matters: name the spy upcaser, fetch_spy, anything that is not the verb. Second, and more usefully, this is a design smell in the DSL, not only in the test. A language where an identifier can silently mean either a verb or a variable will bite your users. If I were shipping this seriously, build_pipeline would compare the resulting step list against the verbs mentioned in the block source, or at minimum the documentation would carry a warning with this example in it.

A duck-typing bug the helpers exposed

Adding Recorder broke four tests with undefined method 'parameters' for #<Recorder>. The validator from Milestone 4 assumed every callable answers parameters, which is true of Procs, lambdas and Methods, and false of an arbitrary object that merely has #call.

    # Anything with #call is a valid step, but only some of those things
    # can be asked about their arguments. Procs and Methods answer
    # #parameters directly; an arbitrary object is asked about its #call;
    # anything else opts out of checking rather than crashing the validator.
    def parameters_of(impl)
      return impl.parameters if impl.respond_to?(:parameters)
      return impl.method(:call).parameters if impl.respond_to?(:method)

      nil
    end

The lesson is general: if you accept a duck, check every feather you use. We advertised "anything responding to #call" and then quietly required a second method. impl.method(:call).parameters is the correct general answer, and returning nil to mean "cannot check" is better than raising, because a validator that crashes on an unusual step is worse than one that declines to judge it.

While investigating I also found a latent hazard in Milestone 7's builder-class cache, which was keyed on registry.object_id. Object ids are recycled after garbage collection, so two different registries can share one and the cache can hand out the wrong verbs. The cache now lives on the registry itself. object_id is not a durable identity, in Ruby or anywhere else with a moving or reusing allocator.

The same test in RSpec
RSpec.describe "research pipeline" do
  include Automation::Testing

  let(:fetch) { stub_step(:fetch, [{ title: "one" }]) }

  subject(:pipeline) do
    build_pipeline("research") { fetch from: "somewhere" }
  end

  it "passes the options through" do
    fetch                                   # force the let to run
    expect(run_pipeline(pipeline)).to be_ok
    expect(fetch.last_options).to eq(from: "somewhere")
  end
end

The helpers work in both because they are plain methods in a module, which is the right way to ship test support: no dependency on a framework, no RSpec.configure in your gem.

One RSpec-specific hazard worth knowing: let is lazy, so a spy defined in a let is not registered until something references it, and a pipeline built before that reference will not see the step. Minitest's eager setup has the opposite trade-off. Neither is wrong; both surprise people once.

Exercise 9

Ship a plugin contract test: a module a third-party plugin author includes to check their plugin behaves like a good citizen. It should verify that the plugin declares a step_name, that every required option is genuinely required (constructing without it and calling raises or the validator complains), that call with an empty input does not raise, that building a pipeline containing it executes nothing, and that it does not mutate its input.

Requirements: the author writes only def plugin_class = MyPlugin and def valid_options = {from: "x"}. Include a test that your own Steps::Filter passes it, and one that a deliberately bad plugin fails it.

Solution 9 — open after trying
module Automation
  module Testing
    # Include this in a test case, define plugin_class and valid_options,
    # and your plugin is checked against the contract every step must meet.
    module PluginContract
      include Automation::Testing

      def test_declares_a_step_name
        assert_kind_of Symbol, plugin_class.step_name
        refute_empty plugin_class.step_name.to_s
      end

      def test_every_required_option_is_enforced
        required = plugin_class.options_spec.select { |_, m| m[:required] }.keys

        required.each do |key|
          plugin_class.register!(test_registry)
          pipeline = build_pipeline("t", strict: false) do
            public_send(plugin_class.step_name, **valid_options.except(key))
          end
          assert_pipeline_invalid pipeline, matching: /#{key}/
        end
      end

      def test_handles_empty_input_without_raising
        plugin_class.call([], **valid_options)
      rescue StandardError => e
        flunk "#{plugin_class} raised on empty input: #{e.class}: #{e.message}"
      end

      def test_does_not_mutate_its_input
        input = [{ title: "a" }, { title: "b" }].freeze
        copy = Marshal.load(Marshal.dump(input))

        plugin_class.call(input, **valid_options)

        assert_equal copy, input, "#{plugin_class} modified the records it was given"
      end

      def test_building_a_pipeline_with_it_executes_nothing
        plugin_class.register!(test_registry)
        calls = 0
        plugin_class.singleton_class.prepend(Module.new do
          define_method(:call) { |*a, **o| calls += 1; super(*a, **o) }
        end)

        build_pipeline("t") { public_send(plugin_class.step_name, **valid_options) }

        assert_equal 0, calls
      end
    end
  end
end

class FilterContractTest < Minitest::Test
  include Automation::Testing::PluginContract

  def plugin_class = Automation::Steps::Filter
  def valid_options = { field: :topic, matching: "AI" }
end

Four things this exercise teaches beyond the code.

  • A contract test is a specification you can execute. Everything a plugin must do is written once, checked automatically, and available to people who have never read your source. This is how a plugin ecosystem stays coherent without a maintainer policing it.
  • The immutability check is the valuable one. A step that mutates the records it is given breaks the next author's pipeline in a way that is very hard to trace, and no amount of documentation prevents it. Freezing the input in the test and comparing a deep copy catches it mechanically.
  • public_send(plugin_class.step_name, ...) is how a test calls a verb whose name it does not know until run time. Inside the DSL block that is exactly what the generated methods make possible.
  • The anonymous module prepended in the last test is the counting trick from Milestone 7, used as a test instrument. Note it permanently modifies the class for the rest of the process, which is acceptable in a test suite and would not be in production code.

Experiment

In test_retry_behaviour_is_testable_without_waiting, change the stub so it needs six attempts before it succeeds (raise IOError, "nope" if attempts < 6) but leave times: 5 unchanged. Run the test: it now fails with the original IOError instead of returning "ok after 6", because times: 5 permits exactly five attempts and the middleware re-raises once attempt >= policy.times. That boundary is exactly what assert_run_failed exists to catch.

Checkpoint

  1. Why does the gem ship testing helpers rather than keeping them in its own test directory?
  2. What is the difference between a spy and a mock, and why does a spy suit pipelines?
  3. Why does each test get its own registry?
  4. Why is backoff a configurable field rather than a fixed schedule?
  5. What went wrong when the validator met a Recorder, and what is the general rule?
  6. Why is object_id a bad cache key?

Milestone 10Pipelines that describe themselves

Goal

Three functions of the AST and nothing else: explain for humans, to_mermaid for documentation, and diff for change review.

Concepts

Reflection over your own data structures, generating diagrams as text, and value equality doing the work of a diff algorithm.

Design

All three functions in this milestone take the same shape: read an AST::PipelineNode, produce a string or a value, and touch nothing else. No new state, no new classes stored anywhere, no changes to how a pipeline is built or run. That restriction is the design.

FunctionReadsProduces
explainsteps, handlers, the registry's own docsa human-readable report
to_mermaidsteps onlyMermaid flowchart text
difftwo pipelines' stepsan added/removed/changed summary

Keeping them as free functions over the AST, rather than methods that reach back into the registry, the runner or the filesystem, is what makes them trustworthy: explain cannot run a step by accident, and diff cannot be fooled by something that happened between building the two pipelines. The cost is that each function has to be told everything it needs (explain takes registry: explicitly, because it is the one function here that reports on more than the AST alone) rather than reaching for a convenient global.

Implementation

    def explain(pipeline, registry: Automation.registry)
      lines = ["#{pipeline.name} (#{pipeline.location})"]

      pipeline.steps.each_with_index do |step, i|
        known = registry.registered?(step.name)
        marker = known ? " " : "?"
        lines << format("  %s%-2d %-40s %s", marker, i + 1, step.to_s, step.location)

        doc = known ? registry.entry(step.name).doc : "UNKNOWN STEP"
        lines << "        #{doc}" if doc
      end

      pipeline.handlers.each do |handler|
        detail = handler.kind == :retry_on ? handler.callable.to_s : "a block"
        lines << format("  * %-42s %s", "#{handler.kind}: #{detail}", handler.location)
      end

      lines.join("\n")
    end
$ automation explain examples/research.rb
research (research.rb:5)
   1  fetch(from: "https://example.invalid/papers.json", limit: 20) research.rb:7
        Fetch a JSON array of records from an HTTP endpoint.
   2  filter(field: :topic, matching: "AI")    research.rb:8
        Keep records whose field matches a value or pattern.
   3  summarize(field: :abstract, max_words: 40) research.rb:9
        Summarise a field of each record into :summary.
   4  save_to(collection: "knowledge_base")    research.rb:10
        Append records to a collection in the knowledge base.
  * retry_on: retry Automation::HttpError up to 3x (exponential) research.rb:6
  * when_failed: a block                       research.rb:12

The Mewlang cat, wearing glasses, looking confidentEvery piece of that output was declared somewhere else for another reason: the step names and options by the DSL, the locations by caller_locations in Milestone 4, the documentation by the doc class macro in Milestone 7, the retry description by RetryPolicy#to_s in Milestone 6. Nothing here is a feature; it is a report over decisions already made. That is what people mean when they say a good data model pays for itself.

Drawing itself

    # Mermaid is renderable by GitHub, GitLab and most documentation tools,
    # so a pipeline can draw itself into a README.
    def to_mermaid(pipeline)
      lines = ["flowchart TD", "  start([#{pipeline.name}])"]
      previous = "start"

      pipeline.steps.each_with_index do |step, i|
        id = "s#{i}"
        label = [step.name, *step.options.map { |k, v| "#{k}=#{v}" }].join("<br/>")
        lines << "  #{id}[\"#{label}\"]"
        lines << "  #{previous} --> #{id}"
        previous = id
      end

      lines << "  #{previous} --> done([done])"
      ...
    end
flowchart TD
  start([research])
  s0["fetch<br/>from=arxiv<br/>since=7d"]
  start --> s0
  s1["filter<br/>topic=AI"]
  s0 --> s1

Generating Mermaid rather than an image means the output is text: diffable, greppable, renderable by GitHub without a build step, and editable by someone who does not have your gem installed. When you need a diagram from a program, emitting a text format that something else renders is almost always better than emitting pixels.

Diffing, for free

    def diff(before, after)
      before_by_name = before.steps.group_by(&:name)
      after_by_name = after.steps.group_by(&:name)

      added = after.steps.reject { |s| before_by_name.key?(s.name) }
      removed = before.steps.reject { |s| after_by_name.key?(s.name) }

      changed = before.steps.filter_map do |old|
        new = after_by_name[old.name]&.first
        next if new.nil?
        next if old.args == new.args && old.options == new.options

        [old, new]
      end
      ...
    end
+ deduplicate()
- filter(topic: "AI")
~ summarize(max_words: 200)  ->  summarize(max_words: 80)

This works because StepNode is a Data, so old.options == new.options compares contents. Had the nodes been ordinary objects without value equality, every comparison would be identity and every step would look changed. The diff is three lines of logic and one line of data modelling done four milestones earlier.

Note what this deliberately is not: a real diff algorithm. It matches steps by name, so a pipeline with two fetch steps will confuse it, and it reports reordering as a boolean rather than as moves. For a change-review tool that is fine; if you needed better, the right move is a proper sequence alignment (the same algorithm git diff uses), not more special cases.

Exercise 10
  1. Handle duplicates. Make diff correct for pipelines containing two steps with the same name. Decide whether to match by position, by options, or by a stable identity you add to StepNode, and write down why.
  2. A review command. Add automation diff FILE PIPELINE --against REF that loads the pipeline from a git revision (git show REF:FILE) and prints the diff. Exit non-zero when anything changed, so it can be a CI check.
  3. Cost annotations. Let a plugin declare cost :network or cost :expensive, and have explain summarise how many network calls a pipeline will make before it runs.
Solution 10 — open after trying

1. The honest answer is to give each step a stable identity at build time:

StepNode = Data.define(:id, :name, :args, :options, :location)
# id: "#{name}@#{location}", assigned by the builder

Matching by position breaks the moment someone inserts a step at the top; matching by options means a step whose options changed looks like an add plus a remove, which is exactly what a diff should avoid saying. An identity derived from the source location is stable across edits elsewhere in the file and is already available. The cost is that a rewritten step needs a new id, and you must decide whether set_options preserves it (it should) while replace does not.

This is the same problem React solves with key props and that database migrations solve with explicit identifiers, and the answer is always the same: if you want to diff a sequence, put identity in the elements rather than inferring it.

2. The interesting part is loading two versions of the same pipeline into one process. Since Automation.define registers by name, loading the old version overwrites the new one. Fix it by loading each into its own registry and pipeline table, or more simply by shelling out to a subprocess that prints to_h as JSON and comparing the data:

old_json = `git show #{ref}:#{file} | automation --format json explain -`

Preferring data over shared process state is the theme of the whole course, and it applies to your own tooling too.

3. cost :network is one more class macro storing metadata, and explain counting it is four lines. The valuable part is what it enables: a policy (Milestone 11) that refuses to define a pipeline making more than N network calls without an explicit retry_on, which is the kind of rule that is impossible to enforce with a YAML file and trivial with an inspectable AST.

Experiment

In to_mermaid, delete the *step.options.map { |k, v| "#{k}=#{v}" } part of the label so each box shows only the step name. Regenerate the diagram for examples/research.rb: every node now reads fetch, filter, summarize, save_to with nothing inside, which is enough to see the shape of a pipeline but not enough to tell which endpoint or field it touches — the reason the label includes options in the first place.

Common mistakes in Milestone 10

  • Building the Mermaid label with a naive string join and forgetting to escape ". An option value containing a quote (matching: 'AI "labs"') closes the node's quoted label early: s0["filter<br/>matching=AI "labs""] is not valid Mermaid, and the diagram silently fails to render with no error from this Ruby code at all. The fix is v.to_s.gsub('"', """) before interpolating, the same escaping discipline the migration script for this very site uses on JSX-significant characters.
  • Calling explain with the wrong registry. The default argument is registry: Automation.registry, the global one; pass a test_registry from Milestone 9 and every step reports UNKNOWN STEP, because explain only knows what the registry you gave it knows.
  • Assuming diff is a real diff. It matches by name, so a pipeline with two fetch steps produces a nonsense comparison. Read Solution 10.1 before relying on it for anything with duplicate step names.

Checkpoint

  1. Where does every piece of the explain output originally come from?
  2. Why emit Mermaid text rather than an image?
  3. Why does the diff work without a diff algorithm?
  4. What breaks the diff, and what is the standard fix for that class of problem?

Milestone 11Pipelines that rewrite themselves

Goal

A rewrite API that produces new pipelines from old ones, organisation-wide policies applied to every pipeline as it is defined, and a pipeline that improves its own successor based on measurements of its last run.

Concepts

Transformation as a pure function, a second small DSL layered on the first, safety limits on generated structure, and the difference between modifying a running program and producing its replacement.

Design

"Dynamically modifiable pipelines" can mean two very different things, and the distinction matters more than the implementation.

Mutate the running pipelineProduce a successor
ConsistencyA run can change under its own feetEach run has one fixed definition
ReproducibilityLogs describe something that no longer existsEvery version is a value you can keep
FailureA bad transformation corrupts a live runA bad transformation is discarded
ComplexityRunner must handle steps appearing mid-runNone: the runner never learns about it

We take the second, and it is not a compromise. Everything people actually want from self-modifying pipelines (adaptive optimisation, feature flags, A/B variants, policy enforcement) is achievable by generating a new pipeline value, and the one thing it rules out (a step rewriting the steps after it while they are in flight) is a debugging nightmare nobody should want.

Implementation

  # Rewrite is a tiny DSL for describing changes to a pipeline. It collects
  # operations and applies them to produce a NEW pipeline; the original is
  # never touched, so a transformation that goes wrong costs nothing.
  class Rewrite
    MAX_STEPS = 500

    def step(name, *args, **options)
      AST::StepNode.new(name: name.to_sym, args: args.freeze,
                        options: options.freeze, location: "(rewritten)")
    end

    def insert_before(name, node) = record { |steps| insert_at(steps, name, node, 0) }
    def insert_after(name, node)  = record { |steps| insert_at(steps, name, node, 1) }
    def append(node)              = record { |steps| steps + [node] }
    def prepend(node)             = record { |steps| [node] + steps }
    def remove(name)              = record { |steps| steps.reject { |s| s.name == name.to_sym } }

    def replace(name, node)
      record { |steps| steps.map { |s| s.name == name.to_sym ? node : s } }
    end

    def set_options(name, **options)
      record do |steps|
        steps.map { |s| s.name == name.to_sym ? s.with(options: s.options.merge(options).freeze) : s }
      end
    end

    def apply
      steps = @operations.reduce(@pipeline.steps) { |acc, op| op.call(acc) }

      if steps.size > MAX_STEPS
        raise Error, "rewrite produced #{steps.size} steps, over the limit of #{MAX_STEPS}"
      end

      @pipeline.with_steps(steps)
    end
  end

Each operation is a lambda from a step list to a step list, collected rather than applied immediately, and then folded in apply. That buys two things: nothing happens until apply, so a rewrite can be inspected or abandoned; and the fold means a failed operation halfway through leaves the original untouched, because we were never mutating it.

MAX_STEPS is the boring safety feature that matters. A policy with a bug can append a step every time it runs, and without a limit the first symptom is memory exhaustion. Any code that generates structure needs a bound on the structure it generates.

  module AST
    PipelineNode.class_eval do
      # pipeline.rewrite { insert_before :summarize, step(:deduplicate) }
      def rewrite(&block)
        rewriter = Rewrite.new(self)
        rewriter.instance_eval(&block)
        rewriter.apply
      end
    end
  end

class_eval reopens the generated Data class to add a method, which is the class-reopening that Ruby is famous for, used here on a class we own. And instance_eval appears a second time, making a second DSL: the vocabulary inside a rewrite block is insert_before, remove, step. Once you have the technique, layering little languages becomes routine, which is a large part of why Ruby codebases look the way they do.

It works, and the original is untouched

--- rewriting: a new pipeline, the original untouched ---
[:fetch, :filter, :summarize, :save_to]
[:fetch, :deduplicate, :summarize, :save_to]
+ deduplicate()
- filter(topic: "AI")
~ summarize(max_words: 200)  ->  summarize(max_words: 80)

--- a rewrite that cannot happen leaves everything alone ---
refused: no step named :nonexistent in research
[:fetch, :filter, :summarize, :save_to]

Policies: rules for every pipeline

  # Policies are rules applied to every pipeline as it is defined: the
  # place for organisation-wide requirements ("always log", "never fetch
  # without a timeout") that individual authors should not have to repeat.
  module Policies
    def apply(pipeline)
      all.reduce(pipeline) do |acc, (name, rule)|
        result = rule.call(acc)
        unless result.is_a?(AST::PipelineNode)
          raise Error, "policy #{name} returned #{result.class}, expected a PipelineNode"
        end

        result
      end
    end
  end
Automation::Policies.register(:always_log_first) do |pl|
  pl.step_names.first == :log_start ? pl : pl.rewrite { prepend step(:log_start) }
end

Automation::Policies.register(:cap_summaries) do |pl|
  pl.find(:summarize) ? pl.rewrite { set_options :summarize, max_words: 50 } : pl
end
[:log_start, :fetch, :summarize]
{:max_words=>50}

A user wrote max_words: 5000 and got 50, and never asked for a logging step but has one. This is genuinely powerful and genuinely dangerous, so two observations.

First, the idempotence check in always_log_first is not optional. Without pl.step_names.first == :log_start ?, every re-application prepends another step, and since policies run on every define, a pipeline rewritten and re-registered grows without bound. MAX_STEPS would eventually stop it, which is a crash rather than an answer. A policy should be a function whose second application changes nothing.

Second, a policy that silently changes what a user wrote is a debugging trap. explain mitigates it (the injected step shows a location of (rewritten)), and a serious version would record which policy made each change and print that. Governance you cannot see is governance people fight.

A pipeline that improves its successor

result = Automation.run(measured)

slow = result.results.select { |r| r.seconds > 0.01 }.map { |r| r.step.name }
successor = slow.reduce(measured) do |acc, name|
  acc.rewrite { insert_before name, step(:cache, for: name) }
end
slow steps: [:slow_step]
[:log_start, :fetch, :cache, :slow_step]
{:for=>:slow_step}

The Mewlang cat, raising a paw in celebrationThe run produced measurements; the measurements produced a transformation; the transformation produced a new pipeline. Nothing was mutated, the original measured is still valid and still describes the run that happened, and the successor can be inspected, diffed, reviewed, persisted or thrown away.

That loop (observe, decide, generate a new version) is the honest form of "a program that modifies itself", and it is the form used by query planners, JIT compilers and autoscalers. The fantasy version, where code edits itself in place while running, is not what any of those systems actually do.

Exercise 11
  1. Fixpoint and idempotence. Apply policies repeatedly until the pipeline stops changing, with a limit of, say, 10 rounds. If the limit is hit, raise an error naming the policies that are still changing things. Then add a test that a deliberately non-idempotent policy is caught rather than looping.
  2. Provenance. Record which policy or rewrite introduced each step, and show it in explain as (added by :always_log_first).
  3. Opt out. Let a pipeline declare skip_policy :cap_summaries, and decide (and document) whether policies should be skippable at all.
Solution 11 — open after trying
    MAX_ROUNDS = 10

    def apply(pipeline)
      MAX_ROUNDS.times do
        after = all.reduce(pipeline) { |acc, (name, rule)| check(name, rule.call(acc)) }
        return after if after == pipeline   # value equality: a fixpoint

        pipeline = after
      end

      culprits = all.reject { |_, rule| rule.call(pipeline) == pipeline }.keys
      raise Error, "policies did not settle after #{MAX_ROUNDS} rounds; " \
                   "not idempotent: #{culprits.join(', ')}"
    end

after == pipeline is the whole trick, and it only works because the AST is made of Data objects with value equality. A fixpoint loop over mutable objects would need a hand-written comparison, and would probably get it wrong.

Naming the culprits in the error is what turns a frustrating message into a fixable one: applying each policy once more and reporting which ones still change something identifies the offender exactly.

Provenance means adding a field to StepNode (added_by, defaulting to nil) and having Rewrite stamp it. The interesting decision is that Rewrite does not know the policy's name, so Policies.apply must pass it in, which means rewrite needs an optional source: argument. That ripple is typical: provenance is cheap to add at the start and awkward to retrofit, which is an argument for putting it in from the beginning of any system that transforms user input.

Opting out deserves a real answer rather than a feature. If policies exist to enforce organisational requirements (audit logging, rate limits, cost caps), a per-pipeline opt-out defeats them, and the correct design is that opting out is itself a policy decision: a list of exemptions held by whoever owns the policy, not a keyword any author can write. If policies exist merely as helpful defaults, an opt-out is fine. Deciding which kind you have is the design work; the code either way is five lines.

Experiment

Lower Rewrite::MAX_STEPS to 3 and re-run the deduplicate rewrite shown earlier, which turns the four-step research pipeline into five. apply now raises Error, "rewrite produced 5 steps, over the limit of 3" instead of returning a new pipeline. Then set the limit back, write a policy with no idempotence check that appends a step every time it runs, and call Automation.define on the same pipeline name three times in a row: instead of the step count growing without bound, it stops dead at the same MAX_STEPS error, which is the crash-rather-than-an-answer failure mode the Design section above warns about.

Common mistakes in Milestone 11

  • Non-idempotent policies. Steps multiply on every define.
  • Mutating the step array inside a rewrite operation instead of returning a new one. FrozenError if you froze, silent corruption if you did not.
  • No bound on generated structure. Add the limit before you need it.
  • Rewriting during a run. Generate a successor instead.
  • Invisible transformations. If explain cannot show that a policy changed something, users will assume the tool is broken.
  • Policies that return something other than a pipeline. Check the type and say which policy misbehaved; a NoMethodError three frames away is not a useful report.

Checkpoint

  1. Why does Rewrite collect operations and apply them at the end, rather than mutating the pipeline as each one is declared?
  2. Why is producing a successor pipeline preferable to mutating the one that is currently running?
  3. What does MAX_STEPS protect against, and why is a crash an acceptable failure mode there?
  4. Why must a policy be idempotent, and what does after == pipeline rely on to detect that?
  5. Why is opting out of a policy a design decision rather than a feature you can add unthinkingly?

Milestone 12Shipping it

Goal

A gemspec with no runtime dependencies, an automation command with four subcommands and honest exit codes, a documented trust boundary between Ruby pipeline files and JSON ones, and the release mechanics.

Concepts

Gem packaging, OptionParser, exit codes as an interface, and keeping a CLI thin enough that the library remains the product.

Design

Shipping a DSL gem is three separate promises to three separate audiences, and this milestone keeps them separate rather than folding them into one file.

PromiseTo whomWhere it lives
What this gem depends on, and which Ruby it needsWhoever runs gem installThe gemspec
How to run a pipeline from a shell, with codes scripts can rely onCI jobs, cron, a human at a terminalThe CLI
Which files are safe to run and which are safe to merely readAnyone integrating untrusted pipeline definitionsload_file / load_data

The CLI in particular is deliberately thin: it parses arguments, calls straight into the library (Automation.run, Automation::Inspector.explain, the loaders), and formats the result. None of the three subsections below contain logic that is not already in the library from an earlier milestone. A CLI that accumulates its own business logic stops being testable the way CLITest below is testable, and becomes the one part of your gem nobody can exercise without shelling out.

Implementation

The gemspec

Gem::Specification.new do |spec|
  spec.name = "automation"
  spec.version = Automation::VERSION
  spec.summary = "A Ruby DSL for describing, inspecting and running automation pipelines."
  spec.homepage = "https://github.com/yourname/automation"
  spec.license = "MIT"
  spec.required_ruby_version = ">= 3.2.0"

  spec.metadata["source_code_uri"] = spec.homepage
  spec.metadata["changelog_uri"] = "#{spec.homepage}/blob/main/CHANGELOG.md"
  spec.metadata["rubygems_mfa_required"] = "true"

  spec.files = Dir["lib/**/*.rb", "exe/*", "README.md", "CHANGELOG.md", "LICENSE.txt"]
  spec.bindir = "exe"
  spec.executables = ["automation"]
  spec.require_paths = ["lib"]

  # No runtime dependencies on purpose: everything used here ships with Ruby.
end

The CLI

  class CLI
    COMMANDS = {
      "run" => :cmd_run,
      "explain" => :cmd_explain,
      "graph" => :cmd_graph,
      "list" => :cmd_list
    }.freeze

    def call(argv)
      options = { dry_run: false, safe: false, vars: {}, verbose: false }
      parser = build_parser(options)
      parser.parse!(argv)

      command = COMMANDS[argv.shift]
      return usage(parser) unless command

      file = argv.shift
      return usage(parser, "a pipeline file is required") unless file

      load_pipelines(file, options)
      send(command, argv.shift, options)
    rescue Automation::InvalidPipeline => e
      warn e.message
      1
    rescue Automation::Error => e
      warn "automation: #{e.message}"
      1
    end

Points worth copying into your own CLIs:

The Mewlang cat, giving an unimpressed side-eyeA bug worth showing: two methods called run

My first version named the entry point run(argv) and the subcommand handler run(name, options). The second definition silently replaced the first, and the delegation I had written to paper over it produced:

exe/automation:42:in `run': super: no superclass method `run' for #<Automation::CLI>

Ruby lets you redefine a method with no warning at all, and the resulting error appears somewhere unrelated. The fix was a dispatch table and cmd_ prefixes, which is what the code above shows. The general habit: when two things in one class want the same name, that is information about the design, not an inconvenience to route around.

The trust boundary, made concrete

  # Load a pipeline file. Ordinary Ruby, so it can do anything Ruby can do:
  # only run files you trust. Use load_data for anything else.
  def self.load_file(path)
    before = pipelines.keys
    Kernel.load(File.expand_path(path))
    pipelines.keys - before
  end

  # The safe path: pipelines as data, no code executed.
  def self.load_data(path)
    require "json"
    data = JSON.parse(File.read(path), symbolize_names: true)
    Array(data[:pipelines] || [data]).map { |h| from_h(h).tap { |pl| pipelines[pl.name] = pl } }
  end

Two loaders, two comments, one flag. automation run pipeline.rb executes Ruby; automation run --safe pipeline.json does not. Both produce the same AST and run through the same engine, which is the payoff for having made the AST the centre of the system.

$ automation run --safe examples/research.json
research_safe: FAILED (1 steps)
  ✗ fetch(from: "http://127.0.0.1:9/papers.json") Automation::HttpError: ...connection refused

The failure is the point: the JSON pipeline really ran, through the real steps, and failed at the network rather than at the parser.

Every command, working

$ automation list examples/research.rb
research             4 steps    research.rb:5

$ automation run examples/research.rb --dry-run
research: ok (4 steps)
  ✓ fetch(from: "https://example.invalid/papers.json", limit: 20) 0.0ms
  ✓ filter(field: :topic, matching: "AI") 0.0ms
  ✓ summarize(field: :abstract, max_words: 40) 0.0ms
  ✓ save_to(collection: "knowledge_base") 0.0ms

$ automation run examples/research.rb
fetch: GET https://example.invalid/papers.json returned transport: Failed to open TCP
       connection to example.invalid:443 (getaddrinfo: Name or service not known)
research: FAILED (1 steps)
  ✗ fetch(...) Automation::HttpError: ... after 3 attempts
$ echo $?
1

Three attempts (the retry_on in the file), the when_failed handler's warning on stderr, the result on stdout, exit code 1.

Releasing

$ bundle exec rake test           # everything green
$ bundle exec rubocop             # clean
$ # bump lib/automation/version.rb, write the CHANGELOG entry
$ git commit -am "v0.1.0" && git tag v0.1.0
$ bundle exec rake release        # builds, tags, pushes to RubyGems

rake release comes from bundler's gem tasks and does the whole sequence, refusing if the working tree is dirty. Two conventions worth following: the version constant is the single source of truth (the gemspec reads it, so they cannot disagree), and the CHANGELOG entry is written before the release, not after, because "what changed" is much easier to answer while you remember.

Semantic versioning for a DSL gem deserves a thought. Your public API is not only your Ruby methods: it is the verbs, their options and their behaviour. Renaming a step option is a breaking change even though no Ruby method signature changed. Being explicit about that in your README saves an argument later.

Exercise 12
  1. Test the CLI. Write an integration test that calls Automation::CLI.new.call(...) with captured stdout and stderr, asserting output and exit codes for: a successful dry run, a validation failure (exit 1), and a bad invocation (exit 2). Use capture_io, which minitest provides.
  2. Machine-readable output. Add --format json so run and explain emit JSON for CI consumption. Decide what the schema is and version it.
  3. A generator. Add automation init writing a starter pipeline file and a plugin skeleton, then verify that the generated files pass automation explain and the plugin contract test from Exercise 9.
Solution 12 — open after trying
class CLITest < Minitest::Test
  def setup
    Automation.pipelines.clear
    @file = File.expand_path("../examples/research.rb", __dir__)
  end

  def test_dry_run_succeeds_and_prints_every_step
    out, err = capture_io do
      assert_equal 0, Automation::CLI.new.call(["run", @file, "--dry-run"])
    end

    assert_match(/research: ok \(4 steps\)/, out)
    assert_match(/✓ fetch/, out)
    assert_empty err
  end

  def test_a_bad_invocation_is_exit_two
    _out, err = capture_io do
      assert_equal 2, Automation::CLI.new.call(["frobnicate", @file])
    end

    assert_match(/usage: automation/, err)
  end

  def test_an_unknown_pipeline_name_is_exit_one
    _out, err = capture_io do
      assert_equal 1, Automation::CLI.new.call(["explain", @file, "nope"])
    end

    assert_match(/no pipeline named "nope"/, err)
  end
end

Three things this makes possible that testing by shelling out does not: it runs in-process so coverage tools see it, it is fast enough to run on every save, and it fails with a real backtrace instead of "expected exit 0, got 1".

The reason it works is the CLI design: call returns an integer and never calls exit. Push side effects to the edges (one exit at the bottom of the executable, one place that writes to streams) and the rest becomes ordinary testable code. That principle is language-independent and it is the single most valuable thing in this milestone.

For part 2, the key decision is that JSON output must be stable: include a "schema": 1 field from the first release, because the moment a CI job parses your output, the format is an API. Nothing is more annoying than a tool that reorders its JSON keys between patch versions.

Experiment

Edit automation.gemspec so required_ruby_version reads ">= 3.4.0", then run gem build automation.gemspec && gem install ./automation-0.1.0.gem on a machine running an older Ruby. Installation refuses outright:

ERROR:  Error installing ./automation-0.1.0.gem:
	automation-0.1.0 requires Ruby version >= 3.4.0. The current ruby version is 3.3.8.

Nothing about your code ran; RubyGems checked the declared floor before extracting a single file. Put the version back to ">= 3.2.0" and the same install succeeds, which is the whole reason this field exists: a clear refusal at install time beats a syntax error the first time someone loads lib/automation/ast.rb and hits Data.define on a Ruby that predates it.

Common mistakes in Milestone 12

  • Two methods with the same name in one class. Ruby redefines silently; see the bug above. A dispatch table with distinct cmd_ method names avoids the collision by construction.
  • Forgetting required_ruby_version. Without it, an unsupported Ruby fails with a raw syntax error deep in your library instead of a clear refusal from gem install.
  • Writing results and errors to the same stream. Send results to $stdout and errors to $stderr via warn, or automation explain x.rb | less mixes the two.
  • Calling exit from inside library or command methods. It makes the CLI untestable in-process; return an integer instead and let one line at the bottom of the executable call exit.
  • Treating load_file and load_data as interchangeable. Only the JSON path is safe for input you did not write yourself; a --safe flag that quietly falls back to Kernel.load defeats the entire trust boundary.

Checkpoint

  1. Why does the CLI return exit codes rather than calling exit directly?
  2. What do the three exit codes mean, and who consumes them?
  3. Why does the gemspec read the version from a constant instead of writing it twice?
  4. What does rubygems_mfa_required protect against?
  5. For a DSL gem, what counts as a breaking change?
  6. What is the practical difference between load_file and load_data?
Why are we using this language here?

Milestone 11 is the last strong Ruby argument in this course. Two DSLs in one system (one for describing pipelines, one for rewriting them), both built from instance_eval and both under thirty lines, with value objects giving fixpoint detection for free. Writing that in a static language means a builder API, an explicit visitor, and generated equality; Ruby lets it stay small enough that a reader can hold it all at once.

Milestones 9, 10 and 12 are Ruby being solid rather than special. The testing helpers rely on duck typing being real, which is nice; the inspector is plain data transformation; the CLI and gemspec are conventional and pleasant. A Python or TypeScript version of these three would look similar and be about as good.

Course-wide honesty: the total bug count I hit while writing these twelve milestones is six, and every one of them was a dynamic-language failure. A shadowed local eating a verb (twice), an exception chain that was never recorded, a prepended method whose self I mistook, instance variables vanishing inside instance_eval, a duck that lacked a feather I used, and a method silently redefined by a second definition of the same name. None would have compiled in Go. All were caught by tests, in seconds, and the fixes were small. That trade, more errors caught later but caught cheaply, with far more expressive power in exchange, is what choosing Ruby actually means.

Repository state after Milestone 12

automation/
├── automation.gemspec        no runtime dependencies
├── Rakefile                  rake test, rake build, rake release
├── CHANGELOG.md              keep-a-changelog, written before release
├── exe/automation            run | explain | graph | list
├── examples/
│   ├── research.rb           a Ruby pipeline (executes code)
│   └── research.json         the same shape as data (executes none)
├── lib/automation/
│   ├── errors, registry, plugin, steps            the vocabulary
│   ├── ast, define, validator, transform          the language
│   ├── context, middleware, retry_policy, runner  the engine
│   ├── inspector                                  explain, mermaid, diff
│   ├── config, http, store, summarizers           the adapters
│   └── testing                                    helpers for your users
└── test/                     34 tests, 94 assertions
$ rake test
34 runs, 94 assertions, 0 failures, 0 errors, 0 skips

The Mewlang cat, happy and celebratingInstalment 9 of the five-course curriculum. Next, and last for Ruby: the advanced phase (refinements, lazy enumerators, Ractors, contract testing, performance), the final challenge with acceptance criteria and a withheld solution, the full knowledge check, and the README, portfolio and interview material. Then Course 3 begins: Perl, and the text archaeologist.

Continue