Milestones 1–4

Instalment 6 · Course 2 (Ruby) · Parts 0–2

A configuration file that is secretly a program

Go was about making concurrency visible. Ruby is about making the shape of a problem visible, by growing the language toward it until the solution reads like a description of itself.

Verification note

Every Ruby example below was executed on Ruby 3.2.3 and the outputs are copied from those runs. The gem and bundler commands in Part 1 are standard, but my sandbox has no access to rubygems.org, so those specific commands are the one thing here I could not run. Everything else you can trust literally.

Course 2 · Part 0What are we building?

The final result

The Mewlang cat, looking up with curiosityA gem called automation. By the end of the course, someone who has never seen your source can write this file:

Automation.define do
  pipeline "research" do
    fetch      "papers", from: "arxiv:cs.AI", since: "7d"
    filter     topic: "AI", min_citations: 5
    summarize  with: :local_model, max_words: 200
    save_to    "knowledge_base"

    retry_on Timeout::Error, times: 3, backoff: :exponential

    when_failed do |error, step|
      notify "me", subject: "#{step.name} failed: #{error.message}"
    end
  end
end

and then drive it from Ruby or the command line:

$ automation run research --dry-run
research (4 steps, 1 failure handler)
  ✓ fetch      papers from=arxiv:cs.AI since=7d      [dry run: would fetch ~40 items]
  ✓ filter     topic=AI min_citations=5              [dry run: would keep ~12 items]
  ✓ summarize  with=local_model max_words=200        [dry run: 12 summaries]
  ✓ save_to    knowledge_base                        [dry run: would write 12 records]

$ automation run research
research: 4 steps, 12 items, 3.2s, saved to knowledge_base

But the interesting part is not that it runs. It is that the pipeline is a data structure the program can examine and change while it is running:

pipeline = Automation["research"]

pipeline.to_h          # => a plain Hash: the whole pipeline as data
pipeline.steps.map(&:name)                    # => [:fetch, :filter, :summarize, :save_to]
pipeline.insert_before(:summarize, :deduplicate)
pipeline.around(:summarize) { |step, input| time_it(step) { step.call(input) } }
pipeline.validate!     # => raises with the line number of the offending step

Automation.register(:deduplicate) do |items, **opts|
  items.uniq { |item| item[:title].downcase }
end

A user of your gem adds a new verb to your language by calling register. No parser change, no code generation, no editing your source. That is the payoff, and it is what "the Ruby program should feel almost like a separate programming language" means in practice.

Why this project is interesting

Every team eventually writes a YAML file that wants to be a program. It starts as six keys; then someone needs a conditional, so you invent if:; then a loop, so you invent for_each:; then variables, so you invent ${}; and within two years you have implemented a bad programming language with no debugger, no tests and error messages that say line 340: unexpected key. GitHub Actions, Kubernetes manifests, Ansible playbooks and CI configs are all at various stages of this journey.

Ruby offers a different deal: start with a real programming language and make it look like configuration. Your users get loops, conditionals, variables, functions and a debugger for free, because they were there all along. The cost, which we will take seriously, is that a config file that is a program can do anything a program can do.

Why Ruby in particular

Why are we using this language here?

The Mewlang cat, wearing glasses, looking confidentHonestly, for the front end of this project Ruby is close to unmatched, and it is the reason Rails, RSpec, Rake, Chef, Puppet, Homebrew, Vagrant and Fastlane all exist in Ruby rather than elsewhere. A generation of tools with a DSL at the front chose Ruby for exactly the features above.

Where Ruby is not the answer, and we will say so at the time:

  • If the pipelines come from untrusted users, an executable DSL is remote code execution with extra steps, and a restricted data format (YAML, JSON, a real parser) is the correct choice. We will build a data-only mode for this reason.
  • If you want errors before the program runs, Ruby cannot help much: a typo in a step name is discovered when that line executes, which may be twenty minutes into a pipeline. Racket's macros do this work at compile time with source locations, which is Course 5, and the contrast is the sharpest in the whole curriculum.
  • If the work is CPU-bound, Ruby is slow and its default interpreter has a global VM lock. Our steps are I/O-bound (HTTP, disk, model calls) so this barely matters, but it would matter if summarising meant running the model yourself.
  • Python can get perhaps 80% of the way with decorators and context managers, and has better data and ML libraries. If your steps are mostly pandas, use Python and accept a clumsier front end.

Architecture we are building toward

   user's pipeline file  (ordinary Ruby, looks like configuration)
              │
              │  blocks + instance_eval
              ▼
   ┌────────────────────┐
   │   DSL surface      │   Automation.define, pipeline, step verbs
   │   (Builder)        │   knows nothing about how steps run
   └─────────┬──────────┘
             │  builds, never executes
             ▼
   ┌────────────────────┐
   │   AST              │   immutable Data nodes:
   │   Pipeline/Step/   │   Pipeline(name, steps, handlers)
   │   Handler          │   Step(name, args, options, source_location)
   └─────────┬──────────┘
             │
             ├──────────────► Validator      unknown steps, bad options,
             │                               reported with file:line
             ├──────────────► Inspector      to_h, to_dot, dry run, diff
             │
             ▼
   ┌────────────────────┐
   │   Runner           │   walks the AST, threads a Context through,
   │   (interpreter)    │   applies middleware, handles failure
   └─────────┬──────────┘
             │  looks up by name
             ▼
   ┌────────────────────┐        ┌──────────────────────────────┐
   │   Step registry    │───────►│  Step implementations         │
   │   (plugins)        │        │  fetch, filter, summarize,    │
   └────────────────────┘        │  save_to, notify, user gems   │
                                 └──────────────┬───────────────┘
                                                ▼
                                 adapters: HTTP, filesystem, SQLite,
                                           summariser, notifier

The single most important line in that diagram is "builds, never executes". The DSL's only job is to turn a block into an AST. Everything else operates on the AST. That separation is what makes dry runs, validation, visualisation, instrumentation, self-modification and testing possible, and a DSL that executes as it parses can have none of them.

The twelve milestones

#MilestoneWhat it teaches
1A gem skeleton and a Step objectobjects, methods, bundler, project layout, irb
2pipeline "x" do ... end that runsblocks, yield, procs and lambdas, the builder pattern
3Bare verbs inside the blockinstance_eval, self, scope gates, what a clean DSL costs
4Build an AST instead of executingimmutable value objects, separating description from action
5The execution enginethe interpreter pattern, a Context, step results, logging
6when_failed, retries, timeoutsexception hierarchies, ensure, retry, recoverable design
7Steps registered at run timedefine_method, method_missing, respond_to_missing?
8Steps that do real workHTTP, files, SQLite, pluggable summarisers, configuration
9Testing a DSLminitest and RSpec, doubles, a DSL for testing the DSL
10Self-inspectionreflection, to_h, dry runs, graph output, instrumentation
11Pipelines that rewrite themselvesAST transformation at run time, middleware, safety limits
12Ship it as a gemgemspec, semantic versioning, a CLI, docs, publishing

What you will know afterwards

How to design an internal DSL that reads like a language and remains debuggable; why the AST is the load-bearing idea rather than the clever syntax; how Ruby's object model actually works (singleton classes, ancestor chains, method lookup) rather than as folklore; when metaprogramming is the right tool and the specific ways it makes code unmaintainable; and how to package and publish a gem other people can extend.


Course 2 · Part 1Install and first program

What Ruby is

Ruby is a dynamically typed, garbage-collected, object-oriented language that runs on an interpreter. Unlike Go, nothing is compiled ahead of time and there is no type checking before execution: a misspelled method name is discovered when that line runs. In exchange you get a language where a program can inspect and modify itself while running, which is the entire basis of this project.

"Dynamically typed" here means types belong to values, not to variables, and are checked when an operation happens. "Object-oriented" in Ruby is stronger than in most languages: everything is an object, including integers, nil, classes themselves, and blocks of code once you capture them.

The standard implementation is CRuby (also called MRI). Alternatives exist (JRuby on the JVM, TruffleRuby) and none of them matter for this course.

Installing

You want Ruby 3.2 or newer, because we use Data.define which arrived in 3.2. Ruby releases on Christmas Day each year; as of my knowledge the current release is 3.4 and 3.5 is likely out by now. Anything from 3.2 up works for everything here. I verified this instalment on 3.2.3.

Do not use the Ruby your operating system ships with. macOS's system Ruby is old and partly restricted; Linux distribution packages lag; both make you use sudo to install gems, which is a mess you will regret. Use a version manager: it installs Rubies in your home directory, lets you have several, and lets a project pin one with a .ruby-version file.

macOS and Linux (recommended: mise, or rbenv)
# mise (formerly rtx): fast, handles many languages, my default suggestion
curl https://mise.run | sh
mise use --global ruby@3.4
ruby -v

# or rbenv, the long-established option
brew install rbenv ruby-build          # macOS
# (on Linux: git clone rbenv and ruby-build, or use your package manager)
rbenv install 3.4.1
rbenv global 3.4.1
echo 'eval "$(rbenv init - bash)"' >> ~/.bashrc   # or zsh

On Linux you also need build tools for gems with C extensions (which SQLite and others use):

sudo apt install build-essential libssl-dev libyaml-dev zlib1g-dev libffi-dev  # Debian/Ubuntu
sudo dnf install gcc make openssl-devel libyaml-devel zlib-devel libffi-devel  # Fedora
Windows
:: RubyInstaller with the DevKit, via winget
winget install RubyInstaller.Ruby.3.4

:: then, when prompted, run the MSYS2 setup (option 3) so native gems build
ridk install

Native Windows Ruby works well these days. That said, Ruby's culture is deeply Unix-shaped (file paths, shelling out, signals, gems that assume a POSIX environment), and if you are also doing the Perl and Erlang courses, WSL2 will save you more time overall than it costs to set up.

Verify
$ ruby -v
ruby 3.2.3 (2024-01-18 revision 52bb2ac0a6) [x86_64-linux-gnu]
$ gem -v
3.4.20
$ irb -v

The toolchain

ToolWhat it doesGo equivalent
rubyRuns a filego run
irbInteractive shell; your main exploration tool(none)
gemInstalls packagesgo get
bundlerResolves and locks a project's dependency setgo.mod/go.sum
rakeTask runner (build, test, release)make
rubocopLinter and formattergofmt + go vet
minitest / rspecTesting; minitest ships with Rubygo test
riOffline documentationgo doc

Note what is different from Go: bundler is not built in (it ships with modern Ruby but is a separate tool), there is no standard formatter that everyone agrees on, and testing has two mainstream frameworks. Ruby's ecosystem has more choice and less consensus than Go's, which is either freedom or fragmentation depending on your temperament.

Creating the project

Bundler generates a gem skeleton, which is the right starting point even though we will not publish for eleven milestones:

$ gem install bundler                     # usually already present
$ bundle gem automation --test=minitest --no-coc --no-mit
Creating gem 'automation'...
      create  automation/Gemfile
      create  automation/lib/automation.rb
      create  automation/lib/automation/version.rb
      create  automation/test/test_automation.rb
      create  automation/automation.gemspec
      create  automation/Rakefile
      create  automation/README.md
      ...
$ cd automation
$ bundle install

(Use --test=rspec if you prefer RSpec. I will use minitest in the examples because it ships with Ruby and needs no installation, and I will show the RSpec equivalent in Milestone 9 where testing gets serious.)

automation/
├── Gemfile                     development dependencies
├── Gemfile.lock                the resolved versions (commit this)
├── automation.gemspec          the package definition
├── Rakefile                    tasks: rake test, rake build, rake release
├── bin/
│   ├── setup                   one-command development setup
│   └── console                 irb with your gem already loaded
├── lib/
│   ├── automation.rb           the entry point everything requires
│   └── automation/
│       └── version.rb          VERSION constant, read by the gemspec
└── test/
    ├── test_helper.rb
    └── test_automation.rb

Three conventions here are enforced by tooling rather than by taste:

Editor setup

Set up RuboCop early (bundle add rubocop --group development) and run it with bundle exec rubocop -a to autocorrect. It is opinionated and configurable via .rubocop.yml; unlike gofmt, you will end up tuning it, and the tuning is a team decision rather than a universal one.

Expect less from your editor than in Go

In a dynamic language, "go to definition" is a guess. When we start defining methods at run time with define_method and method_missing, your editor will not know those methods exist, and neither will your linter. That is a real and permanent cost of the techniques this course teaches, and Milestone 7 discusses how to mitigate it (documenting dynamic methods, respond_to_missing?, and knowing when a plain define_method beats a clever method_missing).

Hello, pipeline

Put this in hello.rb anywhere:

# frozen_string_literal: true

module Automation
  VERSION = "0.0.1"

  def self.greet(name = "world")
    "hello, #{name} (automation #{VERSION})"
  end
end

puts Automation.greet
puts Automation.greet("pipelines")
$ ruby hello.rb
hello, world (automation 0.0.1)
hello, pipelines (automation 0.0.1)

Every line, explained

# frozen_string_literal: true — a magic comment, read by the interpreter rather than ignored. It makes every string literal in this file frozen (immutable), which saves memory (identical literals are shared) and prevents a class of bug where one part of the program mutates a string another part is holding. Put it at the top of every file you write. It has one surprising consequence you will meet in Part 2.

module Automation — a module is a namespace and a bag of methods. Unlike a class, it cannot be instantiated. Top-level constant names must be capitalised; Automation::Runner is how you refer to something nested inside.

VERSION = "0.0.1" — a constant, by virtue of the capital letter. Ruby will warn if you reassign it but will not stop you, which tells you something about the language's attitude: it gives you guidance, not guarantees.

def self.greet(name = "world") — defines a method on the module itself rather than on instances, so you call it as Automation.greet. The self. prefix is how Ruby spells what other languages call static. name = "world" is a default argument.

"hello, #{name} ..." — string interpolation. #{} evaluates any Ruby expression and calls to_s on the result. Only double-quoted strings interpolate; single-quoted ones are literal.

There is no return. A Ruby method returns the value of its last expression. Explicit return exists and is used for early exits, but writing it at the end of a method marks you as a visitor from another language.

puts — a method from Kernel, which is mixed into every object, so it is callable anywhere without a receiver. It prints with a newline. Its cousins: print (no newline), p (prints inspect, the debugging representation, and returns its argument), and pp (pretty-prints nested structures). p is the one you will use most while developing.

irb, your most important tool

$ irb
irb(main):001> require_relative "hello"
=> true
irb(main):002> Automation.greet("irb")
=> "hello, irb (automation 0.0.1)"
irb(main):003> Automation.methods(false)
=> [:greet]
irb(main):004> "hello".methods.size
=> 187
irb(main):005> 5.class.ancestors
=> [Integer, Numeric, Comparable, Object, Kernel, BasicObject]

Go development is edit-compile-run. Ruby development is a conversation: you keep an irb session open, load your code, poke at objects, and ask them what they can do. obj.methods, obj.class, obj.instance_variables and Klass.ancestors are how you learn an unfamiliar library, often faster than reading its documentation. Inside your project, bin/console starts irb with the gem already loaded.

Documentation: ri Array#map in a terminal, rubydoc.info for gems, and ruby-doc.org for the core library. In irb, help "Array#map" works too.

Running and testing

ruby lib/automation.rb            # run a file
bundle exec rake test             # run the test suite
ruby -Ilib test/test_step.rb      # run one test file directly
ruby -Ilib test/test_step.rb -n test_value_equality   # one test
bundle exec rubocop -a            # lint and autocorrect
bin/console                       # irb with your gem loaded

bundle exec runs a command with exactly the gem versions your Gemfile.lock pins. Without it you get whatever is installed globally, which is how "works on my machine" happens in Ruby. Get into the habit.

A first test, in test/test_step.rb:

require "minitest/autorun"

class Step
  attr_reader :name, :options

  def initialize(name, **options)
    @name = name.to_sym
    @options = options.freeze
  end

  def ==(other)
    other.is_a?(Step) && name == other.name && options == other.options
  end
end

class StepTest < Minitest::Test
  def setup
    @step = Step.new("fetch", source: "arxiv")
  end

  def test_name_is_always_a_symbol
    assert_equal :fetch, @step.name
    assert_equal :fetch, Step.new(:fetch).name
  end

  def test_options_are_frozen
    assert @step.options.frozen?, "options should be frozen"
    assert_raises(FrozenError) { @step.options[:source] = "elsewhere" }
  end

  def test_value_equality
    assert_equal Step.new(:fetch, source: "arxiv"), @step
    refute_equal Step.new(:fetch, source: "other"), @step
  end
end
$ ruby test/test_step.rb
Run options: --seed 27134

# Running:
...

Finished in 0.001256s, 2388.1814 runs/s, 4776.3627 assertions/s.
3 runs, 6 assertions, 0 failures, 0 errors, 0 skips

That is real output. Points to notice: setup runs before every test method; test methods must start with test_; the argument order is assert_equal expected, actual (getting it backwards produces confusing failure messages); and assert_raises takes a block, which is your first glimpse of how much of Ruby's design rests on blocks.

Exercise 1

Create the gem skeleton with bundle gem automation. Then add a module method Automation.env that returns a Hash describing the runtime: Ruby version, platform, the gem's own version, and whether $PROGRAM_NAME suggests it is running under a test. Print it nicely from a script. Then write a test asserting the Hash has exactly the keys you expect.

Hints: RUBY_VERSION, RUBY_PLATFORM and $PROGRAM_NAME are globals; Hash#keys; assert_equal on sorted arrays of symbols.

Solution 1 — open after trying
# lib/automation.rb
# frozen_string_literal: true

require_relative "automation/version"

module Automation
  def self.env
    {
      ruby: RUBY_VERSION,
      platform: RUBY_PLATFORM,
      automation: VERSION,
      testing: $PROGRAM_NAME.include?("test")
    }
  end
end
# test/test_env.rb
require "minitest/autorun"
require "automation"

class EnvTest < Minitest::Test
  def test_env_has_expected_keys
    assert_equal %i[automation platform ruby testing], Automation.env.keys.sort
  end

  def test_ruby_version_is_a_string
    assert_kind_of String, Automation.env[:ruby]
    assert_match(/\A\d+\.\d+/, Automation.env[:ruby])
  end
end

Three details worth absorbing. %i[a b c] is shorthand for an array of symbols, and its sibling %w[a b c] makes an array of strings; both appear constantly in real Ruby. require_relative is for files inside your own project (resolved relative to the current file) while require searches the load path and is for gems and standard library. And assert_kind_of rather than assert_equal String, x.class, because the former accepts subclasses, which is usually what you mean.

Common first-day errors
  • The Mewlang cat, giving an unimpressed side-eyecannot load such file -- automation (LoadError) — lib/ is not on the load path. Run with ruby -Ilib ..., or use require_relative, or run through rake/bundle exec, which set it up for you.
  • undefined method 'greet' for Automation:Module — you wrote def greet instead of def self.greet, so it is an instance method on a module that has no instances.
  • uninitialized constant Automation::Runner — the file defining it was never required. Ruby does not autoload by convention (Rails adds that); you must require it.
  • can't modify frozen String — the magic comment is doing its job. Use +"literal" or String.new or, better, stop mutating strings.
  • syntax error, unexpected end-of-input — a missing end. Ruby cannot tell you where you meant to put it, only where the file ran out. Consistent indentation and a good editor are the defence.
  • Using sudo gem install. If you need sudo, your Ruby installation is wrong. Fix the version manager instead.

Checkpoint

  1. Why should you not use the system Ruby?
  2. What is the difference between require and require_relative?
  3. What does bundle exec protect you from?
  4. Where must Automation::Runner live, and what enforces that?
  5. What does the frozen_string_literal magic comment do, and why would you want it?
  6. What is the difference between puts, print, p and pp?

Course 2 · Part 2Language crash course

Only what the DSL needs, which turns out to be most of Ruby's object model and none of its web frameworks. Every output below is copied from an actual run. Keep an irb session open and type the examples.

2.1 Everything is an object, and everything is a method call

p 5.class, "x".class, nil.class, (1..3).class, :sym.class, Integer.class
p 1.+(2), 5.between?(1, 10), nil.to_a, nil.to_s.empty?
Integer
String
NilClass
Range
Symbol
Class
3
true
[]
true
Typical language vs Ruby

In Java, int is not an object and 5.toString() is a syntax error. In Python, 5 is an object but + dispatches through __add__ and classes are instances of type in a way you mostly ignore. Ruby has no exceptions to the rule, which sounds like trivia until you realise that "send a message to an object" being the only mechanism is what lets you intercept all of it. A DSL in Ruby is, at bottom, a set of objects that respond interestingly to messages.

2.2 Symbols, strings, and frozen literals

a, b = "name", "name"
p a.equal?(b), a == b, a.frozen?
p :name.equal?(:name), :name.to_s, "name".to_sym
false      # two separate String objects (in a file without the magic comment)
true       # with the same contents
false
true       # :name is always the same object
"name"
:name

A symbol is an interned, immutable name. :fetch written twice is the same object; "fetch" written twice is two objects. Use symbols for identifiers (method names, hash keys, step names, states) and strings for text (messages, content, user data). Our DSL will normalise every step name to a symbol, which is why Step.new("fetch").name returned :fetch earlier.

The frozen-literal surprise

The Mewlang cat, winking playfullyRun the same snippet in a file that starts with # frozen_string_literal: true and the first line prints true: identical frozen literals are deduplicated into one object. So equal? (object identity) gives different answers depending on a comment at the top of the file. This is worth knowing before it confuses you at 2am. The lesson is not to avoid the magic comment; it is to use == for comparisons and equal? essentially never.

2.3 Truthiness

p [0, "", [], nil, false].map { |v| v ? "truthy" : "falsy" }
["truthy", "truthy", "truthy", "falsy", "falsy"]

Only nil and false are falsy. Zero is true. The empty string is true. The empty array is true. Coming from Python or JavaScript this is the single most likely source of a wrong conditional in your first week, and it cuts both ways: if items does not check for emptiness, it checks for existence. Write if items.empty? when you mean empty.

Two idioms that follow:

@steps ||= []              # assign only if currently nil or false
name = opts[:name] || "unnamed"
count = config&.limit      # safe navigation: nil if config is nil, no NoMethodError

2.4 Methods and their arguments

def describe(name, *rest, retries: 3, **opts)
  "#{name} retries=#{retries} rest=#{rest.inspect} opts=#{opts.inspect}"
end

puts describe("fetch", retries: 5, timeout: 2)

def double(x) = x * 2          # endless method (Ruby 3.0+)
p double(21)
fetch retries=5 rest=[] opts={:timeout=>2}
42

Naming conventions that Ruby programmers treat as meaning, not decoration:

FormMeansExample
name?Returns a booleanempty?, valid?
name!Dangerous: mutates, or raises where the plain version does notsort!, validate!
name=A setter, called as obj.name = xattr_writer generates these
snake_caseMethods and variablessave_to
CamelCaseClasses and modulesStepRegistry
SCREAMINGConstantsVERSION

2.5 Arrays, hashes, and iterating

steps = %w[fetch filter summarize save]
p steps.map(&:upcase)
p steps.select { |s| s.length > 4 }
p steps.each_with_index.map { |s, i| "#{i}:#{s}" }
p steps.reduce("") { |acc, s| acc + s[0] }

config = { topic: "AI", limit: 10 }
p config[:topic], config[:missing], config.fetch(:limit), config.key?(:topic)

nested = { pipeline: { steps: [{ name: "fetch" }] } }
p nested.dig(:pipeline, :steps, 0, :name)
p steps.each_with_object({}) { |s, h| h[s.to_sym] = s.length }
["FETCH", "FILTER", "SUMMARIZE", "SAVE"]
["fetch", "filter", "summarize"]
["0:fetch", "1:filter", "2:summarize", "3:save"]
"ffss"
"AI"
nil
10
true
"fetch"
{:fetch=>5, :filter=>6, :summarize=>9, :save=>4}
Typical language vs Ruby

Python has one iteration protocol and a handful of builtins (map, filter, comprehensions). Ruby has Enumerable, a module with about sixty methods, mixed into anything that defines each. That means group_by, partition, each_slice, tally, sum, min_by, flat_map, zip, lazy and more are available on your own classes the moment you define each and include the module. Learning Enumerable is the single highest-value hour you can spend on Ruby's standard library.

2.6 Blocks

A block is a chunk of code attached to a method call. It is not an argument in the ordinary sense: every Ruby method can take one, and the method decides whether to use it.

def each_step
  return to_enum(:each_step) unless block_given?
  yield "fetch"
  yield "filter"
  :done
end

each_step { |s| print s, " " }
p each_step.to_a

def with_logging(name, &block)
  puts "start #{name}"
  result = block.call
  puts "end #{name}"
  result
end

p with_logging("run") { 40 + 2 }
fetch filter 
["fetch", "filter"]
start run
end run
42

Why blocks matter more in Ruby than closures do elsewhere. Because every method can take exactly one block with no ceremony, Ruby programmers habitually write methods that take a chunk of behaviour: File.open(path) { |f| ... } closes the file afterwards, transaction { ... } commits or rolls back, assert_raises(E) { ... } checks an expectation. The pattern is "the method controls the resource, the block supplies the intent", and our whole DSL is one enormous application of it.

2.7 Procs and lambdas

pr = proc   { |a, b| [a, b] }
la = ->(a, b) { [a, b] }

p pr.call(1), pr.call(1, 2, 3), pr.lambda?, la.lambda?
begin
  la.call(1)
rescue ArgumentError => e
  puts "lambda strict: #{e.message}"
end

def returns_from_lambda
  l = -> { return :from_lambda }
  l.call
  :from_method
end

def returns_from_proc
  pr = proc { return :from_proc }
  pr.call
  :from_method
end

p returns_from_lambda, returns_from_proc
[1, nil]
[1, 2]
false
true
lambda strict: wrong number of arguments (given 1, expected 2)
:from_method
:from_proc

Two differences, and the second one is the dangerous one:

ProcLambda
ArityLenient: missing args become nil, extra ones are droppedStrict: wrong count raises ArgumentError
returnReturns from the enclosing methodReturns from the lambda only
Syntaxproc { }, or a block captured with &->(x) { } or lambda { }

Look at the output again: returns_from_proc returned :from_proc, meaning the return inside the proc terminated the whole method and the last line never ran. The lambda version returned :from_method, because its return only left the lambda. A block captured with &block is a Proc, so a user's when_failed do ... return ... end could return from surprising places. Prefer lambdas when you store user code and call it later, and we will.

Exercise 2.A

The Mewlang cat, thinking with a paw to its chinWrite a method retrying(times:, on: StandardError) that takes a block, calls it, and retries up to times attempts if the block raises an exception of the given class, re-raising if it never succeeds. It should return the block's value on success, and it should be usable as:

result = retrying(times: 3) { flaky_call }

Then extend it to yield the attempt number to the block, so the block can behave differently on a retry. Test it with a counter that fails the first two times.

Solution 2.A — open after trying
def retrying(times:, on: StandardError)
  attempt = 0
  begin
    attempt += 1
    yield attempt
  rescue on => e
    retry if attempt < times
    raise
  end
end

attempts = 0
value = retrying(times: 3) do |n|
  attempts += 1
  raise IOError, "flaky" if n < 3
  "succeeded on attempt #{n}"
end

p value, attempts
# => "succeeded on attempt 3"
# => 3

Four things in eleven lines. retry is a keyword that re-runs the begin block from the top; no loop needed, and no other mainstream language has it. A bare raise inside a rescue re-raises the current exception with its original backtrace, which is what you want; raise e also works but loses nothing only because Ruby is kind. on: defaults to StandardError rather than Exception, because Exception includes SignalException and Interrupt, so rescuing it swallows Ctrl-C. And yield attempt passes a value to the block, which the caller may ignore: blocks are lenient about arity, so retrying(times: 3) { flaky } still works.

This method is, almost unchanged, what Milestone 6 puts behind retry_on.

2.8 Classes

class Step
  attr_reader :name, :options

  def initialize(name, **options)
    @name = name.to_sym
    @options = options.freeze
  end

  def to_s = "#{@name}(#{@options.inspect})"
  def ==(other) = other.is_a?(Step) && name == other.name && options == other.options
  def call(input) = raise(NotImplementedError, "#{self.class} must implement #call")
end

s = Step.new("fetch", source: "arxiv")
puts s
p s == Step.new(:fetch, source: "arxiv"), s.options.frozen?
fetch({:source=>"arxiv"})
true
true
Typical language vs Ruby

There is no new keyword (new is a method on the class object), no field declarations (instance variables appear when assigned), no access modifiers on state (all instance variables are private, always), and no compile-time interface. Visibility applies only to methods, via private and protected, and private is itself a method call that switches a mode for the rest of the class body.

2.9 Modules: namespaces and mixins

module Loggable
  def log(msg) = puts("[#{self.class.name}] #{msg}")
end

module Registry
  def register(name) = (@registered ||= []) << name
  def registered = @registered || []
end

class Pipeline
  include Loggable     # instance methods
  extend  Registry     # class methods
end

Pipeline.new.log("hello")
Pipeline.register(:fetch)
Pipeline.register(:save)
p Pipeline.registered, Pipeline.ancestors.first(4), Pipeline.include?(Loggable)
[Pipeline] hello
[:fetch, :save]
[Pipeline, Loggable, Object, Kernel]
true

include adds instance methods; extend adds methods to the object doing the extending, which for a class means class methods. That one sentence resolves most Ruby module confusion.

ancestors shows the method lookup chain: Ruby searches Pipeline, then Loggable, then Object, then Kernel, then BasicObject, and calls the first matching method. There is no multiple inheritance, but a class can include many modules, and they stack in the chain. This is worth remembering as the honest answer to "does Ruby have multiple inheritance": no, it has a linearised chain you can inject into, which gets you the useful part without the diamond problem.

module Automation
  VERSION = "0.1.0"

  class Error < StandardError; end

  class StepError < Error
    def initialize(step, cause) = super("step #{step} failed: #{cause}")
  end
end

p Automation::VERSION, Automation::StepError.ancestors.first(3)
"0.1.0"
[Automation::StepError, Automation::Error, StandardError]

Every gem should define one base error class inside its namespace and inherit all its others from it, so a user can write rescue Automation::Error and catch everything you raise without catching everything in the world. This is the Ruby equivalent of Go's sentinel error hierarchy and it is not optional in a library.

2.10 self, and the trick the whole DSL rests on

class Builder
  def initialize = @steps = []
  def fetch(what) = @steps << [:fetch, what]
  def steps = @steps

  def build(&block)
    instance_eval(&block)   # self inside the block becomes this builder
    @steps
  end
end

p Builder.new.build { fetch "papers" }
p Builder.new.instance_eval { self.class }
p Builder.new.instance_exec(3) { |n| n * 2 }
[[:fetch, "papers"]]
Builder
6

Look carefully at the first line of output. The block { fetch "papers" } was written at the top level of the file, where no method called fetch exists. It worked because instance_eval evaluated the block with self set to the builder, so the bare call fetch "papers" was sent to the builder.

This is the whole trick. Every Ruby DSL you have ever seen (RSpec's describe/it, Rake's task, a Gemfile's gem, Sinatra's get) is this: a block, evaluated with self pointing at an object that responds to the vocabulary. Milestone 3 does it properly, including the significant problems it causes, which are: the block can no longer see the caller's private methods, method names silently collide, and typos become confusing NoMethodErrors rather than NameErrors.

instance_exec is the same thing but passes arguments to the block. Both are the sharpest tools in Ruby, and both are how you cut yourself.

2.11 Exceptions

attempts = 0
begin
  attempts += 1
  raise Automation::StepError.new(:fetch, "timeout") if attempts < 3
  puts "succeeded on attempt #{attempts}"
rescue Automation::Error => e
  retry if attempts < 3
  puts "giving up: #{e.message}"
ensure
  puts "ensure always runs (attempts=#{attempts})"
end

def risky
  yield
rescue ZeroDivisionError => e
  raise Automation::StepError.new(:divide, e.message)
end

begin
  risky { 1 / 0 }
rescue Automation::StepError => e
  puts "#{e.class}: #{e.message} / cause: #{e.cause.class}"
end
succeeded on attempt 3
ensure always runs (attempts=3)
Automation::StepError: step divide failed: divided by 0 / cause: ZeroDivisionError

2.12 Reflection and the beginnings of metaprogramming

class Ghost
  def initialize = @calls = []

  def method_missing(name, *args, &blk)
    return super unless name.to_s.start_with?("step_")
    @calls << [name, args]
    self
  end

  def respond_to_missing?(name, include_private = false)
    name.to_s.start_with?("step_") || super
  end

  def calls = @calls
end

g = Ghost.new
g.step_fetch("papers").step_save("kb")
p g.calls, g.respond_to?(:step_anything), g.respond_to?(:nope)
g.nope   # => NoMethodError
[[:step_fetch, ["papers"]], [:step_save, ["kb"]]]
true
false
NoMethodError: undefined method `nope' for #<Ghost:0x00...

method_missing is called when an object receives a message it has no method for. Three rules make the difference between a useful ghost and an unmaintainable one:

  1. Handle only what you mean to, and super for the rest. The return super unless line is what preserves the normal NoMethodError for genuine typos. Without it, every misspelling silently succeeds and your users debug ghosts.
  2. Always define respond_to_missing? alongside it. Otherwise respond_to? lies, and so do method(), duck-typing checks and every tool that introspects your object. Notice the output: respond_to?(:step_anything) is true because we defined it.
  3. Returning self makes calls chainable, which is how g.step_fetch(...).step_save(...) works.
class Dynamic
  %i[fetch filter save].each do |name|
    define_method(name) { |arg = nil| "called #{name} with #{arg.inspect}" }
  end
end

d = Dynamic.new
p d.fetch("x"), d.public_send(:filter, 2), Dynamic.instance_methods(false).sort
"called fetch with \"x\""
"called filter with 2"
[:fetch, :filter, :save]

define_method defines a real method from a block, at run time. Compare the two approaches, because Milestone 7 chooses between them constantly:

define_methodmethod_missing
Method really existsYes: shows in methods, respond_to?, documentation toolsNo: must fake it all
SpeedNormal method dispatchSlower: only runs after lookup fails
Needs the names in advanceYesNo
VerdictPrefer thisOnly when the set is genuinely open

send calls a method by name, including private ones; public_send respects visibility and is what you should use when the name comes from user input or a registry.

2.13 Value objects with Data

StepNode = Data.define(:name, :options)

n = StepNode.new(name: :fetch, options: { source: "arxiv" })
p n, n.name, n.frozen?
p n.with(name: :refetch)
n.instance_variable_set(:@name, :hacked)   # => FrozenError
#<data StepNode name=:fetch, options={:source=>"arxiv"}>
:fetch
true
#<data StepNode name=:refetch, options={:source=>"arxiv"}>
FrozenError: can't modify frozen StepNode: #<data Ste...

Data (Ruby 3.2+) creates an immutable value class with readers, value equality, a decent inspect, and with for making modified copies. This is exactly what our AST nodes need, and it arrives with no boilerplate. Its mutable older sibling is Struct, which allows positional construction and assignment; prefer Data for anything representing a description rather than a state.

Immutability in the AST is not aesthetic. In Milestone 11 pipelines rewrite themselves, and "rewrite" will mean "produce a new pipeline from the old one" rather than "mutate in place". That makes the old version still valid, makes diffs possible, and makes a failed transformation harmless.

Exercise 2.B — the capstone of Part 2

Build a miniature of the whole project in about sixty lines. Requirements:

  • A StepNode = Data.define(:name, :options).
  • A Builder that collects step nodes and whose method_missing turns any bare verb into a step, so fetch "papers", from: "arxiv" and summarize max_words: 50 both work. Include respond_to_missing?.
  • A module method Mini.pipeline(name, &block) that evaluates the block against a builder and returns a frozen PipelineNode = Data.define(:name, :steps).
  • A Runner that walks the steps and calls a handler registered by name in a Hash, threading the result of each step into the next. Unknown step names must raise a clear error naming the step and listing the known ones.
  • Tests: that the AST is built correctly without anything executing, that running threads values through, and that an unknown step raises.

The one design rule: building must not execute anything. You should be able to build a pipeline whose steps are all unregistered and only get an error when you run it.

Solution 2.B — open after trying
# frozen_string_literal: true

module Mini
  class Error < StandardError; end

  class UnknownStep < Error
    def initialize(name, known)
      super("unknown step #{name.inspect}; known steps: #{known.sort.join(', ')}")
    end
  end

  StepNode     = Data.define(:name, :options)
  PipelineNode = Data.define(:name, :steps)

  # Builder only collects. It never runs anything, which is what makes
  # dry runs, validation and rewriting possible later.
  class Builder
    attr_reader :steps

    def initialize
      @steps = []
    end

    def method_missing(name, *args, **options, &_block)
      @steps << StepNode.new(name: name, options: options.merge(args: args))
      self
    end

    def respond_to_missing?(_name, _include_private = false) = true
  end

  def self.pipeline(name, &block)
    builder = Builder.new
    builder.instance_eval(&block)
    PipelineNode.new(name: name.to_sym, steps: builder.steps.freeze).freeze
  end

  class Runner
    def initialize(handlers) = @handlers = handlers

    def run(pipeline, input = nil)
      pipeline.steps.reduce(input) do |acc, step|
        handler = @handlers.fetch(step.name) { raise UnknownStep.new(step.name, @handlers.keys) }
        handler.call(acc, **step.options)
      end
    end
  end
end
pipeline = Mini.pipeline("research") do
  fetch "papers", from: "arxiv"
  filter topic: "AI"
  summarize max_words: 50
end

p pipeline.steps.map(&:name)
# => [:fetch, :filter, :summarize]

handlers = {
  fetch:     ->(_input, args:, **) { ["paper about #{args.first}", "paper about cats"] },
  filter:    ->(items, topic:, **) { items.grep(/#{topic}/i) },
  summarize: ->(items, max_words:, **) { items.map { |i| i[0, max_words] } }
}

p Mini::Runner.new(handlers).run(pipeline)
# => ["paper about papers"]

Mini::Runner.new({}).run(pipeline)
# => Mini::UnknownStep: unknown step :fetch; known steps:

Notes on the choices, several of which are the same choices the real project makes.

  • method_missing returning self, and accepting *args, **options so both positional and keyword forms work. Folding args into the options Hash is a shortcut; the real version keeps them separate, because the validator needs to check them differently.
  • respond_to_missing? returning true unconditionally is honest here (the builder really does accept anything) and is exactly what you must not do in the real project, where the registry knows which verbs exist and a typo should be caught.
  • Hash#fetch with a block raises your own error rather than KeyError, and the message lists the known steps. Error messages that tell the user what they could have written instead are the difference between a DSL people like and one they tolerate.
  • reduce threads the accumulator through the steps, which is the entire interpreter in one line. Milestone 5 expands it into a Runner with a Context, logging and middleware, but the shape does not change.
  • Everything is frozen. Try adding a step to a built pipeline and you get a FrozenError immediately rather than a mysterious mutation later.

If you built something close to this, you have already written the skeleton of Milestones 1 through 5, and the rest of the course is about making each piece real: proper errors with source locations, a registry that validates, failure handling, plugins, inspection and packaging.

Common mistakes in Part 2
  • Assuming 0 or "" is falsy. They are not. Only nil and false.
  • Using return at the end of a method. Harmless, but it marks you as a tourist. Worse: return inside a proc exits the enclosing method.
  • hash[:key] where you meant hash.fetch(:key). The nil travels and explodes far from the cause.
  • Mutating a string literal in a file with the frozen magic comment. Use dup or, better, build new strings.
  • method_missing without respond_to_missing?, or without super for unhandled names. Both make objects that lie about themselves.
  • rescue Exception. Now Ctrl-C does not work.
  • Forgetting that class bodies are executable code. attr_reader, include and private are method calls that run when the class is defined, not declarations.
  • Argument order errors: splat must precede keyword arguments, and assert_equal takes expected first.

Part 2 checkpoint

  1. Why is Integer.class equal to Class, and why does that matter for a DSL?
  2. When would you use a Symbol rather than a String?
  3. Which values are falsy in Ruby? Name two idioms that depend on the answer.
  4. What does &:upcase do, mechanically?
  5. Give two differences between a proc and a lambda, and say which one you would store as a user-supplied callback.
  6. What is the difference between include and extend?
  7. What does instance_eval change, and why is that the foundation of Ruby DSLs?
  8. What two methods must you always define together, and what breaks if you do not?
  9. When is define_method better than method_missing?
  10. Why should the AST nodes be immutable?
Why are we using this language here? (Part 2 summary)

You have now seen the four features that make the rest of this course possible: blocks as syntax, instance_eval to redirect self, method_missing and define_method to decide the vocabulary at run time, and class bodies being ordinary executable code. No other mainstream language has all four, and it is why an entire generation of infrastructure tools put a Ruby DSL at the front.

The honest cost, which you have also seen: none of this is checked before it runs. A typo in a step name, a proc whose return escapes, a method_missing that swallows a mistake, a monkey-patched method that changes behaviour three files away. Go's compiler would have caught most of the corresponding mistakes. Ruby gives you expressiveness and hands you the responsibility for correctness, which is why the testing milestone in this course is not optional and why Milestone 4's decision to build an inspectable AST matters so much: it is how we claw back some of what the compiler would have given us.

What is next

Milestone 1 sets up the real gem, defines Step and Registry properly, and gets a test suite running. Milestone 2 introduces blocks as the DSL surface with an explicit builder argument, so you can see what instance_eval buys before we use it. Milestone 3 makes the switch and pays the price. By Milestone 5 you will have an interpreter.

Before then, two things worth doing:

  1. Finish Exercise 2.B. The next instalment assumes you have felt the shape of it.
  2. Run bundle gem automation and commit the skeleton, and spend ten minutes in irb calling .methods on things. Ruby rewards poking at it in a way that compiled languages do not.

The Mewlang cat, viewed from behind, walking awayInstalment 6 of the five-course curriculum. Next: Ruby Milestones 1–4, where the gem gets real, blocks become a DSL, instance_eval earns and costs, and the whole thing turns into an AST.

Continue