Instalment 6 · Course 2 (Ruby) · Parts 0–2
Go was about making concurrency visible. Ruby is about making the shape of a problem visible, by growing the language toward it until the solution reads like a description of itself.
Every Ruby example below was executed on Ruby 3.2.3 and the outputs are copied from those runs. The gem and bundler commands in Part 1 are standard, but my sandbox has no access to rubygems.org, so those specific commands are the one thing here I could not run. Everything else you can trust literally.
A gem called automation. By the end of the course, someone who has never seen your source can write this file:
Automation.define do
pipeline "research" do
fetch "papers", from: "arxiv:cs.AI", since: "7d"
filter topic: "AI", min_citations: 5
summarize with: :local_model, max_words: 200
save_to "knowledge_base"
retry_on Timeout::Error, times: 3, backoff: :exponential
when_failed do |error, step|
notify "me", subject: "#{step.name} failed: #{error.message}"
end
end
end
and then drive it from Ruby or the command line:
$ automation run research --dry-run
research (4 steps, 1 failure handler)
✓ fetch papers from=arxiv:cs.AI since=7d [dry run: would fetch ~40 items]
✓ filter topic=AI min_citations=5 [dry run: would keep ~12 items]
✓ summarize with=local_model max_words=200 [dry run: 12 summaries]
✓ save_to knowledge_base [dry run: would write 12 records]
$ automation run research
research: 4 steps, 12 items, 3.2s, saved to knowledge_base
But the interesting part is not that it runs. It is that the pipeline is a data structure the program can examine and change while it is running:
pipeline = Automation["research"]
pipeline.to_h # => a plain Hash: the whole pipeline as data
pipeline.steps.map(&:name) # => [:fetch, :filter, :summarize, :save_to]
pipeline.insert_before(:summarize, :deduplicate)
pipeline.around(:summarize) { |step, input| time_it(step) { step.call(input) } }
pipeline.validate! # => raises with the line number of the offending step
Automation.register(:deduplicate) do |items, **opts|
items.uniq { |item| item[:title].downcase }
end
A user of your gem adds a new verb to your language by calling register. No parser change, no code generation, no editing your source. That is the payoff, and it is what "the Ruby program should feel almost like a separate programming language" means in practice.
Every team eventually writes a YAML file that wants to be a program. It starts as six keys; then someone needs a conditional, so you invent if:; then a loop, so you invent for_each:; then variables, so you invent ${}; and within two years you have implemented a bad programming language with no debugger, no tests and error messages that say line 340: unexpected key. GitHub Actions, Kubernetes manifests, Ansible playbooks and CI configs are all at various stages of this journey.
Ruby offers a different deal: start with a real programming language and make it look like configuration. Your users get loops, conditionals, variables, functions and a debugger for free, because they were there all along. The cost, which we will take seriously, is that a config file that is a program can do anything a program can do.
pipeline "research" do ... end is ordinary method call syntax. Python's nearest equivalent needs a decorator, a with statement, or a lambda with different indentation rules; JavaScript needs () => {} and a comma. The visual noise difference is small per line and decisive over a whole file.instance_eval changes what self means inside a block. That single feature is why you can write fetch "papers" instead of config.fetch("papers") or p.fetch("papers"). It is the difference between a builder API and something that reads like a language.filter topic: "AI" is a method call that looks like a declaration.topic: "AI", min_citations: 5 is one Hash, and it reads like named fields.method_missing and define_method let the set of available verbs be decided at run time, which is what makes plugins possible without a plugin framework.
Honestly, for the front end of this project Ruby is close to unmatched, and it is the reason Rails, RSpec, Rake, Chef, Puppet, Homebrew, Vagrant and Fastlane all exist in Ruby rather than elsewhere. A generation of tools with a DSL at the front chose Ruby for exactly the features above.
Where Ruby is not the answer, and we will say so at the time:
user's pipeline file (ordinary Ruby, looks like configuration)
│
│ blocks + instance_eval
▼
┌────────────────────┐
│ DSL surface │ Automation.define, pipeline, step verbs
│ (Builder) │ knows nothing about how steps run
└─────────┬──────────┘
│ builds, never executes
▼
┌────────────────────┐
│ AST │ immutable Data nodes:
│ Pipeline/Step/ │ Pipeline(name, steps, handlers)
│ Handler │ Step(name, args, options, source_location)
└─────────┬──────────┘
│
├──────────────► Validator unknown steps, bad options,
│ reported with file:line
├──────────────► Inspector to_h, to_dot, dry run, diff
│
▼
┌────────────────────┐
│ Runner │ walks the AST, threads a Context through,
│ (interpreter) │ applies middleware, handles failure
└─────────┬──────────┘
│ looks up by name
▼
┌────────────────────┐ ┌──────────────────────────────┐
│ Step registry │───────►│ Step implementations │
│ (plugins) │ │ fetch, filter, summarize, │
└────────────────────┘ │ save_to, notify, user gems │
└──────────────┬───────────────┘
▼
adapters: HTTP, filesystem, SQLite,
summariser, notifier
The single most important line in that diagram is "builds, never executes". The DSL's only job is to turn a block into an AST. Everything else operates on the AST. That separation is what makes dry runs, validation, visualisation, instrumentation, self-modification and testing possible, and a DSL that executes as it parses can have none of them.
| # | Milestone | What it teaches |
|---|---|---|
| 1 | A gem skeleton and a Step object | objects, methods, bundler, project layout, irb |
| 2 | pipeline "x" do ... end that runs | blocks, yield, procs and lambdas, the builder pattern |
| 3 | Bare verbs inside the block | instance_eval, self, scope gates, what a clean DSL costs |
| 4 | Build an AST instead of executing | immutable value objects, separating description from action |
| 5 | The execution engine | the interpreter pattern, a Context, step results, logging |
| 6 | when_failed, retries, timeouts | exception hierarchies, ensure, retry, recoverable design |
| 7 | Steps registered at run time | define_method, method_missing, respond_to_missing? |
| 8 | Steps that do real work | HTTP, files, SQLite, pluggable summarisers, configuration |
| 9 | Testing a DSL | minitest and RSpec, doubles, a DSL for testing the DSL |
| 10 | Self-inspection | reflection, to_h, dry runs, graph output, instrumentation |
| 11 | Pipelines that rewrite themselves | AST transformation at run time, middleware, safety limits |
| 12 | Ship it as a gem | gemspec, semantic versioning, a CLI, docs, publishing |
How to design an internal DSL that reads like a language and remains debuggable; why the AST is the load-bearing idea rather than the clever syntax; how Ruby's object model actually works (singleton classes, ancestor chains, method lookup) rather than as folklore; when metaprogramming is the right tool and the specific ways it makes code unmaintainable; and how to package and publish a gem other people can extend.
Ruby is a dynamically typed, garbage-collected, object-oriented language that runs on an interpreter. Unlike Go, nothing is compiled ahead of time and there is no type checking before execution: a misspelled method name is discovered when that line runs. In exchange you get a language where a program can inspect and modify itself while running, which is the entire basis of this project.
"Dynamically typed" here means types belong to values, not to variables, and are checked when an operation happens. "Object-oriented" in Ruby is stronger than in most languages: everything is an object, including integers, nil, classes themselves, and blocks of code once you capture them.
The standard implementation is CRuby (also called MRI). Alternatives exist (JRuby on the JVM, TruffleRuby) and none of them matter for this course.
You want Ruby 3.2 or newer, because we use Data.define which arrived in 3.2. Ruby releases on Christmas Day each year; as of my knowledge the current release is 3.4 and 3.5 is likely out by now. Anything from 3.2 up works for everything here. I verified this instalment on 3.2.3.
Do not use the Ruby your operating system ships with. macOS's system Ruby is old and partly restricted; Linux distribution packages lag; both make you use sudo to install gems, which is a mess you will regret. Use a version manager: it installs Rubies in your home directory, lets you have several, and lets a project pin one with a .ruby-version file.
# mise (formerly rtx): fast, handles many languages, my default suggestion
curl https://mise.run | sh
mise use --global ruby@3.4
ruby -v
# or rbenv, the long-established option
brew install rbenv ruby-build # macOS
# (on Linux: git clone rbenv and ruby-build, or use your package manager)
rbenv install 3.4.1
rbenv global 3.4.1
echo 'eval "$(rbenv init - bash)"' >> ~/.bashrc # or zsh
On Linux you also need build tools for gems with C extensions (which SQLite and others use):
sudo apt install build-essential libssl-dev libyaml-dev zlib1g-dev libffi-dev # Debian/Ubuntu
sudo dnf install gcc make openssl-devel libyaml-devel zlib-devel libffi-devel # Fedora
:: RubyInstaller with the DevKit, via winget
winget install RubyInstaller.Ruby.3.4
:: then, when prompted, run the MSYS2 setup (option 3) so native gems build
ridk install
Native Windows Ruby works well these days. That said, Ruby's culture is deeply Unix-shaped (file paths, shelling out, signals, gems that assume a POSIX environment), and if you are also doing the Perl and Erlang courses, WSL2 will save you more time overall than it costs to set up.
$ ruby -v
ruby 3.2.3 (2024-01-18 revision 52bb2ac0a6) [x86_64-linux-gnu]
$ gem -v
3.4.20
$ irb -v
| Tool | What it does | Go equivalent |
|---|---|---|
ruby | Runs a file | go run |
irb | Interactive shell; your main exploration tool | (none) |
gem | Installs packages | go get |
bundler | Resolves and locks a project's dependency set | go.mod/go.sum |
rake | Task runner (build, test, release) | make |
rubocop | Linter and formatter | gofmt + go vet |
minitest / rspec | Testing; minitest ships with Ruby | go test |
ri | Offline documentation | go doc |
Note what is different from Go: bundler is not built in (it ships with modern Ruby but is a separate tool), there is no standard formatter that everyone agrees on, and testing has two mainstream frameworks. Ruby's ecosystem has more choice and less consensus than Go's, which is either freedom or fragmentation depending on your temperament.
Bundler generates a gem skeleton, which is the right starting point even though we will not publish for eleven milestones:
$ gem install bundler # usually already present
$ bundle gem automation --test=minitest --no-coc --no-mit
Creating gem 'automation'...
create automation/Gemfile
create automation/lib/automation.rb
create automation/lib/automation/version.rb
create automation/test/test_automation.rb
create automation/automation.gemspec
create automation/Rakefile
create automation/README.md
...
$ cd automation
$ bundle install
(Use --test=rspec if you prefer RSpec. I will use minitest in the examples because it ships with Ruby and needs no installation, and I will show the RSpec equivalent in Milestone 9 where testing gets serious.)
automation/
├── Gemfile development dependencies
├── Gemfile.lock the resolved versions (commit this)
├── automation.gemspec the package definition
├── Rakefile tasks: rake test, rake build, rake release
├── bin/
│ ├── setup one-command development setup
│ └── console irb with your gem already loaded
├── lib/
│ ├── automation.rb the entry point everything requires
│ └── automation/
│ └── version.rb VERSION constant, read by the gemspec
└── test/
├── test_helper.rb
└── test_automation.rb
Three conventions here are enforced by tooling rather than by taste:
lib/ is the load path. When your gem is installed, lib/ is added to $LOAD_PATH, so require "automation" finds lib/automation.rb. A file at lib/automation/runner.rb is required as require "automation/runner".Automation::Runner lives in lib/automation/runner.rb. Nothing forces this, and every Ruby programmer will assume it. Gemfile versus gemspec. The gemspec declares what your library needs from its users; the Gemfile declares what your development needs. For a gem, the Gemfile usually just says gemspec, meaning "read the gemspec", plus development-only tools.gem install ruby-lsp and point your LSP client at it. Set up RuboCop early (bundle add rubocop --group development) and run it with bundle exec rubocop -a to autocorrect. It is opinionated and configurable via .rubocop.yml; unlike gofmt, you will end up tuning it, and the tuning is a team decision rather than a universal one.
In a dynamic language, "go to definition" is a guess. When we start defining methods at run time with define_method and method_missing, your editor will not know those methods exist, and neither will your linter. That is a real and permanent cost of the techniques this course teaches, and Milestone 7 discusses how to mitigate it (documenting dynamic methods, respond_to_missing?, and knowing when a plain define_method beats a clever method_missing).
Put this in hello.rb anywhere:
# frozen_string_literal: true
module Automation
VERSION = "0.0.1"
def self.greet(name = "world")
"hello, #{name} (automation #{VERSION})"
end
end
puts Automation.greet
puts Automation.greet("pipelines")
$ ruby hello.rb
hello, world (automation 0.0.1)
hello, pipelines (automation 0.0.1)
# frozen_string_literal: true — a magic comment, read by the interpreter rather than ignored. It makes every string literal in this file frozen (immutable), which saves memory (identical literals are shared) and prevents a class of bug where one part of the program mutates a string another part is holding. Put it at the top of every file you write. It has one surprising consequence you will meet in Part 2.
module Automation — a module is a namespace and a bag of methods. Unlike a class, it cannot be instantiated. Top-level constant names must be capitalised; Automation::Runner is how you refer to something nested inside.
VERSION = "0.0.1" — a constant, by virtue of the capital letter. Ruby will warn if you reassign it but will not stop you, which tells you something about the language's attitude: it gives you guidance, not guarantees.
def self.greet(name = "world") — defines a method on the module itself rather than on instances, so you call it as Automation.greet. The self. prefix is how Ruby spells what other languages call static. name = "world" is a default argument.
"hello, #{name} ..." — string interpolation. #{} evaluates any Ruby expression and calls to_s on the result. Only double-quoted strings interpolate; single-quoted ones are literal.
There is no return. A Ruby method returns the value of its last expression. Explicit return exists and is used for early exits, but writing it at the end of a method marks you as a visitor from another language.
puts — a method from Kernel, which is mixed into every object, so it is callable anywhere without a receiver. It prints with a newline. Its cousins: print (no newline), p (prints inspect, the debugging representation, and returns its argument), and pp (pretty-prints nested structures). p is the one you will use most while developing.
$ irb
irb(main):001> require_relative "hello"
=> true
irb(main):002> Automation.greet("irb")
=> "hello, irb (automation 0.0.1)"
irb(main):003> Automation.methods(false)
=> [:greet]
irb(main):004> "hello".methods.size
=> 187
irb(main):005> 5.class.ancestors
=> [Integer, Numeric, Comparable, Object, Kernel, BasicObject]
Go development is edit-compile-run. Ruby development is a conversation: you keep an irb session open, load your code, poke at objects, and ask them what they can do. obj.methods, obj.class, obj.instance_variables and Klass.ancestors are how you learn an unfamiliar library, often faster than reading its documentation. Inside your project, bin/console starts irb with the gem already loaded.
Documentation: ri Array#map in a terminal, rubydoc.info for gems, and ruby-doc.org for the core library. In irb, help "Array#map" works too.
ruby lib/automation.rb # run a file
bundle exec rake test # run the test suite
ruby -Ilib test/test_step.rb # run one test file directly
ruby -Ilib test/test_step.rb -n test_value_equality # one test
bundle exec rubocop -a # lint and autocorrect
bin/console # irb with your gem loaded
bundle exec runs a command with exactly the gem versions your Gemfile.lock pins. Without it you get whatever is installed globally, which is how "works on my machine" happens in Ruby. Get into the habit.
A first test, in test/test_step.rb:
require "minitest/autorun"
class Step
attr_reader :name, :options
def initialize(name, **options)
@name = name.to_sym
@options = options.freeze
end
def ==(other)
other.is_a?(Step) && name == other.name && options == other.options
end
end
class StepTest < Minitest::Test
def setup
@step = Step.new("fetch", source: "arxiv")
end
def test_name_is_always_a_symbol
assert_equal :fetch, @step.name
assert_equal :fetch, Step.new(:fetch).name
end
def test_options_are_frozen
assert @step.options.frozen?, "options should be frozen"
assert_raises(FrozenError) { @step.options[:source] = "elsewhere" }
end
def test_value_equality
assert_equal Step.new(:fetch, source: "arxiv"), @step
refute_equal Step.new(:fetch, source: "other"), @step
end
end
$ ruby test/test_step.rb
Run options: --seed 27134
# Running:
...
Finished in 0.001256s, 2388.1814 runs/s, 4776.3627 assertions/s.
3 runs, 6 assertions, 0 failures, 0 errors, 0 skips
That is real output. Points to notice: setup runs before every test method; test methods must start with test_; the argument order is assert_equal expected, actual (getting it backwards produces confusing failure messages); and assert_raises takes a block, which is your first glimpse of how much of Ruby's design rests on blocks.
Create the gem skeleton with bundle gem automation. Then add a module method Automation.env that returns a Hash describing the runtime: Ruby version, platform, the gem's own version, and whether $PROGRAM_NAME suggests it is running under a test. Print it nicely from a script. Then write a test asserting the Hash has exactly the keys you expect.
Hints: RUBY_VERSION, RUBY_PLATFORM and $PROGRAM_NAME are globals; Hash#keys; assert_equal on sorted arrays of symbols.
# lib/automation.rb
# frozen_string_literal: true
require_relative "automation/version"
module Automation
def self.env
{
ruby: RUBY_VERSION,
platform: RUBY_PLATFORM,
automation: VERSION,
testing: $PROGRAM_NAME.include?("test")
}
end
end
# test/test_env.rb
require "minitest/autorun"
require "automation"
class EnvTest < Minitest::Test
def test_env_has_expected_keys
assert_equal %i[automation platform ruby testing], Automation.env.keys.sort
end
def test_ruby_version_is_a_string
assert_kind_of String, Automation.env[:ruby]
assert_match(/\A\d+\.\d+/, Automation.env[:ruby])
end
end
Three details worth absorbing. %i[a b c] is shorthand for an array of symbols, and its sibling %w[a b c] makes an array of strings; both appear constantly in real Ruby. require_relative is for files inside your own project (resolved relative to the current file) while require searches the load path and is for gems and standard library. And assert_kind_of rather than assert_equal String, x.class, because the former accepts subclasses, which is usually what you mean.

cannot load such file -- automation (LoadError) — lib/ is not on the load path. Run with ruby -Ilib ..., or use require_relative, or run through rake/bundle exec, which set it up for you.undefined method 'greet' for Automation:Module — you wrote def greet instead of def self.greet, so it is an instance method on a module that has no instances.uninitialized constant Automation::Runner — the file defining it was never required. Ruby does not autoload by convention (Rails adds that); you must require it.can't modify frozen String — the magic comment is doing its job. Use +"literal" or String.new or, better, stop mutating strings.syntax error, unexpected end-of-input — a missing end. Ruby cannot tell you where you meant to put it, only where the file ran out. Consistent indentation and a good editor are the defence.sudo gem install. If you need sudo, your Ruby installation is wrong. Fix the version manager instead.require and require_relative?bundle exec protect you from?Automation::Runner live, and what enforces that?frozen_string_literal magic comment do, and why would you want it?puts, print, p and pp?Only what the DSL needs, which turns out to be most of Ruby's object model and none of its web frameworks. Every output below is copied from an actual run. Keep an irb session open and type the examples.
p 5.class, "x".class, nil.class, (1..3).class, :sym.class, Integer.class
p 1.+(2), 5.between?(1, 10), nil.to_a, nil.to_s.empty?
Integer
String
NilClass
Range
Symbol
Class
3
true
[]
true
5.class is Integer: numbers are objects with methods. There are no primitives. nil.class is NilClass, and nil is a real object you can call methods on. nil.to_a gives [], nil.to_s gives "". This is why Ruby code has fewer nil checks than you would expect.Integer.class is Class: classes are objects too, instances of Class. That fact is what makes define_method and the rest of Milestone 7 possible.1.+(2) works because + is a method. Operators are ordinary methods with special syntax, which means you can define +, ==, <=> and [] on your own types.In Java, int is not an object and 5.toString() is a syntax error. In Python, 5 is an object but + dispatches through __add__ and classes are instances of type in a way you mostly ignore. Ruby has no exceptions to the rule, which sounds like trivia until you realise that "send a message to an object" being the only mechanism is what lets you intercept all of it. A DSL in Ruby is, at bottom, a set of objects that respond interestingly to messages.
a, b = "name", "name"
p a.equal?(b), a == b, a.frozen?
p :name.equal?(:name), :name.to_s, "name".to_sym
false # two separate String objects (in a file without the magic comment)
true # with the same contents
false
true # :name is always the same object
"name"
:name
A symbol is an interned, immutable name. :fetch written twice is the same object; "fetch" written twice is two objects. Use symbols for identifiers (method names, hash keys, step names, states) and strings for text (messages, content, user data). Our DSL will normalise every step name to a symbol, which is why Step.new("fetch").name returned :fetch earlier.
Run the same snippet in a file that starts with # frozen_string_literal: true and the first line prints true: identical frozen literals are deduplicated into one object. So equal? (object identity) gives different answers depending on a comment at the top of the file. This is worth knowing before it confuses you at 2am. The lesson is not to avoid the magic comment; it is to use == for comparisons and equal? essentially never.
p [0, "", [], nil, false].map { |v| v ? "truthy" : "falsy" }
["truthy", "truthy", "truthy", "falsy", "falsy"]
Only nil and false are falsy. Zero is true. The empty string is true. The empty array is true. Coming from Python or JavaScript this is the single most likely source of a wrong conditional in your first week, and it cuts both ways: if items does not check for emptiness, it checks for existence. Write if items.empty? when you mean empty.
Two idioms that follow:
@steps ||= [] # assign only if currently nil or false
name = opts[:name] || "unnamed"
count = config&.limit # safe navigation: nil if config is nil, no NoMethodError
def describe(name, *rest, retries: 3, **opts)
"#{name} retries=#{retries} rest=#{rest.inspect} opts=#{opts.inspect}"
end
puts describe("fetch", retries: 5, timeout: 2)
def double(x) = x * 2 # endless method (Ruby 3.0+)
p double(21)
fetch retries=5 rest=[] opts={:timeout=>2}
42
*rest collects extra positional arguments into an Array (a "splat").retries: 3 is a keyword argument with a default. Since Ruby 3.0, keyword arguments are genuinely separate from positional Hashes, which fixed a decade of subtle bugs.**opts collects unrecognised keyword arguments into a Hash (a "double splat"). This is the feature that makes our DSL's filter topic: "AI", min_citations: 5 work: any options a step wants, without declaring them.def double(x) = x * 2 is an endless method, good for one-liners.def f(a, retries: 3, *rest) is a syntax error, which I confirmed by making exactly that mistake while preparing this instalment.Naming conventions that Ruby programmers treat as meaning, not decoration:
| Form | Means | Example |
|---|---|---|
name? | Returns a boolean | empty?, valid? |
name! | Dangerous: mutates, or raises where the plain version does not | sort!, validate! |
name= | A setter, called as obj.name = x | attr_writer generates these |
snake_case | Methods and variables | save_to |
CamelCase | Classes and modules | StepRegistry |
SCREAMING | Constants | VERSION |
steps = %w[fetch filter summarize save]
p steps.map(&:upcase)
p steps.select { |s| s.length > 4 }
p steps.each_with_index.map { |s, i| "#{i}:#{s}" }
p steps.reduce("") { |acc, s| acc + s[0] }
config = { topic: "AI", limit: 10 }
p config[:topic], config[:missing], config.fetch(:limit), config.key?(:topic)
nested = { pipeline: { steps: [{ name: "fetch" }] } }
p nested.dig(:pipeline, :steps, 0, :name)
p steps.each_with_object({}) { |s, h| h[s.to_sym] = s.length }
["FETCH", "FILTER", "SUMMARIZE", "SAVE"]
["fetch", "filter", "summarize"]
["0:fetch", "1:filter", "2:summarize", "3:save"]
"ffss"
"AI"
nil
10
true
"fetch"
{:fetch=>5, :filter=>6, :summarize=>9, :save=>4}
%w[...] builds an array of strings without quotes or commas.&:upcase is shorthand for { |s| s.upcase }. The & converts a Symbol into a block by calling to_proc on it, which is a small piece of metaprogramming you use daily without noticing.config[:missing] returns nil; config.fetch(:missing) raises KeyError. Use fetch for things that must be present: a nil that travels three layers before exploding is much harder to debug than an immediate KeyError.dig walks nested structures and returns nil rather than raising when a level is missing.each_with_object({}) is the idiomatic "build a collection while iterating"; the object is passed to each iteration and returned at the end. reduce is its cousin where the accumulator is the block's return value, which is why the reduce above works but is easy to get wrong.Python has one iteration protocol and a handful of builtins (map, filter, comprehensions). Ruby has Enumerable, a module with about sixty methods, mixed into anything that defines each. That means group_by, partition, each_slice, tally, sum, min_by, flat_map, zip, lazy and more are available on your own classes the moment you define each and include the module. Learning Enumerable is the single highest-value hour you can spend on Ruby's standard library.
A block is a chunk of code attached to a method call. It is not an argument in the ordinary sense: every Ruby method can take one, and the method decides whether to use it.
def each_step
return to_enum(:each_step) unless block_given?
yield "fetch"
yield "filter"
:done
end
each_step { |s| print s, " " }
p each_step.to_a
def with_logging(name, &block)
puts "start #{name}"
result = block.call
puts "end #{name}"
result
end
p with_logging("run") { 40 + 2 }
fetch filter
["fetch", "filter"]
start run
end run
42
yield calls the block. It is the cheapest way to accept one and needs no parameter.block_given? asks whether a block was passed. Returning to_enum when there is no block is the standard trick that makes your method work both as each_step { ... } and as each_step.to_a, exactly like the built-in collections.&block in the parameter list captures the block as a Proc object you can store, pass on, or call later. This is how our DSL will keep a failure handler for later.yield evaluates to it.{ } for single-line blocks, do ... end for multi-line. They differ in precedence in rare cases, but the real reason is readability, and a DSL always uses do ... end.Why blocks matter more in Ruby than closures do elsewhere. Because every method can take exactly one block with no ceremony, Ruby programmers habitually write methods that take a chunk of behaviour: File.open(path) { |f| ... } closes the file afterwards, transaction { ... } commits or rolls back, assert_raises(E) { ... } checks an expectation. The pattern is "the method controls the resource, the block supplies the intent", and our whole DSL is one enormous application of it.
pr = proc { |a, b| [a, b] }
la = ->(a, b) { [a, b] }
p pr.call(1), pr.call(1, 2, 3), pr.lambda?, la.lambda?
begin
la.call(1)
rescue ArgumentError => e
puts "lambda strict: #{e.message}"
end
def returns_from_lambda
l = -> { return :from_lambda }
l.call
:from_method
end
def returns_from_proc
pr = proc { return :from_proc }
pr.call
:from_method
end
p returns_from_lambda, returns_from_proc
[1, nil]
[1, 2]
false
true
lambda strict: wrong number of arguments (given 1, expected 2)
:from_method
:from_proc
Two differences, and the second one is the dangerous one:
| Proc | Lambda | |
|---|---|---|
| Arity | Lenient: missing args become nil, extra ones are dropped | Strict: wrong count raises ArgumentError |
return | Returns from the enclosing method | Returns from the lambda only |
| Syntax | proc { }, or a block captured with & | ->(x) { } or lambda { } |
Look at the output again: returns_from_proc returned :from_proc, meaning the return inside the proc terminated the whole method and the last line never ran. The lambda version returned :from_method, because its return only left the lambda. A block captured with &block is a Proc, so a user's when_failed do ... return ... end could return from surprising places. Prefer lambdas when you store user code and call it later, and we will.
Write a method retrying(times:, on: StandardError) that takes a block, calls it, and retries up to times attempts if the block raises an exception of the given class, re-raising if it never succeeds. It should return the block's value on success, and it should be usable as:
result = retrying(times: 3) { flaky_call }Then extend it to yield the attempt number to the block, so the block can behave differently on a retry. Test it with a counter that fails the first two times.
def retrying(times:, on: StandardError)
attempt = 0
begin
attempt += 1
yield attempt
rescue on => e
retry if attempt < times
raise
end
end
attempts = 0
value = retrying(times: 3) do |n|
attempts += 1
raise IOError, "flaky" if n < 3
"succeeded on attempt #{n}"
end
p value, attempts
# => "succeeded on attempt 3"
# => 3
Four things in eleven lines. retry is a keyword that re-runs the begin block from the top; no loop needed, and no other mainstream language has it. A bare raise inside a rescue re-raises the current exception with its original backtrace, which is what you want; raise e also works but loses nothing only because Ruby is kind. on: defaults to StandardError rather than Exception, because Exception includes SignalException and Interrupt, so rescuing it swallows Ctrl-C. And yield attempt passes a value to the block, which the caller may ignore: blocks are lenient about arity, so retrying(times: 3) { flaky } still works.
This method is, almost unchanged, what Milestone 6 puts behind retry_on.
class Step
attr_reader :name, :options
def initialize(name, **options)
@name = name.to_sym
@options = options.freeze
end
def to_s = "#{@name}(#{@options.inspect})"
def ==(other) = other.is_a?(Step) && name == other.name && options == other.options
def call(input) = raise(NotImplementedError, "#{self.class} must implement #call")
end
s = Step.new("fetch", source: "arxiv")
puts s
p s == Step.new(:fetch, source: "arxiv"), s.options.frozen?
fetch({:source=>"arxiv"})
true
true
@name is an instance variable. It springs into existence on assignment, is nil if never assigned, and is private to the object: there is no way to read it from outside without a method.attr_reader :name generates a method name that returns @name. Its siblings are attr_writer and attr_accessor. attr_reader is itself a method call, executed when the class body runs, that defines methods. Class bodies are ordinary code, which is the foundation of everything in Milestone 7.initialize is the constructor, called by Step.new.== gives you value equality. Define hash and eql? too if instances will be hash keys or need uniq.raise(NotImplementedError, ...) in a base method is Ruby's abstract method: there is no abstract keyword, and this is the convention.self.class gives an object its own class, so the error message names the actual subclass. There is no new keyword (new is a method on the class object), no field declarations (instance variables appear when assigned), no access modifiers on state (all instance variables are private, always), and no compile-time interface. Visibility applies only to methods, via private and protected, and private is itself a method call that switches a mode for the rest of the class body.
module Loggable
def log(msg) = puts("[#{self.class.name}] #{msg}")
end
module Registry
def register(name) = (@registered ||= []) << name
def registered = @registered || []
end
class Pipeline
include Loggable # instance methods
extend Registry # class methods
end
Pipeline.new.log("hello")
Pipeline.register(:fetch)
Pipeline.register(:save)
p Pipeline.registered, Pipeline.ancestors.first(4), Pipeline.include?(Loggable)
[Pipeline] hello
[:fetch, :save]
[Pipeline, Loggable, Object, Kernel]
true
include adds instance methods; extend adds methods to the object doing the extending, which for a class means class methods. That one sentence resolves most Ruby module confusion.
ancestors shows the method lookup chain: Ruby searches Pipeline, then Loggable, then Object, then Kernel, then BasicObject, and calls the first matching method. There is no multiple inheritance, but a class can include many modules, and they stack in the chain. This is worth remembering as the honest answer to "does Ruby have multiple inheritance": no, it has a linearised chain you can inject into, which gets you the useful part without the diamond problem.
module Automation
VERSION = "0.1.0"
class Error < StandardError; end
class StepError < Error
def initialize(step, cause) = super("step #{step} failed: #{cause}")
end
end
p Automation::VERSION, Automation::StepError.ancestors.first(3)
"0.1.0"
[Automation::StepError, Automation::Error, StandardError]
Every gem should define one base error class inside its namespace and inherit all its others from it, so a user can write rescue Automation::Error and catch everything you raise without catching everything in the world. This is the Ruby equivalent of Go's sentinel error hierarchy and it is not optional in a library.
class Builder
def initialize = @steps = []
def fetch(what) = @steps << [:fetch, what]
def steps = @steps
def build(&block)
instance_eval(&block) # self inside the block becomes this builder
@steps
end
end
p Builder.new.build { fetch "papers" }
p Builder.new.instance_eval { self.class }
p Builder.new.instance_exec(3) { |n| n * 2 }
[[:fetch, "papers"]]
Builder
6
Look carefully at the first line of output. The block { fetch "papers" } was written at the top level of the file, where no method called fetch exists. It worked because instance_eval evaluated the block with self set to the builder, so the bare call fetch "papers" was sent to the builder.
This is the whole trick. Every Ruby DSL you have ever seen (RSpec's describe/it, Rake's task, a Gemfile's gem, Sinatra's get) is this: a block, evaluated with self pointing at an object that responds to the vocabulary. Milestone 3 does it properly, including the significant problems it causes, which are: the block can no longer see the caller's private methods, method names silently collide, and typos become confusing NoMethodErrors rather than NameErrors.
instance_exec is the same thing but passes arguments to the block. Both are the sharpest tools in Ruby, and both are how you cut yourself.
attempts = 0
begin
attempts += 1
raise Automation::StepError.new(:fetch, "timeout") if attempts < 3
puts "succeeded on attempt #{attempts}"
rescue Automation::Error => e
retry if attempts < 3
puts "giving up: #{e.message}"
ensure
puts "ensure always runs (attempts=#{attempts})"
end
def risky
yield
rescue ZeroDivisionError => e
raise Automation::StepError.new(:divide, e.message)
end
begin
risky { 1 / 0 }
rescue Automation::StepError => e
puts "#{e.class}: #{e.message} / cause: #{e.cause.class}"
end
succeeded on attempt 3
ensure always runs (attempts=3)
Automation::StepError: step divide failed: divided by 0 / cause: ZeroDivisionError
rescue SomeError => e catches that class and its subclasses, which is why the error hierarchy matters.retry restarts the begin block. Always bound it with a counter, or you have written an infinite loop with extra steps.ensure runs on every path: success, exception, even an early return. It is Ruby's defer.e.cause is set automatically when you raise inside a rescue: Ruby remembers the original. This is Go's %w wrapping, for free, with no effort from you. begin, so def risky ... rescue ... end needs no explicit block.StandardError, never Exception. A bare rescue => e means StandardError, which is correct. Writing rescue Exception catches Interrupt and SystemExit, so your program ignores Ctrl-C and cannot be killed politely.class Ghost
def initialize = @calls = []
def method_missing(name, *args, &blk)
return super unless name.to_s.start_with?("step_")
@calls << [name, args]
self
end
def respond_to_missing?(name, include_private = false)
name.to_s.start_with?("step_") || super
end
def calls = @calls
end
g = Ghost.new
g.step_fetch("papers").step_save("kb")
p g.calls, g.respond_to?(:step_anything), g.respond_to?(:nope)
g.nope # => NoMethodError
[[:step_fetch, ["papers"]], [:step_save, ["kb"]]]
true
false
NoMethodError: undefined method `nope' for #<Ghost:0x00...
method_missing is called when an object receives a message it has no method for. Three rules make the difference between a useful ghost and an unmaintainable one:
super for the rest. The return super unless line is what preserves the normal NoMethodError for genuine typos. Without it, every misspelling silently succeeds and your users debug ghosts.respond_to_missing? alongside it. Otherwise respond_to? lies, and so do method(), duck-typing checks and every tool that introspects your object. Notice the output: respond_to?(:step_anything) is true because we defined it.self makes calls chainable, which is how g.step_fetch(...).step_save(...) works.class Dynamic
%i[fetch filter save].each do |name|
define_method(name) { |arg = nil| "called #{name} with #{arg.inspect}" }
end
end
d = Dynamic.new
p d.fetch("x"), d.public_send(:filter, 2), Dynamic.instance_methods(false).sort
"called fetch with \"x\""
"called filter with 2"
[:fetch, :filter, :save]
define_method defines a real method from a block, at run time. Compare the two approaches, because Milestone 7 chooses between them constantly:
define_method | method_missing | |
|---|---|---|
| Method really exists | Yes: shows in methods, respond_to?, documentation tools | No: must fake it all |
| Speed | Normal method dispatch | Slower: only runs after lookup fails |
| Needs the names in advance | Yes | No |
| Verdict | Prefer this | Only when the set is genuinely open |
send calls a method by name, including private ones; public_send respects visibility and is what you should use when the name comes from user input or a registry.
StepNode = Data.define(:name, :options)
n = StepNode.new(name: :fetch, options: { source: "arxiv" })
p n, n.name, n.frozen?
p n.with(name: :refetch)
n.instance_variable_set(:@name, :hacked) # => FrozenError
#<data StepNode name=:fetch, options={:source=>"arxiv"}>
:fetch
true
#<data StepNode name=:refetch, options={:source=>"arxiv"}>
FrozenError: can't modify frozen StepNode: #<data Ste...
Data (Ruby 3.2+) creates an immutable value class with readers, value equality, a decent inspect, and with for making modified copies. This is exactly what our AST nodes need, and it arrives with no boilerplate. Its mutable older sibling is Struct, which allows positional construction and assignment; prefer Data for anything representing a description rather than a state.
Immutability in the AST is not aesthetic. In Milestone 11 pipelines rewrite themselves, and "rewrite" will mean "produce a new pipeline from the old one" rather than "mutate in place". That makes the old version still valid, makes diffs possible, and makes a failed transformation harmless.
Build a miniature of the whole project in about sixty lines. Requirements:
StepNode = Data.define(:name, :options).Builder that collects step nodes and whose method_missing turns any bare verb into a step, so fetch "papers", from: "arxiv" and summarize max_words: 50 both work. Include respond_to_missing?.Mini.pipeline(name, &block) that evaluates the block against a builder and returns a frozen PipelineNode = Data.define(:name, :steps).Runner that walks the steps and calls a handler registered by name in a Hash, threading the result of each step into the next. Unknown step names must raise a clear error naming the step and listing the known ones.The one design rule: building must not execute anything. You should be able to build a pipeline whose steps are all unregistered and only get an error when you run it.
# frozen_string_literal: true
module Mini
class Error < StandardError; end
class UnknownStep < Error
def initialize(name, known)
super("unknown step #{name.inspect}; known steps: #{known.sort.join(', ')}")
end
end
StepNode = Data.define(:name, :options)
PipelineNode = Data.define(:name, :steps)
# Builder only collects. It never runs anything, which is what makes
# dry runs, validation and rewriting possible later.
class Builder
attr_reader :steps
def initialize
@steps = []
end
def method_missing(name, *args, **options, &_block)
@steps << StepNode.new(name: name, options: options.merge(args: args))
self
end
def respond_to_missing?(_name, _include_private = false) = true
end
def self.pipeline(name, &block)
builder = Builder.new
builder.instance_eval(&block)
PipelineNode.new(name: name.to_sym, steps: builder.steps.freeze).freeze
end
class Runner
def initialize(handlers) = @handlers = handlers
def run(pipeline, input = nil)
pipeline.steps.reduce(input) do |acc, step|
handler = @handlers.fetch(step.name) { raise UnknownStep.new(step.name, @handlers.keys) }
handler.call(acc, **step.options)
end
end
end
end
pipeline = Mini.pipeline("research") do
fetch "papers", from: "arxiv"
filter topic: "AI"
summarize max_words: 50
end
p pipeline.steps.map(&:name)
# => [:fetch, :filter, :summarize]
handlers = {
fetch: ->(_input, args:, **) { ["paper about #{args.first}", "paper about cats"] },
filter: ->(items, topic:, **) { items.grep(/#{topic}/i) },
summarize: ->(items, max_words:, **) { items.map { |i| i[0, max_words] } }
}
p Mini::Runner.new(handlers).run(pipeline)
# => ["paper about papers"]
Mini::Runner.new({}).run(pipeline)
# => Mini::UnknownStep: unknown step :fetch; known steps:
Notes on the choices, several of which are the same choices the real project makes.
method_missing returning self, and accepting *args, **options so both positional and keyword forms work. Folding args into the options Hash is a shortcut; the real version keeps them separate, because the validator needs to check them differently.respond_to_missing? returning true unconditionally is honest here (the builder really does accept anything) and is exactly what you must not do in the real project, where the registry knows which verbs exist and a typo should be caught.Hash#fetch with a block raises your own error rather than KeyError, and the message lists the known steps. Error messages that tell the user what they could have written instead are the difference between a DSL people like and one they tolerate. reduce threads the accumulator through the steps, which is the entire interpreter in one line. Milestone 5 expands it into a Runner with a Context, logging and middleware, but the shape does not change.FrozenError immediately rather than a mysterious mutation later.If you built something close to this, you have already written the skeleton of Milestones 1 through 5, and the rest of the course is about making each piece real: proper errors with source locations, a registry that validates, failure handling, plugins, inspection and packaging.
0 or "" is falsy. They are not. Only nil and false.return at the end of a method. Harmless, but it marks you as a tourist. Worse: return inside a proc exits the enclosing method.hash[:key] where you meant hash.fetch(:key). The nil travels and explodes far from the cause.dup or, better, build new strings.method_missing without respond_to_missing?, or without super for unhandled names. Both make objects that lie about themselves.rescue Exception. Now Ctrl-C does not work.attr_reader, include and private are method calls that run when the class is defined, not declarations.assert_equal takes expected first.Integer.class equal to Class, and why does that matter for a DSL?&:upcase do, mechanically?include and extend?instance_eval change, and why is that the foundation of Ruby DSLs?define_method better than method_missing?You have now seen the four features that make the rest of this course possible: blocks as syntax, instance_eval to redirect self, method_missing and define_method to decide the vocabulary at run time, and class bodies being ordinary executable code. No other mainstream language has all four, and it is why an entire generation of infrastructure tools put a Ruby DSL at the front.
The honest cost, which you have also seen: none of this is checked before it runs. A typo in a step name, a proc whose return escapes, a method_missing that swallows a mistake, a monkey-patched method that changes behaviour three files away. Go's compiler would have caught most of the corresponding mistakes. Ruby gives you expressiveness and hands you the responsibility for correctness, which is why the testing milestone in this course is not optional and why Milestone 4's decision to build an inspectable AST matters so much: it is how we claw back some of what the compiler would have given us.
Milestone 1 sets up the real gem, defines Step and Registry properly, and gets a test suite running. Milestone 2 introduces blocks as the DSL surface with an explicit builder argument, so you can see what instance_eval buys before we use it. Milestone 3 makes the switch and pays the price. By Milestone 5 you will have an interpreter.
Before then, two things worth doing:
bundle gem automation and commit the skeleton, and spend ten minutes in irb calling .methods on things. Ruby rewards poking at it in a way that compiled languages do not.