Milestones 1–4

Instalment 16 · Course 4 (Erlang) · Parts 0–2

The language that assumes everything will crash

The Mewlang cat, looking up curiouslyA mesh of a thousand independent nodes that send each other messages, and periodically, deliberately, fall over. Supervisors notice. They come back. Nothing else in this curriculum treats failure as a first-class, expected, routinely-exercised code path the way this one does.

Course 4 · Part 0What are we building?

The final result

A simulated network, mesh, of up to several thousand independent nodes, each an Erlang process. Nodes exchange messages, maintain a small amount of state, and periodically misbehave: a node can crash outright, hang and stop responding, send another node a message it cannot parse, or simply vanish. A supervision tree watches every node and restarts the ones that die, with a policy that says precisely how much dying is tolerable before the supervisor gives up and takes the whole subtree down with it.

On top of that: a registry so nodes can find each other by name instead of by process identifier: an observability layer with counters for messages sent, nodes alive, nodes restarted, and message loss: a live dashboard, served over plain HTTP from a process running inside the same system it is reporting on, showing all of that updating in real time in a browser. Then, in the back half of the course, the single-machine simulation becomes a genuine distributed system: several real operating-system processes, each running its own instance of the Erlang virtual machine (the BEAM), connected over a real network, able to survive one of those machines being killed outright.

$ ./bin/mesh start --nodes 2000 --dashboard :8080
mesh: 2000 nodes started, supervisor tree depth 3
mesh: dashboard listening on http://localhost:8080

$ ./bin/mesh chaos --crash-rate 0.02 --partition-every 30s
chaos: injecting failures into a live system

  alive: 1,978   restarting: 14   partitioned: 8   msgs/sec: 40,221

Everything in that output is a real, measured number from a real run once the project exists. Nothing here is a diagram of an idea; every milestone below produces working, running, compiled code.

Why this project is interesting

Every other course in this curriculum treats a crash as a bug to eliminate. This one treats it as a certainty to plan for. That is not a small difference in emphasis — it is a different starting axiom for the whole design. A Go service that panics takes the whole process down with it, and the discipline is to never let a goroutine panic in the first place. An Erlang process that crashes is routine: by the end of Milestone 4 you will have a system where processes die constantly, on purpose, in the test suite, and the assertion is never "no process died" — it is "the system as a whole kept working while processes died."

That reframing is the entire subject of this course. Once you accept that any single process can and will die, the interesting engineering question stops being "how do I prevent this" and becomes "who notices, what do they do about it, and how fast." The answer Erlang and its OTP (Open Telecom Platform) library give is supervision trees: a static, declarative description of which processes depend on which, and exactly what should happen — restart just this one, restart its siblings too, give up entirely — when one of them dies. You will hand-write the naive version in Milestone 4, before you have been told what a supervisor is, discover exactly what it gets wrong, and then meet OTP's supervisor behaviour in Milestone 6 as the answer to a problem you have already felt.

Why Erlang in particular

Three properties of the language are not available, in combination, anywhere else in this curriculum.

Processes are cheap and share nothing. An Erlang process is not an OS thread and not a goroutine — it is a lightweight, independently garbage-collected unit with its own heap, started in microseconds, and a BEAM instance can comfortably run millions of them. Because no process can reach into another's memory — the only way in is a message, copied, into a mailbox — a crashed process cannot corrupt anything belonging to a process that did not crash. Go's goroutines are cheap too, but they share an address space, so a wild pointer write in one can in principle corrupt data another is reading; the go runtime's response to an unrecovered panic in any goroutine is to kill the entire program. Erlang's response to an unhandled exception in a process is to kill exactly that process, and nothing else, ever.

Let it crash is a real engineering strategy, not a slogan. Given the isolation above, the language actively discourages the defensive-programming instinct of catching every possible error and trying to handle it in place. If a node receives a message it cannot make sense of, the idiomatic move is to let the process crash — loudly, with a stack trace, to a log — and let the supervisor restart it into a known-good state, rather than trying to enumerate every way input can be wrong and patch around each one. Milestone 8 makes this concrete: you will feed the mesh genuinely malformed messages and watch the system absorb them by discarding and restarting, rather than by defending.

Distribution was never an afterthought. The message-passing model between two processes on one machine and two processes on different machines is, by design, close to the same API — Erlang was built at Ericsson in the late 1980s specifically to run telephone switches that were not allowed to go down, and those switches were distributed hardware from day one. Milestone 11 turns your single-machine mesh into several real operating-system processes talking over a real network with a change measured in tens of lines, not a rewrite — because the message-passing discipline you have been using since Milestone 3 was already, quietly, the distributed-systems discipline.

Why are we using this language here?

Be honest about the trade this makes. Erlang's processes are wonderful for isolation and terrible for sharing a large, frequently-updated data structure between many workers — there is no equivalent of Go's atomic-read sharded engine from Course 1, because sharing memory is exactly what the model refuses to do. Where Go would shard a struct behind atomics for raw throughput, Erlang would shard the problem into more processes and accept the cost of a message copy at the boundary. For a system whose entire premise is "parts of it are broken right now," that trade is the right one: you are buying certainty about blast radius, and you are paying for it in copying overhead. Both are measured, in this course, rather than asserted.

Architecture we are building toward

                         mesh_sup (one_for_one)
                        /          |            \
              node_sup_1   registry_sup      dashboard_sup
             (simple_one   (name -> pid,    (http listener,
              _for_one)      process groups)  metrics collector)
              /   |   \
          node  node  node  ...  (up to several thousand,
           |     |     |          each a gen_server, each linked
        mailbox mailbox mailbox   to node_sup_1, each able to
                                   crash independently)

  every arrow above is a supervision link: a child's death is
  reported to its supervisor, which decides — by a declared,
  static policy — whether to restart just that child, restart
  every child alongside it, or give up and die itself, which
  reports the death one level further up

Three architectural decisions, explained now because they shape every milestone that follows.

A node is a gen_server, OTP's standard behaviour for "a process with state that answers requests" — you will write the raw receive loop version first in Milestone 3, by hand, before meeting gen_server in Milestone 5 as the same idea with the boilerplate handled for you and a long list of footguns removed.

Every node sits under a supervisor, and supervisors sit under supervisors. This is the tree in "supervision tree": node_sup_1 is a simple_one_for_one supervisor, a template for spawning an unbounded, dynamic number of identical children — exactly what a few thousand interchangeable network nodes need — while mesh_sup at the root uses one_for_one, because the registry, the node pool and the dashboard are three genuinely different things that should not be restarted as a group just because one of them died.

The registry and the dashboard are processes too, supervised exactly like the nodes they serve. There is no special "infrastructure" tier that gets to skip the discipline the rest of the system follows — if the registry crashes, something notices and restarts it, the same as everything else.

The twelve milestones

#MilestoneWhat it teaches
1The shell and a first moduleerl, expressions, atoms, compiling, rebar3
2Functions over datapattern matching, immutability, lists/tuples/maps, recursion, tail calls, guards
3A node as a raw processspawn, !, receive, mailboxes, selective receive
4Let it crashlinks, monitors, exit signals, trap_exit, why defensive code is discouraged
5The node as a gen_serverOTP behaviours, call vs cast, state, timeouts
6Supervision treesrestart strategies, intensity limits, dynamic children, startup ordering
7A network of a thousand nodesregistries, ETS, process groups, routing, message loss
8Chaos injectionrandom crashes, slow handlers, malformed messages, memory pressure
9Observabilitycounters, observer, recon, tracing a live system safely
10The live dashboardan HTTP server in Erlang, streaming updates, aggregation without slowing the system
11Actually distributedmultiple BEAM nodes, cookies, net_kernel, partitions, split brain
12Releases and property testsEUnit, Common Test, PropEr, relx releases, upgrades

What you will know afterwards

Why isolating failure is a design decision with a measurable cost, not a free lunch; how to describe a system's failure-response policy declaratively instead of scattering try/catch through it; how OTP's standard behaviours turn "a process that holds state and answers requests" from a hand-rolled loop into a few callback functions; what actually changes, in code, when a message-passing system stops being single-machine and becomes genuinely distributed; and why a network partition is not a bug you fix but a condition you are required to have an opinion about.


Course 4 · Part 1Install and first program

What Erlang is

Erlang is a dynamically typed, functional, concurrent language built at Ericsson starting in 1986 to program telephone exchanges — hardware that was contractually required to have at most a few minutes of downtime a year. Two consequences follow directly from that origin and explain nearly everything the language does differently from a general-purpose scripting or systems language. First, the runtime (the BEAM, Erlang's virtual machine) was built around lightweight, isolated processes and message passing from the start, because a telephone exchange is naturally a huge number of mostly independent, concurrent calls. Second, the language and its standard library, together known as OTP (Open Telecom Platform), encode decades of "how do you build a system that must not go down" as reusable behaviours — gen_server, supervisor, and others — rather than leaving every team to reinvent them.

Erlang is not "the old version of Elixir." Elixir is a newer language that compiles to the same BEAM bytecode and can call Erlang code directly (and vice versa); Erlang came first, by about 25 years, and this course uses it directly rather than through Elixir's syntax, because the OTP behaviours and the supervision-tree model — the actual subject of this course — are identical either way, and Erlang's syntax, once past the first hour of unfamiliarity, gets out of the way of that subject rather than adding a second one.

Installing

You want a recent OTP — this course was verified on OTP 25 (Erlang/OTP 25, erts-13.1.5) — and, separately, rebar3, the standard build tool and package manager, roughly equivalent to Go's own toolchain or Ruby's Bundler.

Linux
sudo apt install erlang rebar3      # Debian/Ubuntu
sudo dnf install erlang rebar3      # Fedora

erl -eval 'io:format("~s~n", [erlang:system_info(system_version)]), halt().' -noshell
# Erlang/OTP 25 [erts-13.1.5] [source] [64-bit] [smp:N:N] [ds:N:N:N] [jit:ns]
macOS
brew install erlang rebar3
Windows
:: the official installer from erlang.org, or:
winget install ErlangSolutions.Erlang
:: rebar3 is a self-contained escript; download rebar3 and put it on PATH

WSL2 is worth it here, as it was for Perl: observer (Milestone 9) is a GUI tool that is simplest to run from a proper Linux environment, and the distributed-node work in Milestone 11 involves Unix-y networking assumptions that are easiest to reason about on Linux.

When you need a specific version
curl -fsSL https://raw.githubusercontent.com/kerl/kerl/master/kerl -o kerl
chmod +x kerl
./kerl build 25.3.2.9 25.3.2.9
./kerl install 25.3.2.9 ~/erlangs/25.3.2.9
. ~/erlangs/25.3.2.9/activate

kerl is the Erlang-world equivalent of Perl's perlbrew or Go's own version management: it builds a specific OTP release from source into an isolated directory and gives you an activate script to bring it onto your PATH.

The shell is not optional

Every other course in this curriculum treats its language's REPL as a nice-to-have for quick experiments. Erlang's shell, erl, is closer to a permanent fixture of how you work: it can connect to a running production system and let you call functions, inspect process state, and even hot-load new code into it without stopping it — a capability the telephone-exchange origin story makes obvious in retrospect. You will use it constantly, not just at the start.

$ erl
Erlang/OTP 25 [erts-13.1.5] [source] [64-bit] [smp:6:6] [jit:ns]

Eshell V13.1.5  (abort with ^G)
1> 1 + 1.
2
2> X = 40 + 2.
42
3> X = 42.
42
4> X = 43.
** exception error: no match of right hand side value 43

That last line is not a typo in this document — it is the single most important fact about the language, arriving in the fourth line of the shell. = is not assignment. It is pattern matching: X = 42 succeeds because X was unbound and binds it to 42; X = 42 again succeeds because 42 matches what X is already bound to; X = 43 fails, loudly, with an exception, because 43 does not match 42. Once a variable is bound in Erlang, it stays bound to that value for the rest of its scope — there is no reassignment, anywhere, ever. Section 2.1 below is entirely about what this buys you.

Project layout, with rebar3

$ rebar3 new app mesh
$ tree mesh
mesh/
├── rebar.config          dependencies and build config (like a Gemfile or go.mod)
├── src/
│   ├── mesh.app.src       application metadata: name, modules, dependencies
│   ├── mesh_app.erl        the application behaviour's entry point
│   └── mesh_sup.erl        the top-level supervisor (empty for now)
└── test/

rebar3 new app scaffolds an OTP application — not just a script, but a unit with a declared name, a supervision tree entry point, and a place for dependencies, tests and releases to live. That is deliberately more structure than Milestone 1 needs; the scaffold is there because by Milestone 6 you will need every part of it, and generating it once now means never having to retrofit it later.

Hello, mesh

%% src/hello.erl
-module(hello).
-export([world/0]).

world() ->
    io:format("hello, ~s~n", [node()]).
$ erlc hello.erl
$ erl -noshell -eval 'hello:world(), halt().'
hello, nonode@nohost

Every line, explained

The two names for "node", and how to keep them apart

This course's simulated network is made of mesh nodes — ordinary Erlang processes, thousands of them, all living inside one BEAM instance for most of the course. Erlang's own runtime separately has a notion of a distributed node — one running BEAM instance, identified by a name like mesh1@192.168.1.10, that can connect to other BEAM instances over the network. Until Milestone 11 there is exactly one distributed node and thousands of mesh nodes running inside it; from Milestone 11 onward there are several distributed nodes, each still running many mesh nodes. The text says "process" whenever precision matters and "mesh node" for the simulated concept; "node" alone, unqualified, always means the Erlang runtime sense.

Documentation, and getting help without leaving the shell

$ erl
1> h(lists, map).
...
2> lists:map(fun(X) -> X * 2 end, [1,2,3]).
[2,4,6]
$ erl -man gen_server        # the full behaviour reference, as a man page
$ open https://www.erlang.org/doc/                # the official docs, browsable

h/2 inside the shell prints a function's documentation without leaving your session — the closest thing here to Perl's perldoc -f or Go's go doc. The official docs at erlang.org are unusually good and unusually stable: OTP's standard library has unusually few breaking changes release to release, so documentation from three OTP releases ago is still mostly correct today.

Editor setup and formatting

Erlang's editor story is smaller than Go's "every editor has first-class gopls support out of the box," but it is genuinely solid once set up. Two real options, either of which is enough for this course:

Both integrate with rebar3 directly: point either one at a project with a rebar.config already in place (Milestone 1 has you create one) and it finds the project's own dependencies and include paths without further configuration.

$ cat rebar.config
{erl_opts, [debug_info]}.
{plugins, [rebar3_format]}.    %% or: {plugins, [erlfmt]}.

$ rebar3 format          # rebar3_format: reformats in place, project-wide
$ rebar3 fmt -w src/*.erl   # erlfmt: reformats the given files in place
$ rebar3 fmt --check       # erlfmt: reports files that need formatting, changes nothing

Neither formatter ships with rebar3 itself — both are plugins, added once to rebar.config the same way rebar3_proper is added in Milestone 12. They disagree on some stylistic choices (rebar3_format, for instance, tends to keep small map literals on one line where erlfmt more often expands them), so pick one for a given project rather than running both — mixing them means every commit reformats whatever the previous commit's tool did not agree with.

Testing, from the first day

OTP ships EUnit, a lightweight unit-testing framework, in the standard library — no dependency to add, the same way Go ships testing.

%% src/basics.erl
-module(basics).
-export([classify/1]).

classify(N) when N < 0 -> negative;
classify(0) -> zero;
classify(N) when N rem 2 =:= 0 -> even;
classify(_) -> odd.
%% test/basics_tests.erl
-module(basics_tests).
-include_lib("eunit/include/eunit.hrl").

classify_test() ->
    ?assertEqual(negative, basics:classify(-1)),
    ?assertEqual(zero,     basics:classify(0)),
    ?assertEqual(even,     basics:classify(4)),
    ?assertEqual(odd,      basics:classify(7)).
$ rebar3 eunit
===> Running EUnit tests...
  basics_tests: classify_test...[0.001 s] ok
  [done in 0.012 s]
=======================================================
  All 1 tests passed.

-include_lib("eunit/include/eunit.hrl") pulls in the ?assertEqual family of macros — yes, Erlang has a preprocessor and macros starting with ?, inherited from its C heritage and used sparingly outside test code. A function whose name ends in _test (singular) is a single test; one ending in _test_ (trailing underscore) is a test generator, returning a list of tests built programmatically — you will use both shapes throughout this course.

Checkpoint

  1. What does = actually do in Erlang, and why does X = 43 fail after X = 42?
  2. What two, independently meaningful things does a function's arity distinguish that a function's name alone does not?
  3. Name the two things a "node" can mean in this course, and which one node() the function returns.
  4. What is the difference between an EUnit test named foo_test and one named foo_test_?

Course 4 · Part 2Language crash course

2.1 Pattern matching, and the rule with no equivalent in Go, Ruby or Perl

You met it above: = matches rather than assigns. The same mechanism — not a special case of it, the literal same operation — is how you pull values out of compound data, choose between function clauses, and receive messages. There is exactly one destructuring rule in the whole language, used everywhere.

1> {Status, Value} = {ok, 42}.
{ok,42}
2> Status.
ok
3> [First | Rest] = [1,2,3,4].
[1,2,3,4]
4> First.
1
5> Rest.
[2,3,4]
6> {ok, 42} = {ok, 43}.
** exception error: no match of right hand side value {ok,43}

Line 6 is not a corner case to remember — it is the mechanism working as designed. Matching a literal against a value, rather than a fresh variable against it, asserts that the value has exactly that shape. This is how Erlang code routinely reads: a function expects its argument to look like {ok, Result} or its message to look like {From, Ref, Request}, and if it does not, the match fails and the process crashes — which, per this course's whole premise, is the correct response, not a bug to prevent.

Compare this to how a typical language handles the same job:

A typical language vs. Erlang

In Python or Ruby you would write status, value = pair and then a separate if status != "ok": raise ... to assert the shape you expected. Erlang's match is the assertion — there is no separate step, because the destructuring and the shape check are the same operation. The cost is that a bound variable can never be reused for something else in the same scope; the benefit is that "this data has the shape I assumed" is enforced at the exact point you assume it, not several lines later when something built from the wrong assumption finally breaks.

Exercise 2.1

In the shell, bind Point = {3, 4}. Write a pattern match that extracts both coordinates into X and Y in one line, then write an expression using only pattern matching (no if, no function call) that succeeds if and only if the point is exactly the origin, {0, 0}.

Solution
1> Point = {3, 4}.
{3,4}
2> {X, Y} = Point.
{3,4}
3> {0, 0} = Point.
** exception error: no match of right hand side value {3,4}

The third line is the answer to the second half: it succeeds silently for the origin and raises for anything else, which is "succeeds if and only if". Wrapping it in catch {0,0} = Point converts the exception into an ordinary value if you need the result rather than the side effect of not crashing.

2.2 Atoms, tuples and the convention that runs through everything

An atom is a name that is its own value — ok, error, undefined, mesh_node — comparable in spirit to a Ruby symbol, with no equivalent that is quite as central in Go or Perl. Atoms are used everywhere a smaller language would reach for a string constant or a small integer enum, and OTP's standard library leans on one convention hard enough that you should adopt it immediately: a function that can fail returns {ok, Result} on success and {error, Reason} on failure, as a two-element tuple, rather than throwing.

1> file:read_file("does_not_exist.txt").
{error,enoent}
2> file:read_file("/etc/hostname").
{ok,<<"my-machine\n">>}

Matching directly against the shape you expect is the idiomatic way to consume this:

{ok, Contents} = file:read_file(Path),
%% ... use Contents; if the read failed, the match fails and the
%% process crashes here, with a clear reason, which is correct —
%% see the "let it crash" section above.

<<"my-machine\n">> is a binary, Erlang's efficient representation for raw byte data — the type file contents, network payloads and, later in this course, malformed chaos-test messages arrive as.

An {error, Reason} tuple is not raised — it just sits there until you match on it

Coming from a language where a failed file read throws by default, it is easy to assume file:read_file/1 announces failure the same way. It does not — {error, enoent} is an entirely ordinary value, returned exactly like a success would be, and code that does not explicitly check for it will happily keep going with an error tuple sitting where real data was expected:

1> R = file:read_file("does_not_exist.txt").
{error,enoent}
2> {ok, Bin} = R.
** exception error: no match of right hand side value {error,enoent}

The crash above is the correct outcome, and it is the pattern match — {ok, Bin} = R — doing the work of announcing the failure, not the original call. Skip the match (bind Data = file:read_file(Path) and use Data directly, unmatched) and the error tuple flows silently into whatever uses it next, surfacing as a confusing failure far away from where it actually originated.

Exercise 2.2

Write describe_result/1, taking a value shaped like {ok, Value} or {error, Reason} and returning {success, Value} or {failure, Reason} respectively — using pattern matching in the function head, not a guard or an if.

Solution
describe_result({ok, Value}) -> {success, Value};
describe_result({error, Reason}) -> {failure, Reason}.

Verified: describe_result({ok, 42}) gives {success,42}; describe_result({error, not_found}) gives {failure,not_found}. Two clauses, no branching logic written by hand — exactly Section 2.5's "multiple clauses instead of if" idiom, one section early.

2.3 Lists, recursion, and tail calls

Erlang has no loop construct at all — no for, no while. Repetition is always recursion, which sounds alarming until you meet the guarantee that makes it practical: a tail call — a recursive call in the last position of a function clause, with nothing left to do after it returns — is compiled to a jump, not a new stack frame, so a correctly-written recursive loop runs in constant stack space no matter how many times it recurses.

%% NOT tail-recursive: the multiplication happens *after* the
%% recursive call returns, so a stack frame must be kept for it
fact(0) -> 1;
fact(N) -> N * fact(N - 1).

%% tail-recursive: the recursive call is the very last thing this
%% clause does — there is nothing left to do with its result except
%% return it, so no frame needs to be kept
fact_tail(N) -> fact_tail(N, 1).
fact_tail(0, Acc) -> Acc;
fact_tail(N, Acc) -> fact_tail(N - 1, N * Acc).

Both compute the same thing and both are used in this course — fact/1's shape is completely fine for small, bounded recursion, and its accumulator-passing sibling is what you reach for once the recursion depth is genuinely unbounded, which for a mesh of several thousand nodes, it routinely will be.

A typical language vs. Erlang

Go and Ruby both have for loops with mutable loop variables; writing the equivalent accumulation is a variable you reassign each iteration. Erlang cannot reassign a variable at all (Section 2.1), so the accumulator has to be threaded through as an extra function argument instead — which looks like more ceremony for a five-line function and stops looking that way the moment the accumulator needs to be something more interesting than a running total, because it is then just an ordinary function parameter with an ordinary type, not a mutable variable whose type you have to keep straight across dozens of loop iterations by eye.

Exercise 2.3

Write my_reverse/1, reversing a list, tail-recursively (accumulator-passing), without using the built-in lists:reverse/1.

Solution
my_reverse(List) -> my_reverse(List, []).
my_reverse([], Acc) -> Acc;
my_reverse([H | T], Acc) -> my_reverse(T, [H | Acc]).

Each step moves the head of the input onto the front of the accumulator, which is exactly reversal happening one element at a time — verified: my_reverse([1,2,3]) gives [3,2,1].

2.4 Maps

A map is Erlang's key-value structure, added relatively recently (OTP 17, 2015) as a friendlier alternative to the older, more rigid record and property-list idioms for "a bag of named fields" — the shape this course's per-node state will mostly take.

1> Node = #{id => 42, status => alive, energy => 100}.
#{id => 42,status => alive,energy => 100}
2> #{id := Id, status := Status} = Node.
#{energy => 100,id => 42,status => alive}
3> Id.
42
4> Node2 = Node#{status := crashed}.
#{id => 42,status => crashed,energy => 100}

=> is used when constructing or when a key may or may not already be present; := is used when matching or updating a key that must already exist — updating a genuinely new key with := is a runtime error, which is a real, useful guard against typos in a field name silently creating a new field instead of updating the one you meant. Node#{status := crashed} produces a new map, leaving Node itself untouched — maps, like everything else bound to a variable, are immutable once created.

Exercise 2.4

Starting from Node = #{id => 1, status => alive, energy => 100}, write an expression that produces a new map with energy set to 80, using :=. Then predict, and verify, what happens if you try the same update against the key score, which does not exist in Node.

Solution
1> Node = #{id => 1, status => alive, energy => 100}.
#{id => 1,status => alive,energy => 100}
2> Node#{energy := 80}.
#{id => 1,status => alive,energy => 80}
3> Node#{score := 0}.
** exception error: bad key: score
     in function  maps:update/3

The third line is the point of the exercise: := against a key that is not already present is a runtime error, not a silent insert — the guard against a typo'd field name creating a brand-new field instead of updating the one you meant, exactly as Section 2.4 describes. Constructing Node#{score => 0} with => instead would succeed and genuinely add the key, which is the tell for which operator you actually meant to use.

2.5 Functions: multiple clauses and guards

You saw this shape already in classify/1 above: a function can be defined as several clauses, each with its own pattern for the arguments, tried top to bottom until one matches. A guard — the when clause — adds a further condition that must also hold, drawn from a restricted set of side-effect-free, always-fast operations (comparisons, arithmetic, type tests) precisely so that a guard can never itself crash or hang while the runtime is deciding which clause to run.

describe(N) when is_integer(N), N > 0 -> positive_integer;
describe(N) when is_integer(N) -> non_positive_integer;
describe(N) when is_float(N) -> float_value;
describe(N) when is_atom(N) -> atom_value;
describe(_) -> something_else.

This is Erlang's real substitute for both function overloading and a chain of if/ else if — and it is genuinely the idiomatic way to branch, not merely an alternative to it. Milestone 3 onward, nearly every message a node handles is dispatched this way: one function clause per message shape.

A guard cannot call your own functions — the compiler rejects it, it does not just misbehave

The restricted set of operations a guard is allowed to use is enforced at compile time, not left as a style convention to remember: calling an ordinary function you wrote, however small and however obviously side-effect-free, inside a when clause is a compile error, not a warning:

helper(N) -> N > 0.
check(N) when helper(N) -> yes;
check(_) -> no.

$ erlc badguard.erl
badguard.erl:2: call to local/imported function helper/1 is illegal in guard

The fix is either inlining the check directly as a guard expression (N > 0, which is guard-legal), or moving the logic into the function body and checking it there with an if or a nested case — a guard's restriction to a fixed, safe operation set is not negotiable by writing a "safe-looking" helper function around it.

Exercise 2.5

Write is_valid_energy/1, returning true only for integers in the inclusive range 0 to 100, using a guard — no if, no function body logic.

Solution
is_valid_energy(N) when is_integer(N), N >= 0, N =< 100 -> true;
is_valid_energy(_) -> false.

Verified: true for 50, false for both 150 and -1 — three guard conditions joined by commas, all of which must hold, is ordinary boolean "and" inside a guard; a comma-separated guard sequence never short-circuits in a way that matters here because every condition used is cheap and side-effect-free by construction, which is precisely what guards are restricted to in the first place.

2.6 Modules and exports

You have already seen the whole mechanism: -module(name) must match the filename, and -export([f/1, g/2]) lists exactly which name/arity pairs are callable from outside. Two more attributes worth knowing now:

-module(mesh_node).
-behaviour(gen_server).       %% declares intent; checked at compile time
                               %% once the callbacks below exist — Milestone 5

-export([start_link/1, stop/1]).   %% the public API
-export([init/1, handle_call/3]).  %% gen_server callbacks — technically
                                    %% exported so OTP's machinery can call
                                    %% them, not meant for other modules to
                                    %% call directly; the underscore-free
                                    %% naming convention signals that

-behaviour(gen_server) does not exist yet in code you will write until Milestone 5, but it is worth previewing the shape now: a behaviour is a contract — a fixed set of callback functions a module promises to implement — and the compiler warns you at compile time if you declare one and forget a required callback. This is the closest thing in Erlang to Go's interfaces, with one inversion worth noting: a Go interface is satisfied implicitly, by having the right methods; an Erlang behaviour is declared explicitly, and the compiler checks the declaration against the module's actual exports.

The filename and the -module name must match exactly, or nothing compiles

This is one of the first errors most newcomers to Erlang hit, and it looks nothing like the mistake that caused it: save a module declared -module(actualname) in a file named wrongname.erl and the compiler refuses outright, before checking anything else about the code:

$ erlc wrongname.erl
wrongname.beam: Module name 'actualname' does not match file name 'wrongname'

Easy to hit by copy-pasting an existing module as a starting point for a new one and forgetting to update the -module line to match the new filename — the fix is making the two agree, in either direction, not a sign anything else is wrong with the code itself.

Exercise 2.6

Write a module with one exported function that calls a second, unexported "helper" function internally. Confirm the exported function works normally when called from another module, then try calling the helper function directly from that other module and explain what happens.

Solution
-module(expmod).
-export([public_fn/0]).

public_fn() -> internal_fn() + 1.
internal_fn() -> 41.
1> expmod:public_fn().
42
2> expmod:internal_fn().
** exception error: undefined function expmod:internal_fn/0

public_fn/0 calling internal_fn/0 from inside the same module needs no export at all — the export list only governs what code outside the module can reach. Calling internal_fn/0 the same way you would call any other function, from the shell or from a different module, fails with undef, because as far as anything outside expmod is concerned, that function does not exist.

2.7 Processes: spawn, !, and receive

This is the syntax Milestone 3 will build a real node architecture out of. Three primitives, and nothing else is needed to create concurrency in Erlang — no thread pool to configure, no async keyword to remember to add.

-module(echo).
-export([loop/0]).

loop() ->
    receive
        {From, Ref, Msg} ->
            From ! {Ref, {echo, Msg}},
            loop();
        stop ->
            ok
    end.
1> Pid = spawn(echo, loop, []).
<0.94.0>
2> Ref = make_ref().
#Ref<0.123.456.789>
3> Pid ! {self(), Ref, hello}.
{<0.85.0>,#Ref<0.123.456.789>,hello}
4> flush().
Shell got {#Ref<0.123.456.789>,{echo,hello}}
ok

spawn/3 starts a new, independent process running echo:loop() and returns its process identifier (a pid) immediately, without waiting for it to do anything — spawning is not a blocking operation and there is no thread limit to worry about hitting. ! (pronounced "bang") sends a message: it copies the term on its right into the target process's mailbox — an unbounded, per-process FIFO queue that only that process reads from — and returns immediately, whether or not anyone is listening. receive blocks the calling process until a message matching one of its patterns arrives in its own mailbox, then runs the matching clause; a bare loop() as the tail of each clause is exactly the tail-recursion from Section 2.3, and it is what makes this an ongoing server rather than a one-shot responder.

The reply pattern — {From, Ref, Msg} in, {Ref, Reply} back — is not a special language feature, it is a convention, and it is worth understanding why the Ref is there at all rather than replying to From directly. make_ref/0 creates a value guaranteed unique for the lifetime of the runtime; including it lets the caller distinguish a reply to this specific request from some unrelated message that happens to also be addressed to it, which matters the moment a process can have several requests in flight at once. gen_server, met in Milestone 5, does exactly this under the hood, automatically.

Selective receive

receive
    {Ref, Reply} -> Reply     %% only matches a message tagged with THIS Ref
after 1000 ->
    timeout
end

receive does not take the next message in the mailbox unconditionally — it scans for the first message that matches one of its clauses, leaving anything that does not match sitting in the mailbox for a later receive to find. This is selective receive, and the Ref convention above is precisely what makes it possible to say "wait specifically for the reply to this request" in a process that might have other, unrelated messages arriving in the meantime. after 1000 -> bounds the wait to one second, returning timeout if nothing matching arrives in time — without it, a reply that never comes blocks the caller forever.

A mailbox with no selective clause for a message keeps that message forever

If a process's receive only ever matches messages of one shape, and something sends it a message of a different shape, that message is not discarded — it sits in the mailbox, permanently, skipped by every future receive that also does not match it. Enough of these accumulate and the mailbox itself becomes a genuine memory leak, and every receive after it gets slightly slower, because a selective receive has to scan past every unmatched message to find one that does match. Milestone 7 measures exactly this cost once the mesh is under load; the fix, covered there, is a catch-all clause that at least logs and discards anything unrecognised, rather than a mailbox with no fallback at all.

A typical language vs. Erlang

C++ and Java both give you threads that share the process's memory by default — a thread reads and writes the same objects another thread can reach, and correctness depends on remembering to guard every shared access with a mutex, an atomic, or some other explicit synchronisation primitive you have to choose and apply consistently yourself. JavaScript's runtime avoids that specific hazard by having only one thread of execution at a time (concurrency there is about interleaving callbacks and await points on a single thread, not simultaneous memory access), which sidesteps data races but also means one long-running callback blocks everything else. Erlang processes share nothing by default — no object either side can reach through the other, only messages explicitly copied across — so "did I forget to lock this" is not a category of bug that exists here at all; the honest cost is that this copying is real work (a large term sent between processes is genuinely copied, not just referenced), and a design that leans on huge messages passed constantly between processes pays for that isolation in a way a shared-memory design with careful locking would not.

2.8 Links, monitors, and trap_exit

Two mechanisms exist for one process to learn that another one died, and they are not interchangeable — Milestone 4 is built entirely on the distinction.

LinkMonitor
Directionbidirectional — either side dying affects the otherone-directional — only the watcher is notified
Default effect of the other side dyingyou die too, propagating the exit signal onward, unless you set trap_exityou receive a {'DOWN', ...} message; you are never killed by it
Used forsupervision — a supervisor and its children are always linkedobservation without ownership — the registry watching nodes it does not supervise
1> process_flag(trap_exit, true).
false
2> {Pid, Ref} = spawn_monitor(fun() -> exit(boom) end).
{<0.101.0>,#Ref<0.123.456.789>}
3> flush().
Shell got {'DOWN',#Ref<0.123.456.789>,process,<0.101.0>,boom}
ok

process_flag(trap_exit, true) is what turns a link from "their death kills me too" into "their death sends me a plain {'EXIT', Pid, Reason} message instead" — this is exactly the flag every OTP supervisor sets, because a supervisor's entire job is to survive its children's deaths and react to them, which is the opposite of the default linked behaviour. spawn_monitor/1 both spawns and monitors in one call, convenient for exactly the "watch this thing, but I am not responsible for it" relationship the registry will have with nodes in Milestone 7.

2.9 Errors: try/catch, and when not to use it

1> try 1 / 0 catch error:badarith -> {error, division_by_zero} end.
{error,division_by_zero}
2> try throw(custom_signal) catch throw:Reason -> {caught, Reason} end.
{caught,custom_signal}

Erlang has three distinct ways to signal something exceptional — error (a genuine bug — division by zero, a failed pattern match, a bad argument), throw (a value used for non-local control flow, the sender expects it to be caught somewhere), and exit (a deliberate request that a process should stop, which is also what an unhandled error becomes) — and try ... catch Class:Reason -> ... can distinguish between them by matching on Class.

The honest advice, and the whole thesis of "let it crash": reach for try/ catch far less often than instinct suggests. Wrapping every fallible operation defensively is the instinct this language actively argues against — the idiomatic response to "this function can fail in a way I have not planned for" is usually to let the process crash and have a supervisor restart it into clean state, not to catch the failure and attempt to continue in a state you were not prepared for. try/catch earns its place at a genuine boundary: the edge of the system (a network request, user input), or a place where "fail and report a specific, recoverable reason" is itself the correct behaviour rather than "fail and restart."

Catching _:_ hides real bugs as reliably as it hides the failure you meant to catch

A wildcard pattern matches every class and every reason, which means it catches the specific, anticipated failure you were thinking about exactly as well as it catches a genuine programming mistake you were not — a typo'd variable, a bad arithmetic operation, anything:

1> try X = 1 + not_a_number, X catch _:_ -> ok end.
ok

1 + not_a_number is a real bug — adding an integer to an atom — and the wildcard catch above converts it into the same ok a genuinely expected, handled failure would produce, with nothing in the return value to tell them apart. Catching a specific Class:Reason pattern (as the two examples at the top of this section do) lets anything that does not match propagate and crash loudly, which — per this section's own advice — is usually the outcome you want for the case you did not anticipate.

Exercise 2.9

Write an expression that divides two numbers inside a try, catching only error:badarith specifically (not a wildcard), and returns {error, division_by_zero} on failure. Then call it with a division that raises a different error class (throw(not_a_number), say) and confirm your narrow catch does not swallow it.

Solution
safe_divide(A, B) ->
    try A / B
    catch error:badarith -> {error, division_by_zero}
    end.
1> safe_divide(10, 0).
{error,division_by_zero}
2> try throw(not_a_number) catch error:badarith -> {error, division_by_zero} end.
** exception throw: not_a_number

The second call is the point: a throw is a different exception class from error, so a catch pattern narrowed to error:badarith correctly lets it through uncaught, rather than a wildcard silently absorbing an exception the function was never actually written to handle.

2.10 A first taste of behaviours

You will not write a full gen_server until Milestone 5, but it is worth seeing the shape of an OTP behaviour once now, so Milestone 5 is recognising a pattern rather than meeting one cold. A behaviour is a module that implements a fixed set of callbacks; OTP's generic machinery — code you never see or modify — handles the process loop, the message protocol, and a long list of edge cases (what happens if a reply never comes, how to shut down cleanly) uniformly for every module that implements it.

%% the shape you will fill in for real in Milestone 5 — not runnable yet
-module(mesh_node).
-behaviour(gen_server).

init(Args) -> {ok, InitialState}.
handle_call(Request, From, State) -> {reply, Reply, NewState}.
handle_cast(Request, State) -> {noreply, NewState}.

Compare the raw receive loop from Section 2.7 to this: the loop itself, the reply-matching convention, the timeout handling — all of it is gone, replaced by three small functions that only describe what should happen, never how the message actually gets there. That is the entire value proposition of an OTP behaviour, and Milestone 5 is the moment you stop hand-writing the "how" and start trusting a library that has handled it correctly for four decades.

2.11 Testing with EUnit, properly

-module(mesh_node_tests).
-include_lib("eunit/include/eunit.hrl").

%% a fixture: setup, the tests that need it, then teardown — for when a
%% test needs a running process rather than a pure function
spawn_and_stop_test_() ->
    {setup,
     fun() -> {ok, Pid} = mesh_node:start_link(#{id => 1}), Pid end,
     fun(Pid) -> mesh_node:stop(Pid) end,
     fun(Pid) ->
         [?_assert(is_process_alive(Pid)),
          ?_assertEqual(alive, mesh_node:status(Pid))]
     end}.

The {setup, Setup, Cleanup, Tests} tuple is EUnit's fixture shape, roughly parallel to Go's table-driven tests with a shared TestMain, or RSpec's before/after — and it is the shape nearly every test in Milestones 3 onward will use, because nearly every test needs a real, running process to test against.

Part 2 checkpoint

  1. Why is a guard restricted to a small set of side-effect-free operations, rather than allowed to call any function?
  2. What is the actual difference between a linked process and a monitored one, and which relationship does a supervisor have with its children?
  3. Why does an unmatched message sitting forever in a mailbox get worse over time, specifically?
  4. Name the three ways Erlang signals something exceptional, and which one this course's philosophy argues you should reach for least.
  5. What does a behaviour's callback module not have to implement, compared to the raw receive loop version of the same server?

What is next

The Mewlang cat, typing at a laptopMilestone 1 turns the shell experiments above into a real, compiled, tested rebar3 project. Milestone 2 builds the pure data-handling functions the rest of the node will need. Milestone 3 is where a mesh node becomes a real, running process for the first time — everything from Section 2.7 onward, applied.

Instalment 16 of the five-course curriculum. Next: Erlang Milestones 1–4, where the shell experiments above become a real project, and a node becomes a process that can crash on purpose.

Continue