Instalment 16 · Course 4 (Erlang) · Parts 0–2
A mesh of a thousand independent nodes that send each other messages, and periodically, deliberately, fall over. Supervisors notice. They come back. Nothing else in this curriculum treats failure as a first-class, expected, routinely-exercised code path the way this one does.
A simulated network, mesh, of up to several thousand independent nodes, each an Erlang process. Nodes exchange messages, maintain a small amount of state, and periodically misbehave: a node can crash outright, hang and stop responding, send another node a message it cannot parse, or simply vanish. A supervision tree watches every node and restarts the ones that die, with a policy that says precisely how much dying is tolerable before the supervisor gives up and takes the whole subtree down with it.
On top of that: a registry so nodes can find each other by name instead of by process identifier: an observability layer with counters for messages sent, nodes alive, nodes restarted, and message loss: a live dashboard, served over plain HTTP from a process running inside the same system it is reporting on, showing all of that updating in real time in a browser. Then, in the back half of the course, the single-machine simulation becomes a genuine distributed system: several real operating-system processes, each running its own instance of the Erlang virtual machine (the BEAM), connected over a real network, able to survive one of those machines being killed outright.
$ ./bin/mesh start --nodes 2000 --dashboard :8080
mesh: 2000 nodes started, supervisor tree depth 3
mesh: dashboard listening on http://localhost:8080
$ ./bin/mesh chaos --crash-rate 0.02 --partition-every 30s
chaos: injecting failures into a live system
alive: 1,978 restarting: 14 partitioned: 8 msgs/sec: 40,221
Everything in that output is a real, measured number from a real run once the project exists. Nothing here is a diagram of an idea; every milestone below produces working, running, compiled code.
Every other course in this curriculum treats a crash as a bug to eliminate. This one treats it as a certainty to plan for. That is not a small difference in emphasis — it is a different starting axiom for the whole design. A Go service that panics takes the whole process down with it, and the discipline is to never let a goroutine panic in the first place. An Erlang process that crashes is routine: by the end of Milestone 4 you will have a system where processes die constantly, on purpose, in the test suite, and the assertion is never "no process died" — it is "the system as a whole kept working while processes died."
That reframing is the entire subject of this course. Once you accept that any single process can and will die, the interesting engineering question stops being "how do I prevent this" and becomes "who notices, what do they do about it, and how fast." The answer Erlang and its OTP (Open Telecom Platform) library give is supervision trees: a static, declarative description of which processes depend on which, and exactly what should happen — restart just this one, restart its siblings too, give up entirely — when one of them dies. You will hand-write the naive version in Milestone 4, before you have been told what a supervisor is, discover exactly what it gets wrong, and then meet OTP's supervisor behaviour in Milestone 6 as the answer to a problem you have already felt.
Three properties of the language are not available, in combination, anywhere else in this curriculum.
Processes are cheap and share nothing. An Erlang process is not an OS thread and not a goroutine — it is a lightweight, independently garbage-collected unit with its own heap, started in microseconds, and a BEAM instance can comfortably run millions of them. Because no process can reach into another's memory — the only way in is a message, copied, into a mailbox — a crashed process cannot corrupt anything belonging to a process that did not crash. Go's goroutines are cheap too, but they share an address space, so a wild pointer write in one can in principle corrupt data another is reading; the go runtime's response to an unrecovered panic in any goroutine is to kill the entire program. Erlang's response to an unhandled exception in a process is to kill exactly that process, and nothing else, ever.
Let it crash is a real engineering strategy, not a slogan. Given the isolation above, the language actively discourages the defensive-programming instinct of catching every possible error and trying to handle it in place. If a node receives a message it cannot make sense of, the idiomatic move is to let the process crash — loudly, with a stack trace, to a log — and let the supervisor restart it into a known-good state, rather than trying to enumerate every way input can be wrong and patch around each one. Milestone 8 makes this concrete: you will feed the mesh genuinely malformed messages and watch the system absorb them by discarding and restarting, rather than by defending.
Distribution was never an afterthought. The message-passing model between two processes on one machine and two processes on different machines is, by design, close to the same API — Erlang was built at Ericsson in the late 1980s specifically to run telephone switches that were not allowed to go down, and those switches were distributed hardware from day one. Milestone 11 turns your single-machine mesh into several real operating-system processes talking over a real network with a change measured in tens of lines, not a rewrite — because the message-passing discipline you have been using since Milestone 3 was already, quietly, the distributed-systems discipline.
Be honest about the trade this makes. Erlang's processes are wonderful for isolation and terrible for sharing a large, frequently-updated data structure between many workers — there is no equivalent of Go's atomic-read sharded engine from Course 1, because sharing memory is exactly what the model refuses to do. Where Go would shard a struct behind atomics for raw throughput, Erlang would shard the problem into more processes and accept the cost of a message copy at the boundary. For a system whose entire premise is "parts of it are broken right now," that trade is the right one: you are buying certainty about blast radius, and you are paying for it in copying overhead. Both are measured, in this course, rather than asserted.
mesh_sup (one_for_one)
/ | \
node_sup_1 registry_sup dashboard_sup
(simple_one (name -> pid, (http listener,
_for_one) process groups) metrics collector)
/ | \
node node node ... (up to several thousand,
| | | each a gen_server, each linked
mailbox mailbox mailbox to node_sup_1, each able to
crash independently)
every arrow above is a supervision link: a child's death is
reported to its supervisor, which decides — by a declared,
static policy — whether to restart just that child, restart
every child alongside it, or give up and die itself, which
reports the death one level further up
Three architectural decisions, explained now because they shape every milestone that follows.
A node is a gen_server, OTP's standard behaviour for "a process with state that answers requests" — you will write the raw receive loop version first in Milestone 3, by hand, before meeting gen_server in Milestone 5 as the same idea with the boilerplate handled for you and a long list of footguns removed.
Every node sits under a supervisor, and supervisors sit under supervisors. This is the tree in "supervision tree": node_sup_1 is a simple_one_for_one supervisor, a template for spawning an unbounded, dynamic number of identical children — exactly what a few thousand interchangeable network nodes need — while mesh_sup at the root uses one_for_one, because the registry, the node pool and the dashboard are three genuinely different things that should not be restarted as a group just because one of them died.
The registry and the dashboard are processes too, supervised exactly like the nodes they serve. There is no special "infrastructure" tier that gets to skip the discipline the rest of the system follows — if the registry crashes, something notices and restarts it, the same as everything else.
| # | Milestone | What it teaches |
|---|---|---|
| 1 | The shell and a first module | erl, expressions, atoms, compiling, rebar3 |
| 2 | Functions over data | pattern matching, immutability, lists/tuples/maps, recursion, tail calls, guards |
| 3 | A node as a raw process | spawn, !, receive, mailboxes, selective receive |
| 4 | Let it crash | links, monitors, exit signals, trap_exit, why defensive code is discouraged |
| 5 | The node as a gen_server | OTP behaviours, call vs cast, state, timeouts |
| 6 | Supervision trees | restart strategies, intensity limits, dynamic children, startup ordering |
| 7 | A network of a thousand nodes | registries, ETS, process groups, routing, message loss |
| 8 | Chaos injection | random crashes, slow handlers, malformed messages, memory pressure |
| 9 | Observability | counters, observer, recon, tracing a live system safely |
| 10 | The live dashboard | an HTTP server in Erlang, streaming updates, aggregation without slowing the system |
| 11 | Actually distributed | multiple BEAM nodes, cookies, net_kernel, partitions, split brain |
| 12 | Releases and property tests | EUnit, Common Test, PropEr, relx releases, upgrades |
Why isolating failure is a design decision with a measurable cost, not a free lunch; how to describe a system's failure-response policy declaratively instead of scattering try/catch through it; how OTP's standard behaviours turn "a process that holds state and answers requests" from a hand-rolled loop into a few callback functions; what actually changes, in code, when a message-passing system stops being single-machine and becomes genuinely distributed; and why a network partition is not a bug you fix but a condition you are required to have an opinion about.
Erlang is a dynamically typed, functional, concurrent language built at Ericsson starting in 1986 to program telephone exchanges — hardware that was contractually required to have at most a few minutes of downtime a year. Two consequences follow directly from that origin and explain nearly everything the language does differently from a general-purpose scripting or systems language. First, the runtime (the BEAM, Erlang's virtual machine) was built around lightweight, isolated processes and message passing from the start, because a telephone exchange is naturally a huge number of mostly independent, concurrent calls. Second, the language and its standard library, together known as OTP (Open Telecom Platform), encode decades of "how do you build a system that must not go down" as reusable behaviours — gen_server, supervisor, and others — rather than leaving every team to reinvent them.
Erlang is not "the old version of Elixir." Elixir is a newer language that compiles to the same BEAM bytecode and can call Erlang code directly (and vice versa); Erlang came first, by about 25 years, and this course uses it directly rather than through Elixir's syntax, because the OTP behaviours and the supervision-tree model — the actual subject of this course — are identical either way, and Erlang's syntax, once past the first hour of unfamiliarity, gets out of the way of that subject rather than adding a second one.
You want a recent OTP — this course was verified on OTP 25 (Erlang/OTP 25, erts-13.1.5) — and, separately, rebar3, the standard build tool and package manager, roughly equivalent to Go's own toolchain or Ruby's Bundler.
sudo apt install erlang rebar3 # Debian/Ubuntu
sudo dnf install erlang rebar3 # Fedora
erl -eval 'io:format("~s~n", [erlang:system_info(system_version)]), halt().' -noshell
# Erlang/OTP 25 [erts-13.1.5] [source] [64-bit] [smp:N:N] [ds:N:N:N] [jit:ns]
brew install erlang rebar3
:: the official installer from erlang.org, or:
winget install ErlangSolutions.Erlang
:: rebar3 is a self-contained escript; download rebar3 and put it on PATH
WSL2 is worth it here, as it was for Perl: observer (Milestone 9) is a GUI tool that is simplest to run from a proper Linux environment, and the distributed-node work in Milestone 11 involves Unix-y networking assumptions that are easiest to reason about on Linux.
curl -fsSL https://raw.githubusercontent.com/kerl/kerl/master/kerl -o kerl
chmod +x kerl
./kerl build 25.3.2.9 25.3.2.9
./kerl install 25.3.2.9 ~/erlangs/25.3.2.9
. ~/erlangs/25.3.2.9/activate
kerl is the Erlang-world equivalent of Perl's perlbrew or Go's own version management: it builds a specific OTP release from source into an isolated directory and gives you an activate script to bring it onto your PATH.
Every other course in this curriculum treats its language's REPL as a nice-to-have for quick experiments. Erlang's shell, erl, is closer to a permanent fixture of how you work: it can connect to a running production system and let you call functions, inspect process state, and even hot-load new code into it without stopping it — a capability the telephone-exchange origin story makes obvious in retrospect. You will use it constantly, not just at the start.
$ erl
Erlang/OTP 25 [erts-13.1.5] [source] [64-bit] [smp:6:6] [jit:ns]
Eshell V13.1.5 (abort with ^G)
1> 1 + 1.
2
2> X = 40 + 2.
42
3> X = 42.
42
4> X = 43.
** exception error: no match of right hand side value 43
That last line is not a typo in this document — it is the single most important fact about the language, arriving in the fourth line of the shell. = is not assignment. It is pattern matching: X = 42 succeeds because X was unbound and binds it to 42; X = 42 again succeeds because 42 matches what X is already bound to; X = 43 fails, loudly, with an exception, because 43 does not match 42. Once a variable is bound in Erlang, it stays bound to that value for the rest of its scope — there is no reassignment, anywhere, ever. Section 2.1 below is entirely about what this buys you.
rebar3$ rebar3 new app mesh
$ tree mesh
mesh/
├── rebar.config dependencies and build config (like a Gemfile or go.mod)
├── src/
│ ├── mesh.app.src application metadata: name, modules, dependencies
│ ├── mesh_app.erl the application behaviour's entry point
│ └── mesh_sup.erl the top-level supervisor (empty for now)
└── test/
rebar3 new app scaffolds an OTP application — not just a script, but a unit with a declared name, a supervision tree entry point, and a place for dependencies, tests and releases to live. That is deliberately more structure than Milestone 1 needs; the scaffold is there because by Milestone 6 you will need every part of it, and generating it once now means never having to retrofit it later.
%% src/hello.erl
-module(hello).
-export([world/0]).
world() ->
io:format("hello, ~s~n", [node()]).
$ erlc hello.erl
$ erl -noshell -eval 'hello:world(), halt().'
hello, nonode@nohost
-module(hello). declares the module name, and it must match the filename (hello.erl) — the compiler enforces this, unlike Go's package-directory convention which is a strong suggestion rather than a rule.-export([world/0]). lists which functions are visible outside this module, by name and arity — world/0 is "the zero-argument function named world", a genuinely different, independently-exportable function from a hypothetical world/1. Nothing not listed here is callable from outside the module at all; there is no equivalent of Go's capitalisation convention, this is enforced by the compiler.io:format("hello, ~s~n", [node()]) calls the format function in the standard library's io module. ~s is a format directive meaning "a string here"; ~n is a portable newline; the second argument is always a list of values, one per directive, which is why a single argument still needs its own brackets.node() returns the name of the current Erlang node — the runtime concept, not this course's "network node" concept, and the collision between those two uses of the word "node" is real and worth flagging now rather than mid-sentence in Milestone 11. nonode@nohost is what an Erlang VM is called when it was started without a name, which is the normal case until Milestone 11 gives every VM instance a real one.This course's simulated network is made of mesh nodes — ordinary Erlang processes, thousands of them, all living inside one BEAM instance for most of the course. Erlang's own runtime separately has a notion of a distributed node — one running BEAM instance, identified by a name like mesh1@192.168.1.10, that can connect to other BEAM instances over the network. Until Milestone 11 there is exactly one distributed node and thousands of mesh nodes running inside it; from Milestone 11 onward there are several distributed nodes, each still running many mesh nodes. The text says "process" whenever precision matters and "mesh node" for the simulated concept; "node" alone, unqualified, always means the Erlang runtime sense.
$ erl
1> h(lists, map).
...
2> lists:map(fun(X) -> X * 2 end, [1,2,3]).
[2,4,6]
$ erl -man gen_server # the full behaviour reference, as a man page
$ open https://www.erlang.org/doc/ # the official docs, browsable
h/2 inside the shell prints a function's documentation without leaving your session — the closest thing here to Perl's perldoc -f or Go's go doc. The official docs at erlang.org are unusually good and unusually stable: OTP's standard library has unusually few breaking changes release to release, so documentation from three OTP releases ago is still mostly correct today.
Erlang's editor story is smaller than Go's "every editor has first-class gopls support out of the box," but it is genuinely solid once set up. Two real options, either of which is enough for this course:
erlang_ls extension (marketplace id erlang-ls.erlang-ls) wraps the Erlang Language Server — inline compile errors as you type, go-to-definition, autocomplete over both the standard library and your own project's modules. WhatsApp's newer ELP (Erlang Language Platform) targets the same job and scales better to very large codebases; for a project this size either works, and erlang_ls is the more widely documented default.erlang-mode ships inside the Erlang/OTP source distribution itself (lib/tools/emacs/erlang.el), not as a separate package to track — install Erlang from your package manager as shown above and it is usually already on disk somewhere under the installation, just needing to be added to Emacs's load path. It predates every other option here by decades and much of the wider Erlang community's own tooling still assumes it as the baseline.Both integrate with rebar3 directly: point either one at a project with a rebar.config already in place (Milestone 1 has you create one) and it finds the project's own dependencies and include paths without further configuration.
$ cat rebar.config
{erl_opts, [debug_info]}.
{plugins, [rebar3_format]}. %% or: {plugins, [erlfmt]}.
$ rebar3 format # rebar3_format: reformats in place, project-wide
$ rebar3 fmt -w src/*.erl # erlfmt: reformats the given files in place
$ rebar3 fmt --check # erlfmt: reports files that need formatting, changes nothing
Neither formatter ships with rebar3 itself — both are plugins, added once to rebar.config the same way rebar3_proper is added in Milestone 12. They disagree on some stylistic choices (rebar3_format, for instance, tends to keep small map literals on one line where erlfmt more often expands them), so pick one for a given project rather than running both — mixing them means every commit reformats whatever the previous commit's tool did not agree with.
OTP ships EUnit, a lightweight unit-testing framework, in the standard library — no dependency to add, the same way Go ships testing.
%% src/basics.erl
-module(basics).
-export([classify/1]).
classify(N) when N < 0 -> negative;
classify(0) -> zero;
classify(N) when N rem 2 =:= 0 -> even;
classify(_) -> odd.
%% test/basics_tests.erl
-module(basics_tests).
-include_lib("eunit/include/eunit.hrl").
classify_test() ->
?assertEqual(negative, basics:classify(-1)),
?assertEqual(zero, basics:classify(0)),
?assertEqual(even, basics:classify(4)),
?assertEqual(odd, basics:classify(7)).
$ rebar3 eunit
===> Running EUnit tests...
basics_tests: classify_test...[0.001 s] ok
[done in 0.012 s]
=======================================================
All 1 tests passed.
-include_lib("eunit/include/eunit.hrl") pulls in the ?assertEqual family of macros — yes, Erlang has a preprocessor and macros starting with ?, inherited from its C heritage and used sparingly outside test code. A function whose name ends in _test (singular) is a single test; one ending in _test_ (trailing underscore) is a test generator, returning a list of tests built programmatically — you will use both shapes throughout this course.
= actually do in Erlang, and why does X = 43 fail after X = 42?node() the function returns.foo_test and one named foo_test_?You met it above: = matches rather than assigns. The same mechanism — not a special case of it, the literal same operation — is how you pull values out of compound data, choose between function clauses, and receive messages. There is exactly one destructuring rule in the whole language, used everywhere.
1> {Status, Value} = {ok, 42}.
{ok,42}
2> Status.
ok
3> [First | Rest] = [1,2,3,4].
[1,2,3,4]
4> First.
1
5> Rest.
[2,3,4]
6> {ok, 42} = {ok, 43}.
** exception error: no match of right hand side value {ok,43}
Line 6 is not a corner case to remember — it is the mechanism working as designed. Matching a literal against a value, rather than a fresh variable against it, asserts that the value has exactly that shape. This is how Erlang code routinely reads: a function expects its argument to look like {ok, Result} or its message to look like {From, Ref, Request}, and if it does not, the match fails and the process crashes — which, per this course's whole premise, is the correct response, not a bug to prevent.
Compare this to how a typical language handles the same job:
In Python or Ruby you would write status, value = pair and then a separate if status != "ok": raise ... to assert the shape you expected. Erlang's match is the assertion — there is no separate step, because the destructuring and the shape check are the same operation. The cost is that a bound variable can never be reused for something else in the same scope; the benefit is that "this data has the shape I assumed" is enforced at the exact point you assume it, not several lines later when something built from the wrong assumption finally breaks.
In the shell, bind Point = {3, 4}. Write a pattern match that extracts both coordinates into X and Y in one line, then write an expression using only pattern matching (no if, no function call) that succeeds if and only if the point is exactly the origin, {0, 0}.
1> Point = {3, 4}.
{3,4}
2> {X, Y} = Point.
{3,4}
3> {0, 0} = Point.
** exception error: no match of right hand side value {3,4}
The third line is the answer to the second half: it succeeds silently for the origin and raises for anything else, which is "succeeds if and only if". Wrapping it in catch {0,0} = Point converts the exception into an ordinary value if you need the result rather than the side effect of not crashing.
An atom is a name that is its own value — ok, error, undefined, mesh_node — comparable in spirit to a Ruby symbol, with no equivalent that is quite as central in Go or Perl. Atoms are used everywhere a smaller language would reach for a string constant or a small integer enum, and OTP's standard library leans on one convention hard enough that you should adopt it immediately: a function that can fail returns {ok, Result} on success and {error, Reason} on failure, as a two-element tuple, rather than throwing.
1> file:read_file("does_not_exist.txt").
{error,enoent}
2> file:read_file("/etc/hostname").
{ok,<<"my-machine\n">>}
Matching directly against the shape you expect is the idiomatic way to consume this:
{ok, Contents} = file:read_file(Path),
%% ... use Contents; if the read failed, the match fails and the
%% process crashes here, with a clear reason, which is correct —
%% see the "let it crash" section above.
<<"my-machine\n">> is a binary, Erlang's efficient representation for raw byte data — the type file contents, network payloads and, later in this course, malformed chaos-test messages arrive as.
{error, Reason} tuple is not raised — it just sits there until you match on itComing from a language where a failed file read throws by default, it is easy to assume file:read_file/1 announces failure the same way. It does not — {error, enoent} is an entirely ordinary value, returned exactly like a success would be, and code that does not explicitly check for it will happily keep going with an error tuple sitting where real data was expected:
1> R = file:read_file("does_not_exist.txt").
{error,enoent}
2> {ok, Bin} = R.
** exception error: no match of right hand side value {error,enoent}
The crash above is the correct outcome, and it is the pattern match — {ok, Bin} = R — doing the work of announcing the failure, not the original call. Skip the match (bind Data = file:read_file(Path) and use Data directly, unmatched) and the error tuple flows silently into whatever uses it next, surfacing as a confusing failure far away from where it actually originated.
Write describe_result/1, taking a value shaped like {ok, Value} or {error, Reason} and returning {success, Value} or {failure, Reason} respectively — using pattern matching in the function head, not a guard or an if.
describe_result({ok, Value}) -> {success, Value};
describe_result({error, Reason}) -> {failure, Reason}.
Verified: describe_result({ok, 42}) gives {success,42}; describe_result({error, not_found}) gives {failure,not_found}. Two clauses, no branching logic written by hand — exactly Section 2.5's "multiple clauses instead of if" idiom, one section early.
Erlang has no loop construct at all — no for, no while. Repetition is always recursion, which sounds alarming until you meet the guarantee that makes it practical: a tail call — a recursive call in the last position of a function clause, with nothing left to do after it returns — is compiled to a jump, not a new stack frame, so a correctly-written recursive loop runs in constant stack space no matter how many times it recurses.
%% NOT tail-recursive: the multiplication happens *after* the
%% recursive call returns, so a stack frame must be kept for it
fact(0) -> 1;
fact(N) -> N * fact(N - 1).
%% tail-recursive: the recursive call is the very last thing this
%% clause does — there is nothing left to do with its result except
%% return it, so no frame needs to be kept
fact_tail(N) -> fact_tail(N, 1).
fact_tail(0, Acc) -> Acc;
fact_tail(N, Acc) -> fact_tail(N - 1, N * Acc).
Both compute the same thing and both are used in this course — fact/1's shape is completely fine for small, bounded recursion, and its accumulator-passing sibling is what you reach for once the recursion depth is genuinely unbounded, which for a mesh of several thousand nodes, it routinely will be.
Go and Ruby both have for loops with mutable loop variables; writing the equivalent accumulation is a variable you reassign each iteration. Erlang cannot reassign a variable at all (Section 2.1), so the accumulator has to be threaded through as an extra function argument instead — which looks like more ceremony for a five-line function and stops looking that way the moment the accumulator needs to be something more interesting than a running total, because it is then just an ordinary function parameter with an ordinary type, not a mutable variable whose type you have to keep straight across dozens of loop iterations by eye.
Write my_reverse/1, reversing a list, tail-recursively (accumulator-passing), without using the built-in lists:reverse/1.
my_reverse(List) -> my_reverse(List, []).
my_reverse([], Acc) -> Acc;
my_reverse([H | T], Acc) -> my_reverse(T, [H | Acc]).
Each step moves the head of the input onto the front of the accumulator, which is exactly reversal happening one element at a time — verified: my_reverse([1,2,3]) gives [3,2,1].
A map is Erlang's key-value structure, added relatively recently (OTP 17, 2015) as a friendlier alternative to the older, more rigid record and property-list idioms for "a bag of named fields" — the shape this course's per-node state will mostly take.
1> Node = #{id => 42, status => alive, energy => 100}.
#{id => 42,status => alive,energy => 100}
2> #{id := Id, status := Status} = Node.
#{energy => 100,id => 42,status => alive}
3> Id.
42
4> Node2 = Node#{status := crashed}.
#{id => 42,status => crashed,energy => 100}
=> is used when constructing or when a key may or may not already be present; := is used when matching or updating a key that must already exist — updating a genuinely new key with := is a runtime error, which is a real, useful guard against typos in a field name silently creating a new field instead of updating the one you meant. Node#{status := crashed} produces a new map, leaving Node itself untouched — maps, like everything else bound to a variable, are immutable once created.
Starting from Node = #{id => 1, status => alive, energy => 100}, write an expression that produces a new map with energy set to 80, using :=. Then predict, and verify, what happens if you try the same update against the key score, which does not exist in Node.
1> Node = #{id => 1, status => alive, energy => 100}.
#{id => 1,status => alive,energy => 100}
2> Node#{energy := 80}.
#{id => 1,status => alive,energy => 80}
3> Node#{score := 0}.
** exception error: bad key: score
in function maps:update/3
The third line is the point of the exercise: := against a key that is not already present is a runtime error, not a silent insert — the guard against a typo'd field name creating a brand-new field instead of updating the one you meant, exactly as Section 2.4 describes. Constructing Node#{score => 0} with => instead would succeed and genuinely add the key, which is the tell for which operator you actually meant to use.
You saw this shape already in classify/1 above: a function can be defined as several clauses, each with its own pattern for the arguments, tried top to bottom until one matches. A guard — the when clause — adds a further condition that must also hold, drawn from a restricted set of side-effect-free, always-fast operations (comparisons, arithmetic, type tests) precisely so that a guard can never itself crash or hang while the runtime is deciding which clause to run.
describe(N) when is_integer(N), N > 0 -> positive_integer;
describe(N) when is_integer(N) -> non_positive_integer;
describe(N) when is_float(N) -> float_value;
describe(N) when is_atom(N) -> atom_value;
describe(_) -> something_else.
This is Erlang's real substitute for both function overloading and a chain of if/ else if — and it is genuinely the idiomatic way to branch, not merely an alternative to it. Milestone 3 onward, nearly every message a node handles is dispatched this way: one function clause per message shape.
The restricted set of operations a guard is allowed to use is enforced at compile time, not left as a style convention to remember: calling an ordinary function you wrote, however small and however obviously side-effect-free, inside a when clause is a compile error, not a warning:
helper(N) -> N > 0.
check(N) when helper(N) -> yes;
check(_) -> no.
$ erlc badguard.erl
badguard.erl:2: call to local/imported function helper/1 is illegal in guard
The fix is either inlining the check directly as a guard expression (N > 0, which is guard-legal), or moving the logic into the function body and checking it there with an if or a nested case — a guard's restriction to a fixed, safe operation set is not negotiable by writing a "safe-looking" helper function around it.
Write is_valid_energy/1, returning true only for integers in the inclusive range 0 to 100, using a guard — no if, no function body logic.
is_valid_energy(N) when is_integer(N), N >= 0, N =< 100 -> true;
is_valid_energy(_) -> false.
Verified: true for 50, false for both 150 and -1 — three guard conditions joined by commas, all of which must hold, is ordinary boolean "and" inside a guard; a comma-separated guard sequence never short-circuits in a way that matters here because every condition used is cheap and side-effect-free by construction, which is precisely what guards are restricted to in the first place.
You have already seen the whole mechanism: -module(name) must match the filename, and -export([f/1, g/2]) lists exactly which name/arity pairs are callable from outside. Two more attributes worth knowing now:
-module(mesh_node).
-behaviour(gen_server). %% declares intent; checked at compile time
%% once the callbacks below exist — Milestone 5
-export([start_link/1, stop/1]). %% the public API
-export([init/1, handle_call/3]). %% gen_server callbacks — technically
%% exported so OTP's machinery can call
%% them, not meant for other modules to
%% call directly; the underscore-free
%% naming convention signals that
-behaviour(gen_server) does not exist yet in code you will write until Milestone 5, but it is worth previewing the shape now: a behaviour is a contract — a fixed set of callback functions a module promises to implement — and the compiler warns you at compile time if you declare one and forget a required callback. This is the closest thing in Erlang to Go's interfaces, with one inversion worth noting: a Go interface is satisfied implicitly, by having the right methods; an Erlang behaviour is declared explicitly, and the compiler checks the declaration against the module's actual exports.
-module name must match exactly, or nothing compilesThis is one of the first errors most newcomers to Erlang hit, and it looks nothing like the mistake that caused it: save a module declared -module(actualname) in a file named wrongname.erl and the compiler refuses outright, before checking anything else about the code:
$ erlc wrongname.erl
wrongname.beam: Module name 'actualname' does not match file name 'wrongname'
Easy to hit by copy-pasting an existing module as a starting point for a new one and forgetting to update the -module line to match the new filename — the fix is making the two agree, in either direction, not a sign anything else is wrong with the code itself.
Write a module with one exported function that calls a second, unexported "helper" function internally. Confirm the exported function works normally when called from another module, then try calling the helper function directly from that other module and explain what happens.
-module(expmod).
-export([public_fn/0]).
public_fn() -> internal_fn() + 1.
internal_fn() -> 41.
1> expmod:public_fn().
42
2> expmod:internal_fn().
** exception error: undefined function expmod:internal_fn/0
public_fn/0 calling internal_fn/0 from inside the same module needs no export at all — the export list only governs what code outside the module can reach. Calling internal_fn/0 the same way you would call any other function, from the shell or from a different module, fails with undef, because as far as anything outside expmod is concerned, that function does not exist.
spawn, !, and receiveThis is the syntax Milestone 3 will build a real node architecture out of. Three primitives, and nothing else is needed to create concurrency in Erlang — no thread pool to configure, no async keyword to remember to add.
-module(echo).
-export([loop/0]).
loop() ->
receive
{From, Ref, Msg} ->
From ! {Ref, {echo, Msg}},
loop();
stop ->
ok
end.
1> Pid = spawn(echo, loop, []).
<0.94.0>
2> Ref = make_ref().
#Ref<0.123.456.789>
3> Pid ! {self(), Ref, hello}.
{<0.85.0>,#Ref<0.123.456.789>,hello}
4> flush().
Shell got {#Ref<0.123.456.789>,{echo,hello}}
ok
spawn/3 starts a new, independent process running echo:loop() and returns its process identifier (a pid) immediately, without waiting for it to do anything — spawning is not a blocking operation and there is no thread limit to worry about hitting. ! (pronounced "bang") sends a message: it copies the term on its right into the target process's mailbox — an unbounded, per-process FIFO queue that only that process reads from — and returns immediately, whether or not anyone is listening. receive blocks the calling process until a message matching one of its patterns arrives in its own mailbox, then runs the matching clause; a bare loop() as the tail of each clause is exactly the tail-recursion from Section 2.3, and it is what makes this an ongoing server rather than a one-shot responder.
The reply pattern — {From, Ref, Msg} in, {Ref, Reply} back — is not a special language feature, it is a convention, and it is worth understanding why the Ref is there at all rather than replying to From directly. make_ref/0 creates a value guaranteed unique for the lifetime of the runtime; including it lets the caller distinguish a reply to this specific request from some unrelated message that happens to also be addressed to it, which matters the moment a process can have several requests in flight at once. gen_server, met in Milestone 5, does exactly this under the hood, automatically.
receive
{Ref, Reply} -> Reply %% only matches a message tagged with THIS Ref
after 1000 ->
timeout
end
receive does not take the next message in the mailbox unconditionally — it scans for the first message that matches one of its clauses, leaving anything that does not match sitting in the mailbox for a later receive to find. This is selective receive, and the Ref convention above is precisely what makes it possible to say "wait specifically for the reply to this request" in a process that might have other, unrelated messages arriving in the meantime. after 1000 -> bounds the wait to one second, returning timeout if nothing matching arrives in time — without it, a reply that never comes blocks the caller forever.
If a process's receive only ever matches messages of one shape, and something sends it a message of a different shape, that message is not discarded — it sits in the mailbox, permanently, skipped by every future receive that also does not match it. Enough of these accumulate and the mailbox itself becomes a genuine memory leak, and every receive after it gets slightly slower, because a selective receive has to scan past every unmatched message to find one that does match. Milestone 7 measures exactly this cost once the mesh is under load; the fix, covered there, is a catch-all clause that at least logs and discards anything unrecognised, rather than a mailbox with no fallback at all.
C++ and Java both give you threads that share the process's memory by default — a thread reads and writes the same objects another thread can reach, and correctness depends on remembering to guard every shared access with a mutex, an atomic, or some other explicit synchronisation primitive you have to choose and apply consistently yourself. JavaScript's runtime avoids that specific hazard by having only one thread of execution at a time (concurrency there is about interleaving callbacks and await points on a single thread, not simultaneous memory access), which sidesteps data races but also means one long-running callback blocks everything else. Erlang processes share nothing by default — no object either side can reach through the other, only messages explicitly copied across — so "did I forget to lock this" is not a category of bug that exists here at all; the honest cost is that this copying is real work (a large term sent between processes is genuinely copied, not just referenced), and a design that leans on huge messages passed constantly between processes pays for that isolation in a way a shared-memory design with careful locking would not.
trap_exitTwo mechanisms exist for one process to learn that another one died, and they are not interchangeable — Milestone 4 is built entirely on the distinction.
| Link | Monitor | |
|---|---|---|
| Direction | bidirectional — either side dying affects the other | one-directional — only the watcher is notified |
| Default effect of the other side dying | you die too, propagating the exit signal onward, unless you set trap_exit | you receive a {'DOWN', ...} message; you are never killed by it |
| Used for | supervision — a supervisor and its children are always linked | observation without ownership — the registry watching nodes it does not supervise |
1> process_flag(trap_exit, true).
false
2> {Pid, Ref} = spawn_monitor(fun() -> exit(boom) end).
{<0.101.0>,#Ref<0.123.456.789>}
3> flush().
Shell got {'DOWN',#Ref<0.123.456.789>,process,<0.101.0>,boom}
ok
process_flag(trap_exit, true) is what turns a link from "their death kills me too" into "their death sends me a plain {'EXIT', Pid, Reason} message instead" — this is exactly the flag every OTP supervisor sets, because a supervisor's entire job is to survive its children's deaths and react to them, which is the opposite of the default linked behaviour. spawn_monitor/1 both spawns and monitors in one call, convenient for exactly the "watch this thing, but I am not responsible for it" relationship the registry will have with nodes in Milestone 7.
try/catch, and when not to use it1> try 1 / 0 catch error:badarith -> {error, division_by_zero} end.
{error,division_by_zero}
2> try throw(custom_signal) catch throw:Reason -> {caught, Reason} end.
{caught,custom_signal}
Erlang has three distinct ways to signal something exceptional — error (a genuine bug — division by zero, a failed pattern match, a bad argument), throw (a value used for non-local control flow, the sender expects it to be caught somewhere), and exit (a deliberate request that a process should stop, which is also what an unhandled error becomes) — and try ... catch Class:Reason -> ... can distinguish between them by matching on Class.
The honest advice, and the whole thesis of "let it crash": reach for try/ catch far less often than instinct suggests. Wrapping every fallible operation defensively is the instinct this language actively argues against — the idiomatic response to "this function can fail in a way I have not planned for" is usually to let the process crash and have a supervisor restart it into clean state, not to catch the failure and attempt to continue in a state you were not prepared for. try/catch earns its place at a genuine boundary: the edge of the system (a network request, user input), or a place where "fail and report a specific, recoverable reason" is itself the correct behaviour rather than "fail and restart."
_:_ hides real bugs as reliably as it hides the failure you meant to catch A wildcard pattern matches every class and every reason, which means it catches the specific, anticipated failure you were thinking about exactly as well as it catches a genuine programming mistake you were not — a typo'd variable, a bad arithmetic operation, anything:
1> try X = 1 + not_a_number, X catch _:_ -> ok end.
ok
1 + not_a_number is a real bug — adding an integer to an atom — and the wildcard catch above converts it into the same ok a genuinely expected, handled failure would produce, with nothing in the return value to tell them apart. Catching a specific Class:Reason pattern (as the two examples at the top of this section do) lets anything that does not match propagate and crash loudly, which — per this section's own advice — is usually the outcome you want for the case you did not anticipate.
Write an expression that divides two numbers inside a try, catching only error:badarith specifically (not a wildcard), and returns {error, division_by_zero} on failure. Then call it with a division that raises a different error class (throw(not_a_number), say) and confirm your narrow catch does not swallow it.
safe_divide(A, B) ->
try A / B
catch error:badarith -> {error, division_by_zero}
end.
1> safe_divide(10, 0).
{error,division_by_zero}
2> try throw(not_a_number) catch error:badarith -> {error, division_by_zero} end.
** exception throw: not_a_number
The second call is the point: a throw is a different exception class from error, so a catch pattern narrowed to error:badarith correctly lets it through uncaught, rather than a wildcard silently absorbing an exception the function was never actually written to handle.
You will not write a full gen_server until Milestone 5, but it is worth seeing the shape of an OTP behaviour once now, so Milestone 5 is recognising a pattern rather than meeting one cold. A behaviour is a module that implements a fixed set of callbacks; OTP's generic machinery — code you never see or modify — handles the process loop, the message protocol, and a long list of edge cases (what happens if a reply never comes, how to shut down cleanly) uniformly for every module that implements it.
%% the shape you will fill in for real in Milestone 5 — not runnable yet
-module(mesh_node).
-behaviour(gen_server).
init(Args) -> {ok, InitialState}.
handle_call(Request, From, State) -> {reply, Reply, NewState}.
handle_cast(Request, State) -> {noreply, NewState}.
Compare the raw receive loop from Section 2.7 to this: the loop itself, the reply-matching convention, the timeout handling — all of it is gone, replaced by three small functions that only describe what should happen, never how the message actually gets there. That is the entire value proposition of an OTP behaviour, and Milestone 5 is the moment you stop hand-writing the "how" and start trusting a library that has handled it correctly for four decades.
-module(mesh_node_tests).
-include_lib("eunit/include/eunit.hrl").
%% a fixture: setup, the tests that need it, then teardown — for when a
%% test needs a running process rather than a pure function
spawn_and_stop_test_() ->
{setup,
fun() -> {ok, Pid} = mesh_node:start_link(#{id => 1}), Pid end,
fun(Pid) -> mesh_node:stop(Pid) end,
fun(Pid) ->
[?_assert(is_process_alive(Pid)),
?_assertEqual(alive, mesh_node:status(Pid))]
end}.
The {setup, Setup, Cleanup, Tests} tuple is EUnit's fixture shape, roughly parallel to Go's table-driven tests with a shared TestMain, or RSpec's before/after — and it is the shape nearly every test in Milestones 3 onward will use, because nearly every test needs a real, running process to test against.
receive loop version of the same server?
Milestone 1 turns the shell experiments above into a real, compiled, tested rebar3 project. Milestone 2 builds the pure data-handling functions the rest of the node will need. Milestone 3 is where a mesh node becomes a real, running process for the first time — everything from Section 2.7 onward, applied.