Milestones 9–12

Instalment 20 · Course 4 (Erlang) · Advanced phase and finish

What an experienced Erlang programmer wires together next, and whether you can now explain any of it

Five advanced topics with working code, one substantial final challenge about failover across real distributed nodes with its solution withheld, a knowledge check of forty-one questions, and everything you need to put mesh on GitHub and defend it in an interview.

Part AThe advanced phase

The Mewlang cat, wearing glasses, looking confidentThe mesh works: it supervises thousands of nodes, survives chaos injected on purpose, reports itself over HTTP, and runs across real, separate machines. These five topics are what you would reach for if this were a system you had to actually operate, not a system you had to finish.

A1 · Wiring the real supervision tree

Every supervisor so far has been started and exercised individually, from the shell. A real OTP application starts its whole tree from one place, deterministically, in dependency order.

%% src/mesh_sup.erl
-module(mesh_sup).
-behaviour(supervisor).
-export([start_link/0, init/1]).

start_link() ->
    supervisor:start_link({local, ?MODULE}, ?MODULE, []).

init([]) ->
    SupFlags = #{strategy => one_for_one, intensity => 5, period => 10},
    Children = [
        #{id => mesh_metrics,   start => {mesh_metrics, start_link, []}},
        #{id => node_sup,       start => {node_sup, start_link, []}, type => supervisor},
        #{id => mesh_dashboard, start => {mesh_dashboard, start_link, [8080]}}
    ],
    {ok, {SupFlags, Children}}.

Order is not cosmetic. Children start left to right and stop right to left, so mesh_metrics — which everything else calls into — is guaranteed running before node_sup starts spawning nodes that might immediately increment a counter, and mesh_dashboard, which reads metrics, starts last, after there is something meaningful to read. Getting this order wrong is a real, common bug: a child whose init/1 happens to call another child that has not started yet fails on startup, non-deterministically, depending on how fast each child happens to initialise — which is exactly the kind of bug that passes in development and fails once in a hundred production restarts.

What to measure: total startup time for the whole tree at your target node count, and whether it is dominated by node_sup spawning thousands of children or by something else — Milestone 7's 13ms-for-2,000-nodes number says it should not be the nodes themselves.

A2 · Process groups instead of scanning the whole registry

mesh_registry:broadcast/1 from Milestone 7's exercise walks every entry in one ETS table. That is fine at a few thousand nodes and starts to be the wrong tool past that, because every broadcast is O(n) work on the broadcasting process specifically. OTP's pg module (process groups, distribution-aware since OTP 23) exists for exactly "a dynamic set of processes I want to message as a group":

pg:start_link(),
pg:join(mesh_nodes, self()),           %% called from inside mesh_node:init/1
[Pid ! Msg || Pid <- pg:get_members(mesh_nodes)].

The advantage over the hand-rolled ETS version is not raw speed at this scale — measured, they are close — it is that pg groups are distribution-aware for free: a process group spans every connected distributed node automatically, so pg:get_members/1 called from nodeb in Milestone 11's two-node setup returns members living on nodea too, with no extra code. The hand-rolled ETS registry from Milestone 7 is local to whichever node it is running on; making it distributed would mean building the exact replication pg already has.

What to measure: broadcast latency at 2,000 and 20,000 members, both ways, and whether pg's distribution awareness changes anything about your own registry design once you have it.

A3 · Structured logs with logger

OTP's built-in logger (standard since OTP 21) gives structured, leveled logging with no dependency, the same shape Go's log/slog gave Course 1's colony:

logger:info("node started", #{id => Id, pid => self()}),
logger:warning("restart intensity exceeded", #{supervisor => node_sup, window_s => 5}),
2026-09-22T13:44:10.552Z info: node started id=42 pid=<0.94.0>

Two rules, matching the ones this curriculum has already argued for twice, in two other languages: never log per message — a mesh of two thousand nodes routing thousands of messages a second will produce far more log volume than the metrics counters ever will, for far less signal — and log the exceptional, count the routine, which is precisely why mesh_metrics:inc/1 exists as a separate mechanism from logging at all.

A4 · Deterministic chaos, properly replayable

Milestone 8's rand:seed/2 makes an attack pattern reproducible if nothing else about the run varies — Exercise 8's third part already found the gap: the input (which ids exist) has to be deterministic too. The complete fix records the schedule itself rather than trusting reseeded randomness to reproduce it exactly across Erlang/OTP versions, whose rand algorithm choices are not guaranteed stable forever:

build_schedule(Seed, NumAttacks, Ids) ->
    rand:seed(exsplus, {Seed, Seed, Seed}),
    [{lists:nth(rand:uniform(length(Ids)), Ids), rand:uniform()}
     || _ <- lists:seq(1, NumAttacks)].

A schedule built once and written to disk (file:write_file/2 with term_to_binary/1, or plain text via ~p) replays identically regardless of what rand does internally in a future OTP release — the same lesson Go's Course 1 learned building its own Schedule struct, arrived at for the same reason: a schedule is data, and data outlives the specific random-number generator that produced it.

A5 · Shipping it

FROM erlang:25-slim AS build
WORKDIR /src
COPY . .
RUN rebar3 release

FROM debian:bookworm-slim
COPY --from=build /src/_build/default/rel/mesh /opt/mesh
EXPOSE 8080
USER nobody
ENTRYPOINT ["/opt/mesh/bin/mesh", "foreground"]

A release (Milestone 12) is already the deployment artefact — this Dockerfile does almost nothing beyond copying it into a smaller base image. foreground rather than daemon for the entrypoint matters specifically in a container: a container's process manager expects the entrypoint to be the running process (PID 1), not a script that forks a daemon and exits immediately, which from the container runtime's point of view looks like the container finishing its work and stopping.

The BEAM itself can die, and something outside it has to notice

Every supervisor in this course watches processes inside one BEAM instance. Nothing built so far watches the BEAM instance itself — if the VM process is killed outright (kill -9, an out-of-memory kill, a host reboot), every supervision tree inside it dies with it, and by definition nothing inside that VM can be the thing that notices and restarts it. Erlang's own answer is heart, a small, separate OS-level watchdog process, started automatically when a release is booted with -heart, whose only job is monitoring the main BEAM process from outside it and restarting the whole thing if it stops responding. In a container, the more common answer is to let the orchestrator (Kubernetes, systemd, Docker's own restart policy) play that role instead — but the principle is the same one Course 1's Go colony met with Restart=always in a systemd unit: a supervision tree can only supervise what is inside its own process; something else has to watch the process itself.


Part BThe final challenge

Everything up to here had a solution a few paragraphs later. This one does not, and it is deliberately at the edge of what you can now do.

Failover across real distributed nodes, with no double ownership

Milestone 11 connected two BEAM instances and sent one message between them. Now make them cooperate for real: split the mesh's id space across two or more distributed Erlang nodes, and when one goes down, have another take over serving its ids — without ever having two live nodes simultaneously answering for the same id.

Requirements

  1. Run 3 distributed nodes (start with -sname on one machine, same as Milestone 11). Statically assign each a contiguous band of ids at startup (node 1: ids 1–999, node 2: 1000–1999, node 3: 2000–2999).
  2. Each node runs its own mesh_sup tree, owning only its own band.
  3. A client process (anywhere) can ask any connected node "what is the status of id N", and get a correct answer regardless of which node actually owns that id — the client should never need to know the sharding scheme.
  4. Kill a node outright (erlang:halt() called on it, or killing its OS process). Within a bounded time, exactly one surviving node must take over serving status queries for the dead node's band — not necessarily its exact process state (that was lost, honestly, with the dead node), but at minimum a correct "this id exists in a dead band, here is what we last knew" answer.
  5. Restart the dead node. It must not resume owning its old band while another node is actively covering for it — reconcile explicitly, on a schedule you define and document, not by accident of timing.

Constraints

Acceptance criteria

  1. No double ownership, ever, under a script that kills and restarts nodes randomly for 60 seconds. Log every ownership claim with a timestamp and assert, after the run, that no two claims for the same band overlap in time by more than your documented transition bound.
  2. Every status query, from any node, for any id, gets a response — including during a failover window, where "band N is currently being taken over, try again shortly" is an acceptable honest answer, but silence or a crash is not.
  3. A killed node's band is covered within a documented, bounded time after the kill is detected — state the number and justify it against your detection mechanism's own latency.
  4. rebar3 do eunit, ct passes, including at least one Common Test suite that actually starts multiple distributed nodes (ct_slave or peer, OTP's own facilities for spawning real additional nodes from inside a test).
  5. A documented failure analysis, in the same shape as Go Course 1's final challenge: every point something can go wrong, and why the design is safe (or explicitly not) at that point.

Hints, in increasing order of how much they give away

Attempt it before reading on. Even a partial implementation with an honest account of what does not fully work is worth more than the section below.

Solution — only look after trying

The design: global names as ownership claims

Each band gets a well-known global name, {band_owner, N}. A node claims a band by registering that name for a local process — band_owner_proc — using global:register_name/2. global is exactly the primitive this needs: the name is unique across every connected node without any node acting as a single coordinator, and — critically — if the process holding the name dies, the name is released automatically, which is what turns "node 2 crashed" into "band 2's ownership becomes claimable again" with no extra code.

%% src/band_owner.erl
-module(band_owner).
-behaviour(gen_server).
-export([start_link/1, claim/1, whois/1]).
-export([init/1, handle_call/3, handle_info/2]).

start_link(Band) -> gen_server:start_link(?MODULE, Band, []).

claim(Band) ->
    Name = {band_owner, Band},
    case global:register_name(Name, self()) of
        yes -> owned;
        no  -> {owned_by, global:whereis_name(Name)}
    end.

whois(Band) ->
    case global:whereis_name({band_owner, Band}) of
        undefined -> unowned;
        Pid -> {ok, Pid}
    end.

init(Band) ->
    net_kernel:monitor_nodes(true),
    self() ! {try_claim, Band},
    {ok, #{band => Band, owned => false}}.

handle_info({try_claim, Band}, State) ->
    case claim(Band) of
        owned ->
            logger:info("band claimed", #{band => Band, node => node()}),
            {noreply, State#{owned => true}};
        {owned_by, _Other} ->
            erlang:send_after(2000, self(), {try_claim, Band}),
            {noreply, State}
    end;
handle_info({nodedown, _Node}, #{band := Band} = State) ->
    %% a node vanished; whether it was THIS band's owner is unknown from
    %% here directly, so re-attempt the claim -- if this node already
    %% owns it, global:register_name/2 on the same name is a safe no-op
    %% for the current holder
    self() ! {try_claim, Band},
    {noreply, State}.

Every node runs one band_owner process per band that is not necessarily its own — in practice, one per band total across the whole cluster, all attempting to claim every band, with global's uniqueness guarantee meaning only one attempt per band actually succeeds. This directly answers the "no single point of failure for the decision" constraint: every node is independently trying to claim every band, all the time; nothing about the mechanism depends on any one specific node being alive to arbitrate.

The slow-versus-dead ambiguity, named rather than solved

{nodedown, Node} fires on a clean TCP disconnect or a missed heartbeat past net_kernel's own timeout — it cannot, and does not claim to, distinguish "that node's BEAM process is genuinely gone" from "that node is still running but is unreachable right now." This is the FLP-impossibility fact the Go course's final challenge met from a different angle: no failure detector in an asynchronous network can be both accurate and complete. The design above resolves the ambiguity the same way global itself does: by trusting the connection state, accepting that a node which is merely slow to respond but not actually disconnected will not trigger nodedown at all (correct — it is not down), and accepting that a node which reconnects after a network hiccup, without ever actually crashing, will find its global registration already gone and will simply re-claim its own band on the next {try_claim, ...} retry, which is safe precisely because re-claiming an already-owned-by-you name is a no-op.

Reconciliation on restart

A restarted node's band_owner for its own band runs init/1 exactly like every other node's copy: attempt to claim, and if another node already holds it (because it covered during the outage), back off and retry every 2 seconds. There is no special "I used to own this" logic at all — the restarted node is, deliberately, treated identically to any other node that wants a band it does not currently hold. This is the direct answer to requirement 5: reconciliation is not a separate mechanism bolted on for the restart case, it is the same claim loop every node always runs, which is what makes it safe rather than merely convenient.

What is still wrong with this, and you should say so in your README

  • Losing the band means losing its state. The surviving node that takes over a band did not inherit the dead node's mesh nodes' actual in-memory state — requirement 4 explicitly allowed "here is what we last knew" rather than perfect continuity, and this design takes that allowance fully: there is no state replication here at all, only ownership replication. A production version would need the owning node to periodically checkpoint each mesh node's state somewhere a successor could read it, which is a materially larger project.
  • global itself does not scale past a few hundred nodes gracefully — it is a full-mesh, all-nodes-agree protocol, fine for this challenge's three nodes and a genuine limitation the documentation for global states plainly for larger clusters.
  • The 2-second retry interval is a real, tunable trade between failover latency (how long a band goes uncovered) and claim-storm cost (how much distribution traffic every node generates trying to claim bands it does not hold) — the number was chosen, not derived, and that should be stated rather than implied to be correct.
  • No protection against a genuinely malicious or badly time-skewed node claiming a band it should not — global trusts every connected node equally, which is consistent with this course's whole cookie-based trust model (Milestone 11) and worth naming as a boundary rather than leaving implicit.

If you can explain why "no single point of failure for the decision" specifically ruled out a designated coordinator, and why that is the same FLP-flavoured impossibility the Go course's own final challenge met, you are ahead of most candidates with "distributed systems" on their CV.


Part CKnowledge check

C1 · Twenty conceptual questions

  1. What does = actually do in Erlang, precisely, and why is calling it "assignment" wrong?
  2. Why is a tail call compiled to a jump rather than a new stack frame, and what class of program does that guarantee make practical that would not be otherwise?
  3. Explain the difference between a link and a monitor in terms of direction and default effect, and name a place in this project each was the correct choice.
  4. Why does trap_exit change what a link does, rather than simply ignoring exit signals altogether?
  5. What is the actual difference between error, throw, and exit as ways to signal something exceptional?
  6. Why does this course's philosophy argue for reaching for try/catch less often than instinct suggests, and where does it say the boundary still belongs?
  7. What does an OTP behaviour actually check at compile time, and how does that differ from how a Go interface is satisfied?
  8. Explain simple_one_for_one versus plain one_for_one, and which shape of child population each fits.
  9. What do a supervisor's intensity and period actually bound together, and what happens once that bound is exceeded?
  10. Why is a process's mailbox described as unbounded, and what real cost follows from that when a caster outpaces its target?
  11. What is selective receive, and why does the Ref-tagging convention specifically enable it for request/reply?
  12. Why is ETS the right tool for a registry that many processes read and write concurrently, compared to funnelling every lookup through one process?
  13. What does a network partition actually do to the processes on either side of it, and what does it not do?
  14. Why can Erlang's distribution layer route a message to a remote pid with no protocol code written by the application, where Go's Milestone 12 needed an explicit wire protocol?
  15. What does the shared cookie between two distributed nodes actually authorise, and what does it not protect against?
  16. Why is tracing every call to a hot function on a live system dangerous in a way that is specific to what tracing costs, not just "debugging in production is risky" in the abstract?
  17. What is the difference between what EUnit, Common Test, and PropEr are each best suited to test?
  18. Why does a property-based test search for a counterexample rather than checking specific examples, and what kind of bug is it more likely to find than a hand-written test suite?
  19. Explain why heart (or an external process manager) is necessary even though every supervisor tree in this project is, itself, a fault-tolerance mechanism.
  20. State, in your own words, what this course's central comparison to the Perl course's "a line you cannot parse is still evidence" rule actually is, and why both are correct for what each system is protecting.
Answers to C1
  1. Pattern matching: it binds an unbound variable to a value, or asserts (raising an exception if false) that an already-bound variable equals the value on the right. "Assignment" implies reassignment is possible, which it is not — a bound variable stays bound to that value for the rest of its scope.
  2. A tail call reuses the current stack frame instead of pushing a new one, because nothing remains to be done with its result except return it. This makes unbounded recursion — the only looping construct Erlang has — run in constant stack space, which is what makes "loop by recursion" a practical default rather than a recipe for stack overflow on any sufficiently long-running process.
  3. A link is bidirectional and, by default, propagates a crash to both sides; a monitor is one-directional and only ever notifies, never kills. Links: every supervisor-to-child relationship in this project. Monitors: the registry's cleanup-on-death watcher in Milestone 7, which observes nodes it does not own or supervise.
  4. trap_exit converts what would be a fatal propagated exit signal into an ordinary {'EXIT', Pid, Reason} message delivered to the trapping process's own mailbox — the link (and the guarantee that a linked process's death is always reported) is preserved; only the default "and then you die too" behaviour is replaced with "and then you get to decide."
  5. error signals a genuine bug or invalid operation; throw signals a value meant for non-local control flow that the thrower expects someone to catch; exit is a deliberate request that a process terminate, which is also what an unhandled error becomes once it propagates past the top of a process.
  6. Because catching a failure defensively and attempting to continue means continuing in a state nobody designed for or tested; letting the process crash and restart into the known-good state init/1 establishes is usually safer. The boundary is genuine edges — user input, a network response, anywhere "fail with a specific, recoverable reason" is itself the correct behaviour rather than a defensive reflex.
  7. It checks, at compile time, that every callback the behaviour requires is actually exported by the module declaring it, and warns if one is missing. A Go interface is satisfied implicitly, by having the right method set, checked only where it is used; an Erlang behaviour is declared explicitly and checked against its own declaration regardless of whether anything yet calls it.
  8. simple_one_for_one is a template for an unbounded, dynamically-sized population of identical children, added and removed at runtime — this project's node pool. one_for_one is a fixed, statically-declared list of distinct children known at startup — mesh_sup's own children (metrics, node_sup, dashboard).
  9. The number of restarts tolerated within a rolling time window. Exceeding it means the supervisor concludes restarting is not fixing anything, stops trying, and terminates itself, reporting the failure one level further up the tree.
  10. Because nothing bounds how many messages can queue in it — a producer that outpaces its consumer grows the mailbox without limit, and every receive against that mailbox gets slower too, because a selective receive has to scan past everything already queued that does not match.
  11. Scanning a mailbox for the first message matching one of a receive's clauses, leaving non-matching messages queued for a later receive. The Ref, unique per request, lets a process wait specifically for the reply to this call while other, unrelated messages sit safely unmatched in the same mailbox.
  12. ETS lets any process read or write the table directly and concurrently, with atomic per-key operations like update_counter/3 — funnelling every access through one owning process turns that process into a serialisation point and a single point of failure for every reader, which is the exact problem ETS avoids.
  13. It stops the processes on either side from being able to see or message each other; it does not kill, corrupt, or pause any process on either side — each half keeps running, internally consistent, believing itself to be the whole system.
  14. Because a pid is a location-transparent reference the runtime itself knows how to route to, locally or across a connection, with the distribution protocol built into the BEAM. Go's channels are a purely local, in-process primitive with no distributed counterpart, so "distribute it" meant designing and implementing a wire protocol by hand.
  15. That two nodes trust each other enough to connect and exchange messages at all. It does not encrypt traffic between them, and does not by itself protect against a node on the same network that has obtained the cookie — it is a shared secret, not a full authentication and encryption scheme.
  16. Tracing generates one trace message per matched call, delivered to the tracer process — tracing every call to a function invoked thousands of times a second can produce trace volume the tracer cannot keep up with, which is a new, self-inflicted load and memory problem layered on top of whatever you were originally trying to diagnose.
  17. EUnit: fast, function-level unit tests, including ones needing a running process via its fixture shape. Common Test: heavier, suite-level tests with proper setup/teardown, suited to integration-shaped scenarios like starting a real supervision tree. PropEr: properties that should hold across a whole class of generated inputs, not specific examples.
  18. Because it generates many random inputs and searches for one that breaks the stated property, rather than checking only the inputs a person thought to write down — it is more likely to find an edge case (a boundary value, an unusual combination) that never occurred to the test's author.
  19. A supervisor tree can only supervise processes running inside its own BEAM instance; if the BEAM process itself is killed or the host crashes, nothing inside that VM can be the thing that notices, because there is no "inside" left running. Something outside the VM — heart, a container orchestrator, systemd — has to watch the VM process itself.
  20. Strata protects data: a malformed record is still evidence, so every parser must always return something rather than lose input. A mesh node protects overall system availability: a node that crashed and restarted cleanly is cheap and expected, while a node that survived in a corrupted, half-understood state to keep processing is the actual danger. Both are legitimate fault tolerance strategies; the difference is what each system considers the unacceptable outcome.

C2 · Ten code-reading questions

Predict the output of each, then check. All ten were run to confirm the answers.

%% 1
X = 1 == 1.0,
Y = 1 =:= 1.0,
io:format("~p ~p~n", [X, Y]).

%% 2
F = fun() ->
    L = [N * 2 || N <- [1,2,3,4], N rem 2 =/= 0],
    L
end,
io:format("~p~n", [F()]).

%% 3
io:format("~p~n", [length([1,2,3|[4,5]])]).

%% 4
A = try error(oops) catch _:_ -> caught end,
io:format("~p~n", [A]).

%% 5
R = (catch 1/0),
io:format("~p~n", [element(1, R)]).

%% 6
M0 = #{a => 1},
M1 = M0#{a => 2},
io:format("~p ~p~n", [M0, M1]).

%% 7
loop(0) -> done;
loop(N) -> loop(N - 1).
io:format("~p~n", [loop(3)]).

%% 8
Pid = spawn(fun() -> receive _ -> ok end end),
Alive1 = is_process_alive(Pid),
exit(Pid, kill),
timer:sleep(10),
Alive2 = is_process_alive(Pid),
io:format("~p ~p~n", [Alive1, Alive2]).

%% 9
{ok, S} = {ok, hello},
Result = case S of
    hello -> matched_atom;
    _ -> something_else
end,
io:format("~p~n", [Result]).

%% 10
G = fun(N) when N > 0, N < 10 -> small;
        (N) when N >= 10 -> big;
        (_) -> negative_or_zero
    end,
io:format("~p ~p ~p~n", [G(5), G(50), G(-1)]).
Answers to C2
1.  true false
2.  [2,6]
3.  5
4.  caught
5.  'EXIT'
6.  #{a => 1} #{a => 2}
7.  done
8.  true false
9.  matched_atom
10. small big negative_or_zero
  1. == compares by value across types (1 and 1.0 are numerically equal); =:= also requires the same type, and an integer is never the same type as a float.
  2. The filter keeps odd numbers from the input (1 and 3), and the comprehension doubles each, giving [2, 6] — note the output order matches the input order, not the doubled values' magnitude.
  3. [1,2,3|[4,5]] is exactly the same list as [1,2,3,4,5] — the | syntax conses onto the front of whatever list follows it, and a list is a valid thing to cons onto just as a single element is.
  4. try ... catch Class:Reason -> ... with a wildcard pattern catches any exception class, converting the crash into the ordinary value caught.
  5. A bare catch Expr (not try) converts an exception into a value shaped {'EXIT', Reason} rather than propagating it; element(1, R) extracts the tag, which is the atom 'EXIT'.
  6. Map update syntax (M0#{a := 2} or, as here, => which both inserts and updates) produces a new map; M0, already bound, is completely unaffected by anything done to build M1.
  7. Tail-recursive countdown to the base case, returning the atom done — included as a reminder that a "loop" in Erlang is just an ordinary function, with an ordinary return value, nothing special about it syntactically.
  8. The process is genuinely alive before exit/2 (it is blocked in receive, which is a normal, alive state, not a dead one) and genuinely dead 10ms after an exit(Pid, kill) — kill specifically is a reason that cannot be trapped even by a process with trap_exit set, unlike an ordinary exit reason.
  9. Matching {ok, S} against the literal tuple {ok, hello} binds S to the atom hello; the case then matches the first clause exactly.
  10. Three function clauses in one anonymous fun, dispatched by guard exactly like named-function clauses — the syntax (semicolons between clauses, all sharing one fun ... end) is the only thing that looks unfamiliar; the dispatch mechanism is identical to describe/1 from the instalment's Section 2.5.

C3 · Five debugging exercises

Each gives a symptom and a suspect. Diagnose before opening the answer.

  1. Symptom: a node's mailbox grows without bound over a long chaos run, and the whole process gets visibly slower over time, though it never crashes.
    handle_call(get_status, _From, State) ->
        #{status := S} = State,
        {reply, S, State}.
    %% no handle_cast/2 clause defined at all
  2. Symptom: two chaos runs with the same seed crash different sets of node ids.
    Ids = [Id || {Id, _} <- ets:tab2list(mesh_registry)],  % order not guaranteed
    rand:seed(exsplus, {Seed, Seed, Seed}),
    [attack(Id) || Id <- Ids].
  3. Symptom: node_sup occasionally fails to start with {error, {already_started, Pid}}, only under a fast test script that starts and stops it repeatedly.
    stop_sup() ->
        exit(whereis(node_sup), kill),
        ok.   % returns immediately; caller assumes the supervisor is gone
  4. Symptom: the dashboard's /metrics endpoint occasionally hangs indefinitely for one specific client, never timing out, blocking that one connection forever.
    handle(Sock) ->
        {ok, Req} = gen_tcp:recv(Sock, 0),   % no timeout argument
        ...
  5. Symptom: a supervisor configured with intensity => 3, period => 5 shuts down after only one crash during a specific test, not three.
    ChildSpec = #{id => mesh_node, start => {mesh_node, start_link, []},
                  restart => permanent, shutdown => 100},
    %% test calls mesh_node:stop/1 (a deliberate, clean stop) three times,
    %% then crashes it once, expecting 3 tolerated restarts still available
Answers to C3
  1. No catch-all in handle_cast/2 (or missing entirely). A gen_server with no matching handle_cast/2 clause for an incoming cast crashes on that cast — which would at least be visible — but a module that defines no handle_cast/2 at all still compiles (the behaviour only warns, does not require every callback if defaults exist) and silently accumulates unhandled casts exactly like a receive with no matching clause, per the instalment's mailbox warning. Fix: add an explicit catch-all clause that at least logs and discards.
  2. ets:tab2list/1 has no guaranteed order. The chaos schedule seeds rand deterministically, but iterates the ids in whatever order ETS happens to return them, which is not specified to be stable across runs or even across the same table's internal rehashing. Fix: sort the ids explicitly before building the schedule, so the sequence of "which id gets attacked Nth" is deterministic, not merely the random numbers are.
  3. Killing a process does not mean it has finished terminating yet. exit(Pid, kill) sends an asynchronous signal and returns immediately — the target process is scheduled to die but has not necessarily been fully cleaned up (its registered name released, in particular) by the time the caller proceeds to start a new one under the same name. Fix: monitor the process and wait for its 'DOWN' message before considering it gone, exactly the pattern mesh_watcher already used for the opposite direction.
  4. gen_tcp:recv/2 with no timeout blocks forever if the client opens a connection and then never sends a complete request — a slow-loris-shaped client, deliberate or not. Fix: gen_tcp:recv(Sock, 0, 5000), and handle the {error, timeout} case by closing the socket, exactly the discipline every gen_server:call/2 already has built in by default.
  5. restart => permanent instead of transient. A permanent child is restarted on any termination, including the clean, deliberate stops the test performed — each of those three intentional stops counted against the intensity budget exactly like a real crash would, leaving none left for the genuine crash that followed. This is precisely why Milestone 6 chose transient for mesh nodes: a deliberately-stopped node should not consume restart budget meant for genuine failures.

C4 · Five implementation exercises

  1. A bounded mailbox pattern. gen_server mailboxes are unbounded by design; implement admission control in front of one — a wrapper that checks process_info(Pid, message_queue_len) before casting, and returns {error, overloaded} rather than casting past a configurable threshold. Measure whether this actually protects a deliberately slow-handler node from an unbounded mailbox under chaos.
  2. Rolling restart. Write a function that restarts every node in node_sup one at a time, waiting for each replacement to report alive before moving to the next, with a configurable delay between each — the operational tool you would actually want before deploying a code change to a live mesh.
  3. A gen_statem version of mesh_node. OTP's gen_statem behaviour models a process as an explicit state machine rather than a bag of callbacks over one opaque state term. Rebuild mesh_node with alive and crashed as explicit states, and compare: what does making the states explicit catch that the map-based status field could not?
  4. Cluster-wide metrics. Using pg or global from Part A, aggregate mesh_metrics:snapshot/0 across every connected distributed node into one combined view, callable from any single node.
  5. A minimal relup. Using relx's upgrade support, ship a trivial code change (a new log line in mesh_node) as a hot upgrade to a running release — no restart, the running system picks up the new code while its supervision tree and all live processes keep running. This is genuinely fiddly to get right; treat getting it to work at all as the win.

C5 · One substantial challenge

Distinct from the final challenge in Part B, and smaller, but not easy.

Build a mailbox-depth-aware load shedder. Give every node's gen_server a way to report its own mailbox depth on every message handled (process_info(self(), message_queue_len), cheap enough to call routinely), publish it to mesh_metrics, and have mesh_registry:route/2 refuse to route a new message to a node whose reported depth is above a threshold, returning {error, node_overloaded} instead of adding to an already-backed-up mailbox.

Requirements: the threshold must be per-node, not global, since Milestone 8's slow-handler attack targets individual nodes, not the whole mesh; the mechanism must not itself become a bottleneck at 2,000 nodes (measure the added cost of a depth check on every routed message); and you must demonstrate, under chaos, that a genuinely overloaded node's mailbox depth stays bounded rather than growing without limit the way an unprotected one does. Hint: reporting depth on every handled message is itself extra work on the hot path — consider reporting it periodically instead, and what staleness that trades away.

C6 · You should now be able to explain

C7 · You should now be able to implement


Part DShipping it: README, portfolio, interview

D1 · README draft

# mesh

A simulated network of thousands of supervised Erlang processes that crash,
restart, partition and recover — a laboratory for OTP supervision, and,
in its final form, a genuinely distributed, multi-node fault-tolerant system.

No third-party dependencies beyond PropEr (dev/test only). Standard OTP.

## What it does

- Each mesh node is a gen_server, supervised by a simple_one_for_one tree
  that tolerates a bounded number of crashes before giving up and
  escalating — measured: exactly 3 restarts tolerated, the 4th within the
  same 5-second window brings the supervisor down, on command.
- An ETS-backed registry finds any node by id, with automatic,
  monitor-driven cleanup on crash — no stale entries survive a restart.
- A seeded chaos module attacks the live mesh: random crashes, slow
  handlers, malformed messages — fully reproducible from one integer.
- A dashboard, served over plain HTTP from inside the system it reports
  on, with a streaming (SSE) live view — no external web framework.
- Genuinely distributed: multiple real BEAM nodes, connected over a real
  network, sharing message-passing code that never needed to change to
  become distribution-aware.
- Ships as a standalone relx release: no Erlang installation required on
  the target machine.

## Quick start

    rebar3 release
    _build/default/rel/mesh/bin/mesh daemon
    curl http://localhost:8080/metrics
    _build/default/rel/mesh/bin/mesh stop

## Architecture

    mesh_sup (one_for_one)
      |- mesh_metrics        counters, ETS-backed
      |- node_sup (simple_one_for_one)
      |    `- mesh_node x N   gen_server, supervised, crashable on purpose
      `- mesh_dashboard       HTTP + SSE, reads mesh_metrics

## Testing

    rebar3 do eunit, ct, proper

## Known limitations

- The dashboard has no auth; it is meant for a trusted network only.
- Distribution trusts the shared cookie; there is no additional
  authentication or encryption between nodes.
- global-based failover (advanced phase) does not scale past a few
  hundred nodes gracefully — this is a stated limitation of `global`
  itself, not a bug in this project's use of it.

## Licence

MIT

Three deliberate choices, matching every README in this curriculum so far: it leads with measured numbers (the exact restart-intensity threshold, not an adjective); it states what it depends on and what it does not; and it has a known limitations section that names the real, specific boundary of global rather than implying the failover mechanism is production-ready as built.

D2 · GitHub project description

A simulated distributed network in Erlang/OTP: supervised processes that crash and recover on purpose, a live HTTP dashboard, and real multi-node failover — a laboratory for "let it crash" as an engineering strategy, not a slogan.

Topics: erlang, otp, gen-server, supervisor, fault-tolerance, distributed-systems, chaos-engineering, observability, property-based-testing.

D3 · Performance considerations

D4 · Security considerations

D5 · What to put in your portfolio

Do not present this as "a chat simulation of network failures." Present it as what it is: a study of failure as a first-class, deliberately-exercised code path, ending in a real, multi-node distributed system with a documented, honest failure analysis. The narrative that makes it interesting is the escalation from hand-rolled to OTP-standard, twice.

  1. A hand-rolled process and a hand-rolled watcher, each working, each with a named, specific gap.
  2. The same two problems, solved by gen_server and supervisor — same external behaviour, the gaps closed, measured: exactly 3 restarts tolerated, the 4th one not.
  3. A registry at real scale (2,000 nodes, 13ms), with the stale-entry bug found and fixed honestly.
  4. Chaos injection absorbed silently by the supervision tree — the mesh's population recovers to full strength with no manual intervention, verified by counting live children, not by assuming.
  5. Real distribution: two separate BEAM instances, a real message crossing a real network boundary, with no code written specifically to make that possible.
  6. The final challenge's honest failure analysis, including what still is not solved.

Keep a docs/ folder with the restart-intensity transcript, the 2,000-node timing, and one architecture diagram. A reviewer who spends ninety seconds on your repository should come away knowing you tested failure on purpose and measured what it actually cost, not just that the happy path works.

D6 · Interview questions someone could ask, and what a good answer contains

QuestionWhat a strong answer includes
Walk me through the supervision design.The tree shape, why simple_one_for_one fits a dynamic node pool where one_for_one fits the root, and the exact restart-intensity numbers, measured.
Why not just catch every exception defensively?The "let it crash" argument, stated precisely — a caught, patched-around failure continues in an untested state; a crash-and-restart returns to a known-good one. And the honest boundary: genuine edges still need explicit error handling.
Tell me about a bug you found.The registry's stale-entry leak from a missing cleanup path, or the process-dictionary id counter's hidden assumption of single-process use — how it was found, and the fix.
How does this differ from how you'd do fault tolerance in Go?Structural process isolation versus disciplined defer recover(); an unrecovered panic takes the whole Go program down where an Erlang crash takes down exactly one process. Both real trade-offs, neither universally superior.
What happens during a network partition?Both sides keep running, independently, each internally consistent; the danger is exclusively in reconciling state accumulated on both sides afterward, which this project's final challenge handles for pure ownership (via global) and explicitly does not solve for replicated state.
How would you make this production-ready?Authentication on the dashboard and the distribution port, TLS between distributed nodes, request-size limits on the hand-rolled HTTP server, state checkpointing for real failover continuity, and replacing global with something proven at larger scale if the cluster ever needs to grow past a few hundred nodes.
Exactly-once delivery: how, in this system?You cannot, structurally, the same as every other course in this curriculum that met the question — at-least-once plus an idempotent receiver is the honest answer, and this project does not currently implement it anywhere that would need it, which is worth saying plainly rather than implying it does.
When would you not use Erlang for this?Anywhere CPU-bound numeric throughput on one core is the actual bottleneck — the BEAM is built for concurrency and soft real-time responsiveness, not for winning a tight numeric loop against a language that compiles closer to the metal. Name what Erlang wins on here specifically: isolation, supervision, and distribution built into the runtime rather than bolted on.
What would you do differently?Build the deterministic chaos schedule (Part A4) from the start rather than discovering the seeded-but-not-fully-reproducible gap after the fact; design the metrics set before counters accumulated ad hoc; and decide the state-checkpointing story for failover before, not after, building the ownership-transfer mechanism.

D7 · Extensions worth building


Course 4 completeWhat you built

An OTP application spanning a dozen modules, a supervision tree tested to its actual restart-intensity limit rather than assumed to work, a registry proven at 2,000 nodes with a real bug found and fixed, a seeded and reproducible chaos-injection framework, a dashboard serving a live system's own health over HTTP from inside it, a real multi-node distributed cluster, and a documented, honest failover design with its limitations stated rather than hidden. More importantly: the instinct to let something crash on purpose and trust the supervisor, rather than reaching first for a defensive catch.

The central question of this curriculum was what kinds of problems does this language make unusually natural to solve? Erlang's answer, stated as precisely as this project allows: problems where the cost of one component failing must never become the cost of the whole system failing, where isolation is worth paying a copying and messaging overhead for, and where "this will be distributed eventually" is true often enough that the language should not treat distribution as a bolt-on afterthought. Not the fastest, not the simplest to learn in an afternoon, not the language you reach for when the whole problem is a tight numeric loop. The one where letting things break, on purpose, correctly, is how the system stays up.

Courses so far have paired concurrency (Go, Erlang) against expressiveness and language design (Ruby, Racket), with Perl's text-forensics work sitting adjacent to both. Course 5 closes the curriculum with Racket, and the same comparison this course drew against Go's channels and Perl's parsing philosophy gets drawn one more time, from the opposite direction: instead of a language whose runtime already gives you the primitive you need (processes, supervision), Racket is a language built around the idea that if it does not give you the primitive you need, you can build it — as a real language of your own, checked, hygienic, and yours.

Instalment 20 of the five-course curriculum, and the end of Course 4. Next: Course 5, Racket, and the Language Factory. Parts 0–2 first (what we are building, installation and raco, the language crash course), then twelve milestones ending in real, working #lang implementations.

Next: Racket instalment