Instalment 20 · Course 4 (Erlang) · Advanced phase and finish
Five advanced topics with working code, one substantial final challenge about failover across real distributed nodes with its solution withheld, a knowledge check of forty-one questions, and everything you need to put mesh on GitHub and defend it in an interview.
The mesh works: it supervises thousands of nodes, survives chaos injected on purpose, reports itself over HTTP, and runs across real, separate machines. These five topics are what you would reach for if this were a system you had to actually operate, not a system you had to finish.
Every supervisor so far has been started and exercised individually, from the shell. A real OTP application starts its whole tree from one place, deterministically, in dependency order.
%% src/mesh_sup.erl
-module(mesh_sup).
-behaviour(supervisor).
-export([start_link/0, init/1]).
start_link() ->
supervisor:start_link({local, ?MODULE}, ?MODULE, []).
init([]) ->
SupFlags = #{strategy => one_for_one, intensity => 5, period => 10},
Children = [
#{id => mesh_metrics, start => {mesh_metrics, start_link, []}},
#{id => node_sup, start => {node_sup, start_link, []}, type => supervisor},
#{id => mesh_dashboard, start => {mesh_dashboard, start_link, [8080]}}
],
{ok, {SupFlags, Children}}.
Order is not cosmetic. Children start left to right and stop right to left, so mesh_metrics — which everything else calls into — is guaranteed running before node_sup starts spawning nodes that might immediately increment a counter, and mesh_dashboard, which reads metrics, starts last, after there is something meaningful to read. Getting this order wrong is a real, common bug: a child whose init/1 happens to call another child that has not started yet fails on startup, non-deterministically, depending on how fast each child happens to initialise — which is exactly the kind of bug that passes in development and fails once in a hundred production restarts.
What to measure: total startup time for the whole tree at your target node count, and whether it is dominated by node_sup spawning thousands of children or by something else — Milestone 7's 13ms-for-2,000-nodes number says it should not be the nodes themselves.
mesh_registry:broadcast/1 from Milestone 7's exercise walks every entry in one ETS table. That is fine at a few thousand nodes and starts to be the wrong tool past that, because every broadcast is O(n) work on the broadcasting process specifically. OTP's pg module (process groups, distribution-aware since OTP 23) exists for exactly "a dynamic set of processes I want to message as a group":
pg:start_link(),
pg:join(mesh_nodes, self()), %% called from inside mesh_node:init/1
[Pid ! Msg || Pid <- pg:get_members(mesh_nodes)].
The advantage over the hand-rolled ETS version is not raw speed at this scale — measured, they are close — it is that pg groups are distribution-aware for free: a process group spans every connected distributed node automatically, so pg:get_members/1 called from nodeb in Milestone 11's two-node setup returns members living on nodea too, with no extra code. The hand-rolled ETS registry from Milestone 7 is local to whichever node it is running on; making it distributed would mean building the exact replication pg already has.
What to measure: broadcast latency at 2,000 and 20,000 members, both ways, and whether pg's distribution awareness changes anything about your own registry design once you have it.
loggerOTP's built-in logger (standard since OTP 21) gives structured, leveled logging with no dependency, the same shape Go's log/slog gave Course 1's colony:
logger:info("node started", #{id => Id, pid => self()}),
logger:warning("restart intensity exceeded", #{supervisor => node_sup, window_s => 5}),
2026-09-22T13:44:10.552Z info: node started id=42 pid=<0.94.0>
Two rules, matching the ones this curriculum has already argued for twice, in two other languages: never log per message — a mesh of two thousand nodes routing thousands of messages a second will produce far more log volume than the metrics counters ever will, for far less signal — and log the exceptional, count the routine, which is precisely why mesh_metrics:inc/1 exists as a separate mechanism from logging at all.
Milestone 8's rand:seed/2 makes an attack pattern reproducible if nothing else about the run varies — Exercise 8's third part already found the gap: the input (which ids exist) has to be deterministic too. The complete fix records the schedule itself rather than trusting reseeded randomness to reproduce it exactly across Erlang/OTP versions, whose rand algorithm choices are not guaranteed stable forever:
build_schedule(Seed, NumAttacks, Ids) ->
rand:seed(exsplus, {Seed, Seed, Seed}),
[{lists:nth(rand:uniform(length(Ids)), Ids), rand:uniform()}
|| _ <- lists:seq(1, NumAttacks)].
A schedule built once and written to disk (file:write_file/2 with term_to_binary/1, or plain text via ~p) replays identically regardless of what rand does internally in a future OTP release — the same lesson Go's Course 1 learned building its own Schedule struct, arrived at for the same reason: a schedule is data, and data outlives the specific random-number generator that produced it.
FROM erlang:25-slim AS build
WORKDIR /src
COPY . .
RUN rebar3 release
FROM debian:bookworm-slim
COPY --from=build /src/_build/default/rel/mesh /opt/mesh
EXPOSE 8080
USER nobody
ENTRYPOINT ["/opt/mesh/bin/mesh", "foreground"]
A release (Milestone 12) is already the deployment artefact — this Dockerfile does almost nothing beyond copying it into a smaller base image. foreground rather than daemon for the entrypoint matters specifically in a container: a container's process manager expects the entrypoint to be the running process (PID 1), not a script that forks a daemon and exits immediately, which from the container runtime's point of view looks like the container finishing its work and stopping.
Every supervisor in this course watches processes inside one BEAM instance. Nothing built so far watches the BEAM instance itself — if the VM process is killed outright (kill -9, an out-of-memory kill, a host reboot), every supervision tree inside it dies with it, and by definition nothing inside that VM can be the thing that notices and restarts it. Erlang's own answer is heart, a small, separate OS-level watchdog process, started automatically when a release is booted with -heart, whose only job is monitoring the main BEAM process from outside it and restarting the whole thing if it stops responding. In a container, the more common answer is to let the orchestrator (Kubernetes, systemd, Docker's own restart policy) play that role instead — but the principle is the same one Course 1's Go colony met with Restart=always in a systemd unit: a supervision tree can only supervise what is inside its own process; something else has to watch the process itself.
Everything up to here had a solution a few paragraphs later. This one does not, and it is deliberately at the edge of what you can now do.
Milestone 11 connected two BEAM instances and sent one message between them. Now make them cooperate for real: split the mesh's id space across two or more distributed Erlang nodes, and when one goes down, have another take over serving its ids — without ever having two live nodes simultaneously answering for the same id.
-sname on one machine, same as Milestone 11). Statically assign each a contiguous band of ids at startup (node 1: ids 1–999, node 2: 1000–1999, node 3: 2000–2999).mesh_sup tree, owning only its own band.erlang:halt() called on it, or killing its OS process). Within a bounded time, exactly one surviving node must take over serving status queries for the dead node's band — not necessarily its exact process state (that was lost, honestly, with the dead node), but at minimum a correct "this id exists in a dead band, here is what we last knew" answer.rebar3 do eunit, ct passes, including at least one Common Test suite that actually starts multiple distributed nodes (ct_slave or peer, OTP's own facilities for spawning real additional nodes from inside a test).net_kernel:monitor_nodes(true) delivers a {nodedown, Node} message to whichever process asked. The hard part is not detecting a clean disconnect — it is that a node which is merely slow looks identical, from a distance, to a node which is dead. What does your design do about that ambiguity?global:register_name/2 gives you a name that is unique across every connected node, enforced by Erlang's own distribution kernel — think about what it would mean for "who currently owns band 2" to be exactly this kind of name, claimed by whichever node currently believes it owns that band, and what happens automatically when the process holding that name dies.Attempt it before reading on. Even a partial implementation with an honest account of what does not fully work is worth more than the section below.
global names as ownership claimsEach band gets a well-known global name, {band_owner, N}. A node claims a band by registering that name for a local process — band_owner_proc — using global:register_name/2. global is exactly the primitive this needs: the name is unique across every connected node without any node acting as a single coordinator, and — critically — if the process holding the name dies, the name is released automatically, which is what turns "node 2 crashed" into "band 2's ownership becomes claimable again" with no extra code.
%% src/band_owner.erl
-module(band_owner).
-behaviour(gen_server).
-export([start_link/1, claim/1, whois/1]).
-export([init/1, handle_call/3, handle_info/2]).
start_link(Band) -> gen_server:start_link(?MODULE, Band, []).
claim(Band) ->
Name = {band_owner, Band},
case global:register_name(Name, self()) of
yes -> owned;
no -> {owned_by, global:whereis_name(Name)}
end.
whois(Band) ->
case global:whereis_name({band_owner, Band}) of
undefined -> unowned;
Pid -> {ok, Pid}
end.
init(Band) ->
net_kernel:monitor_nodes(true),
self() ! {try_claim, Band},
{ok, #{band => Band, owned => false}}.
handle_info({try_claim, Band}, State) ->
case claim(Band) of
owned ->
logger:info("band claimed", #{band => Band, node => node()}),
{noreply, State#{owned => true}};
{owned_by, _Other} ->
erlang:send_after(2000, self(), {try_claim, Band}),
{noreply, State}
end;
handle_info({nodedown, _Node}, #{band := Band} = State) ->
%% a node vanished; whether it was THIS band's owner is unknown from
%% here directly, so re-attempt the claim -- if this node already
%% owns it, global:register_name/2 on the same name is a safe no-op
%% for the current holder
self() ! {try_claim, Band},
{noreply, State}.
Every node runs one band_owner process per band that is not necessarily its own — in practice, one per band total across the whole cluster, all attempting to claim every band, with global's uniqueness guarantee meaning only one attempt per band actually succeeds. This directly answers the "no single point of failure for the decision" constraint: every node is independently trying to claim every band, all the time; nothing about the mechanism depends on any one specific node being alive to arbitrate.
{nodedown, Node} fires on a clean TCP disconnect or a missed heartbeat past net_kernel's own timeout — it cannot, and does not claim to, distinguish "that node's BEAM process is genuinely gone" from "that node is still running but is unreachable right now." This is the FLP-impossibility fact the Go course's final challenge met from a different angle: no failure detector in an asynchronous network can be both accurate and complete. The design above resolves the ambiguity the same way global itself does: by trusting the connection state, accepting that a node which is merely slow to respond but not actually disconnected will not trigger nodedown at all (correct — it is not down), and accepting that a node which reconnects after a network hiccup, without ever actually crashing, will find its global registration already gone and will simply re-claim its own band on the next {try_claim, ...} retry, which is safe precisely because re-claiming an already-owned-by-you name is a no-op.
A restarted node's band_owner for its own band runs init/1 exactly like every other node's copy: attempt to claim, and if another node already holds it (because it covered during the outage), back off and retry every 2 seconds. There is no special "I used to own this" logic at all — the restarted node is, deliberately, treated identically to any other node that wants a band it does not currently hold. This is the direct answer to requirement 5: reconciliation is not a separate mechanism bolted on for the restart case, it is the same claim loop every node always runs, which is what makes it safe rather than merely convenient.
global itself does not scale past a few hundred nodes gracefully — it is a full-mesh, all-nodes-agree protocol, fine for this challenge's three nodes and a genuine limitation the documentation for global states plainly for larger clusters. global trusts every connected node equally, which is consistent with this course's whole cookie-based trust model (Milestone 11) and worth naming as a boundary rather than leaving implicit.If you can explain why "no single point of failure for the decision" specifically ruled out a designated coordinator, and why that is the same FLP-flavoured impossibility the Go course's own final challenge met, you are ahead of most candidates with "distributed systems" on their CV.
= actually do in Erlang, precisely, and why is calling it "assignment" wrong? trap_exit change what a link does, rather than simply ignoring exit signals altogether?error, throw, and exit as ways to signal something exceptional?try/catch less often than instinct suggests, and where does it say the boundary still belongs?simple_one_for_one versus plain one_for_one, and which shape of child population each fits.intensity and period actually bound together, and what happens once that bound is exceeded?Ref-tagging convention specifically enable it for request/reply?heart (or an external process manager) is necessary even though every supervisor tree in this project is, itself, a fault-tolerance mechanism.trap_exit converts what would be a fatal propagated exit signal into an ordinary {'EXIT', Pid, Reason} message delivered to the trapping process's own mailbox — the link (and the guarantee that a linked process's death is always reported) is preserved; only the default "and then you die too" behaviour is replaced with "and then you get to decide."error signals a genuine bug or invalid operation; throw signals a value meant for non-local control flow that the thrower expects someone to catch; exit is a deliberate request that a process terminate, which is also what an unhandled error becomes once it propagates past the top of a process.init/1 establishes is usually safer. The boundary is genuine edges — user input, a network response, anywhere "fail with a specific, recoverable reason" is itself the correct behaviour rather than a defensive reflex.simple_one_for_one is a template for an unbounded, dynamically-sized population of identical children, added and removed at runtime — this project's node pool. one_for_one is a fixed, statically-declared list of distinct children known at startup — mesh_sup's own children (metrics, node_sup, dashboard).receive against that mailbox gets slower too, because a selective receive has to scan past everything already queued that does not match.receive's clauses, leaving non-matching messages queued for a later receive. The Ref, unique per request, lets a process wait specifically for the reply to this call while other, unrelated messages sit safely unmatched in the same mailbox.update_counter/3 — funnelling every access through one owning process turns that process into a serialisation point and a single point of failure for every reader, which is the exact problem ETS avoids.heart, a container orchestrator, systemd — has to watch the VM process itself.Predict the output of each, then check. All ten were run to confirm the answers.
%% 1
X = 1 == 1.0,
Y = 1 =:= 1.0,
io:format("~p ~p~n", [X, Y]).
%% 2
F = fun() ->
L = [N * 2 || N <- [1,2,3,4], N rem 2 =/= 0],
L
end,
io:format("~p~n", [F()]).
%% 3
io:format("~p~n", [length([1,2,3|[4,5]])]).
%% 4
A = try error(oops) catch _:_ -> caught end,
io:format("~p~n", [A]).
%% 5
R = (catch 1/0),
io:format("~p~n", [element(1, R)]).
%% 6
M0 = #{a => 1},
M1 = M0#{a => 2},
io:format("~p ~p~n", [M0, M1]).
%% 7
loop(0) -> done;
loop(N) -> loop(N - 1).
io:format("~p~n", [loop(3)]).
%% 8
Pid = spawn(fun() -> receive _ -> ok end end),
Alive1 = is_process_alive(Pid),
exit(Pid, kill),
timer:sleep(10),
Alive2 = is_process_alive(Pid),
io:format("~p ~p~n", [Alive1, Alive2]).
%% 9
{ok, S} = {ok, hello},
Result = case S of
hello -> matched_atom;
_ -> something_else
end,
io:format("~p~n", [Result]).
%% 10
G = fun(N) when N > 0, N < 10 -> small;
(N) when N >= 10 -> big;
(_) -> negative_or_zero
end,
io:format("~p ~p ~p~n", [G(5), G(50), G(-1)]).
1. true false
2. [2,6]
3. 5
4. caught
5. 'EXIT'
6. #{a => 1} #{a => 2}
7. done
8. true false
9. matched_atom
10. small big negative_or_zero
== compares by value across types (1 and 1.0 are numerically equal); =:= also requires the same type, and an integer is never the same type as a float. [2, 6] — note the output order matches the input order, not the doubled values' magnitude.[1,2,3|[4,5]] is exactly the same list as [1,2,3,4,5] — the | syntax conses onto the front of whatever list follows it, and a list is a valid thing to cons onto just as a single element is.try ... catch Class:Reason -> ... with a wildcard pattern catches any exception class, converting the crash into the ordinary value caught.catch Expr (not try) converts an exception into a value shaped {'EXIT', Reason} rather than propagating it; element(1, R) extracts the tag, which is the atom 'EXIT'.M0#{a := 2} or, as here, => which both inserts and updates) produces a new map; M0, already bound, is completely unaffected by anything done to build M1.done — included as a reminder that a "loop" in Erlang is just an ordinary function, with an ordinary return value, nothing special about it syntactically.exit/2 (it is blocked in receive, which is a normal, alive state, not a dead one) and genuinely dead 10ms after an exit(Pid, kill) — kill specifically is a reason that cannot be trapped even by a process with trap_exit set, unlike an ordinary exit reason.{ok, S} against the literal tuple {ok, hello} binds S to the atom hello; the case then matches the first clause exactly.fun ... end) is the only thing that looks unfamiliar; the dispatch mechanism is identical to describe/1 from the instalment's Section 2.5.Each gives a symptom and a suspect. Diagnose before opening the answer.
handle_call(get_status, _From, State) ->
#{status := S} = State,
{reply, S, State}.
%% no handle_cast/2 clause defined at allIds = [Id || {Id, _} <- ets:tab2list(mesh_registry)], % order not guaranteed
rand:seed(exsplus, {Seed, Seed, Seed}),
[attack(Id) || Id <- Ids].node_sup occasionally fails to start with {error, {already_started, Pid}}, only under a fast test script that starts and stops it repeatedly.stop_sup() ->
exit(whereis(node_sup), kill),
ok. % returns immediately; caller assumes the supervisor is gone/metrics endpoint occasionally hangs indefinitely for one specific client, never timing out, blocking that one connection forever.handle(Sock) ->
{ok, Req} = gen_tcp:recv(Sock, 0), % no timeout argument
...intensity => 3, period => 5 shuts down after only one crash during a specific test, not three.ChildSpec = #{id => mesh_node, start => {mesh_node, start_link, []},
restart => permanent, shutdown => 100},
%% test calls mesh_node:stop/1 (a deliberate, clean stop) three times,
%% then crashes it once, expecting 3 tolerated restarts still availablehandle_cast/2 (or missing entirely). A gen_server with no matching handle_cast/2 clause for an incoming cast crashes on that cast — which would at least be visible — but a module that defines no handle_cast/2 at all still compiles (the behaviour only warns, does not require every callback if defaults exist) and silently accumulates unhandled casts exactly like a receive with no matching clause, per the instalment's mailbox warning. Fix: add an explicit catch-all clause that at least logs and discards.ets:tab2list/1 has no guaranteed order. The chaos schedule seeds rand deterministically, but iterates the ids in whatever order ETS happens to return them, which is not specified to be stable across runs or even across the same table's internal rehashing. Fix: sort the ids explicitly before building the schedule, so the sequence of "which id gets attacked Nth" is deterministic, not merely the random numbers are.exit(Pid, kill) sends an asynchronous signal and returns immediately — the target process is scheduled to die but has not necessarily been fully cleaned up (its registered name released, in particular) by the time the caller proceeds to start a new one under the same name. Fix: monitor the process and wait for its 'DOWN' message before considering it gone, exactly the pattern mesh_watcher already used for the opposite direction. gen_tcp:recv/2 with no timeout blocks forever if the client opens a connection and then never sends a complete request — a slow-loris-shaped client, deliberate or not. Fix: gen_tcp:recv(Sock, 0, 5000), and handle the {error, timeout} case by closing the socket, exactly the discipline every gen_server:call/2 already has built in by default.restart => permanent instead of transient. A permanent child is restarted on any termination, including the clean, deliberate stops the test performed — each of those three intentional stops counted against the intensity budget exactly like a real crash would, leaving none left for the genuine crash that followed. This is precisely why Milestone 6 chose transient for mesh nodes: a deliberately-stopped node should not consume restart budget meant for genuine failures.gen_server mailboxes are unbounded by design; implement admission control in front of one — a wrapper that checks process_info(Pid, message_queue_len) before casting, and returns {error, overloaded} rather than casting past a configurable threshold. Measure whether this actually protects a deliberately slow-handler node from an unbounded mailbox under chaos.node_sup one at a time, waiting for each replacement to report alive before moving to the next, with a configurable delay between each — the operational tool you would actually want before deploying a code change to a live mesh.gen_statem version of mesh_node. OTP's gen_statem behaviour models a process as an explicit state machine rather than a bag of callbacks over one opaque state term. Rebuild mesh_node with alive and crashed as explicit states, and compare: what does making the states explicit catch that the map-based status field could not?pg or global from Part A, aggregate mesh_metrics:snapshot/0 across every connected distributed node into one combined view, callable from any single node.relup. Using relx's upgrade support, ship a trivial code change (a new log line in mesh_node) as a hot upgrade to a running release — no restart, the running system picks up the new code while its supervision tree and all live processes keep running. This is genuinely fiddly to get right; treat getting it to work at all as the win.Distinct from the final challenge in Part B, and smaller, but not easy.
Build a mailbox-depth-aware load shedder. Give every node's gen_server a way to report its own mailbox depth on every message handled (process_info(self(), message_queue_len), cheap enough to call routinely), publish it to mesh_metrics, and have mesh_registry:route/2 refuse to route a new message to a node whose reported depth is above a threshold, returning {error, node_overloaded} instead of adding to an already-backed-up mailbox.
Requirements: the threshold must be per-node, not global, since Milestone 8's slow-handler attack targets individual nodes, not the whole mesh; the mechanism must not itself become a bottleneck at 2,000 nodes (measure the added cost of a depth check on every routed message); and you must demonstrate, under chaos, that a genuinely overloaded node's mailbox depth stays bounded rather than growing without limit the way an unprotected one does. Hint: reporting depth on every handled message is itself extra work on the hot path — consider reporting it periodically instead, and what staleness that trades away.
gen_server specifically removes from hand-written process code.gen_server-based process with a clean call/cast API.simple_one_for_one supervisor for a dynamic pool, and a one_for_one root supervisor wiring several different kinds of children together in dependency order.gen_tcp, including a streaming (SSE) response.relx release you can start, ping, and stop as a daemon.# mesh
A simulated network of thousands of supervised Erlang processes that crash,
restart, partition and recover — a laboratory for OTP supervision, and,
in its final form, a genuinely distributed, multi-node fault-tolerant system.
No third-party dependencies beyond PropEr (dev/test only). Standard OTP.
## What it does
- Each mesh node is a gen_server, supervised by a simple_one_for_one tree
that tolerates a bounded number of crashes before giving up and
escalating — measured: exactly 3 restarts tolerated, the 4th within the
same 5-second window brings the supervisor down, on command.
- An ETS-backed registry finds any node by id, with automatic,
monitor-driven cleanup on crash — no stale entries survive a restart.
- A seeded chaos module attacks the live mesh: random crashes, slow
handlers, malformed messages — fully reproducible from one integer.
- A dashboard, served over plain HTTP from inside the system it reports
on, with a streaming (SSE) live view — no external web framework.
- Genuinely distributed: multiple real BEAM nodes, connected over a real
network, sharing message-passing code that never needed to change to
become distribution-aware.
- Ships as a standalone relx release: no Erlang installation required on
the target machine.
## Quick start
rebar3 release
_build/default/rel/mesh/bin/mesh daemon
curl http://localhost:8080/metrics
_build/default/rel/mesh/bin/mesh stop
## Architecture
mesh_sup (one_for_one)
|- mesh_metrics counters, ETS-backed
|- node_sup (simple_one_for_one)
| `- mesh_node x N gen_server, supervised, crashable on purpose
`- mesh_dashboard HTTP + SSE, reads mesh_metrics
## Testing
rebar3 do eunit, ct, proper
## Known limitations
- The dashboard has no auth; it is meant for a trusted network only.
- Distribution trusts the shared cookie; there is no additional
authentication or encryption between nodes.
- global-based failover (advanced phase) does not scale past a few
hundred nodes gracefully — this is a stated limitation of `global`
itself, not a bug in this project's use of it.
## Licence
MIT
Three deliberate choices, matching every README in this curriculum so far: it leads with measured numbers (the exact restart-intensity threshold, not an adjective); it states what it depends on and what it does not; and it has a known limitations section that names the real, specific boundary of global rather than implying the failover mechanism is production-ready as built.
A simulated distributed network in Erlang/OTP: supervised processes that crash and recover on purpose, a live HTTP dashboard, and real multi-node failover — a laboratory for "let it crash" as an engineering strategy, not a slogan.
Topics: erlang, otp, gen-server, supervisor, fault-tolerance, distributed-systems, chaos-engineering, observability, property-based-testing.
global does not scale indefinitely — its own documentation says so, and the final challenge's solution inherits that limit honestly rather than hiding it.erl -remsh) onto a live node can do anything to it — the same operational power that makes live tracing and observer so useful is a real attack surface if the distribution port is reachable by anyone who should not have it. Bind it to a private network or a VPN, never expose it publicly.gen_tcp gets none of a real framework's hardening for free — request size limits, malformed-header handling, slow-client protection beyond the one timeout this course added. Treat the hand-rolled server as a teaching tool, not a production HTTP stack.mesh_chaos against anything you do not own and fully control.Do not present this as "a chat simulation of network failures." Present it as what it is: a study of failure as a first-class, deliberately-exercised code path, ending in a real, multi-node distributed system with a documented, honest failure analysis. The narrative that makes it interesting is the escalation from hand-rolled to OTP-standard, twice.
gen_server and supervisor — same external behaviour, the gaps closed, measured: exactly 3 restarts tolerated, the 4th one not.Keep a docs/ folder with the restart-intensity transcript, the 2,000-node timing, and one architecture diagram. A reviewer who spends ninety seconds on your repository should come away knowing you tested failure on purpose and measured what it actually cost, not just that the happy path works.
| Question | What a strong answer includes |
|---|---|
| Walk me through the supervision design. | The tree shape, why simple_one_for_one fits a dynamic node pool where one_for_one fits the root, and the exact restart-intensity numbers, measured. |
| Why not just catch every exception defensively? | The "let it crash" argument, stated precisely — a caught, patched-around failure continues in an untested state; a crash-and-restart returns to a known-good one. And the honest boundary: genuine edges still need explicit error handling. |
| Tell me about a bug you found. | The registry's stale-entry leak from a missing cleanup path, or the process-dictionary id counter's hidden assumption of single-process use — how it was found, and the fix. |
| How does this differ from how you'd do fault tolerance in Go? | Structural process isolation versus disciplined defer recover(); an unrecovered panic takes the whole Go program down where an Erlang crash takes down exactly one process. Both real trade-offs, neither universally superior. |
| What happens during a network partition? | Both sides keep running, independently, each internally consistent; the danger is exclusively in reconciling state accumulated on both sides afterward, which this project's final challenge handles for pure ownership (via global) and explicitly does not solve for replicated state. |
| How would you make this production-ready? | Authentication on the dashboard and the distribution port, TLS between distributed nodes, request-size limits on the hand-rolled HTTP server, state checkpointing for real failover continuity, and replacing global with something proven at larger scale if the cluster ever needs to grow past a few hundred nodes. |
| Exactly-once delivery: how, in this system? | You cannot, structurally, the same as every other course in this curriculum that met the question — at-least-once plus an idempotent receiver is the honest answer, and this project does not currently implement it anywhere that would need it, which is worth saying plainly rather than implying it does. |
| When would you not use Erlang for this? | Anywhere CPU-bound numeric throughput on one core is the actual bottleneck — the BEAM is built for concurrency and soft real-time responsiveness, not for winning a tight numeric loop against a language that compiles closer to the metal. Name what Erlang wins on here specifically: isolation, supervision, and distribution built into the runtime rather than bolted on. |
| What would you do differently? | Build the deterministic chaos schedule (Part A4) from the start rather than discovering the seeded-but-not-fully-reproducible gap after the fact; design the metrics set before counters accumulated ad hoc; and decide the state-checkpointing story for failover before, not after, building the ownership-transfer mechanism. |
EventSource page from Milestone 10's exercise.gen_statem throughout, per Exercise C4.3 — genuinely worth doing for the whole node, not just as an exercise, once you have felt what explicit states catch.global failover mechanism at larger scale — Raft-based leader election is the standard next step, and building even a minimal version is one of the most educationally valuable extensions in this curriculum.An OTP application spanning a dozen modules, a supervision tree tested to its actual restart-intensity limit rather than assumed to work, a registry proven at 2,000 nodes with a real bug found and fixed, a seeded and reproducible chaos-injection framework, a dashboard serving a live system's own health over HTTP from inside it, a real multi-node distributed cluster, and a documented, honest failover design with its limitations stated rather than hidden. More importantly: the instinct to let something crash on purpose and trust the supervisor, rather than reaching first for a defensive catch.
The central question of this curriculum was what kinds of problems does this language make unusually natural to solve? Erlang's answer, stated as precisely as this project allows: problems where the cost of one component failing must never become the cost of the whole system failing, where isolation is worth paying a copying and messaging overhead for, and where "this will be distributed eventually" is true often enough that the language should not treat distribution as a bolt-on afterthought. Not the fastest, not the simplest to learn in an afternoon, not the language you reach for when the whole problem is a tight numeric loop. The one where letting things break, on purpose, correctly, is how the system stays up.
Courses so far have paired concurrency (Go, Erlang) against expressiveness and language design (Ruby, Racket), with Perl's text-forensics work sitting adjacent to both. Course 5 closes the curriculum with Racket, and the same comparison this course drew against Go's channels and Perl's parsing philosophy gets drawn one more time, from the opposite direction: instead of a language whose runtime already gives you the primitive you need (processes, supervision), Racket is a language built around the idea that if it does not give you the primitive you need, you can build it — as a real language of your own, checked, hygienic, and yours.
Next: Racket instalment