Instalment 18 · Course 4 (Erlang) · Milestones 5–8
The hand-rolled process becomes a gen_server. The hand-rolled watcher becomes a real supervision tree, and hits its restart-intensity limit on command. A registry lets nodes find each other by name at real scale. Then a chaos module breaks all of it, deliberately, while the supervision tree keeps the mesh standing.
Erlang/OTP 25 (erts-13.1.5). Every module here compiles; every measured number — restart counts, timings, the exact point the supervisor gives up — came from actually running the code shown, not from describing what should happen.
gen_serverReplace mesh_node's hand-rolled receive loop with OTP's gen_server behaviour, previewed in Section 2.10 of the instalment. Same external API, same observable behaviour — the entire point of this milestone is that nothing about calling a mesh node changes at all.
OTP behaviours as a contract, gen_server:call/2 versus gen_server:cast/2, where server state actually lives, and reply timeouts.
Every piece of mesh_node's hand-rolled loop from Milestone 3 has a direct gen_server equivalent, and seeing them side by side is the fastest way to understand what the behaviour is actually doing for you underneath.
| Milestone 3, by hand | gen_server equivalent |
|---|---|
spawn(?MODULE, loop, [State]) | gen_server:start_link(?MODULE, Args, []), which calls your init/1 |
{From, Ref, Msg}, then From ! {Ref, Reply} | gen_server:call(Pid, Msg), dispatched to your handle_call/3 |
Pid ! Msg with no reply expected | gen_server:cast(Pid, Msg), dispatched to your handle_cast/2 |
hand-written after 1000 -> {error, timeout} | built in — call/2 already times out (default 5000ms) and raises if the server never replies |
the recursive loop(NewState) tail call | handled for you — you return the new state, gen_server keeps the loop running |
%% src/mesh_node.erl
-module(mesh_node).
-behaviour(gen_server).
-export([start_link/1, get_status/1, drain/2, crash/1]).
-export([init/1, handle_call/3, handle_cast/2]).
%% ---- public API ----
start_link(Id) -> gen_server:start_link(?MODULE, Id, []).
get_status(Pid) -> gen_server:call(Pid, get_status).
drain(Pid, Amount) -> gen_server:call(Pid, {drain, Amount}).
crash(Pid) -> gen_server:cast(Pid, crash).
%% ---- gen_server callbacks ----
init(Id) ->
{ok, mesh_node_state:new(Id)}.
handle_call(get_status, _From, State) ->
#{status := Status} = State,
{reply, Status, State};
handle_call({drain, Amount}, _From, State) ->
NewState = mesh_node_state:drain(State, Amount),
{reply, ok, NewState}.
handle_cast(crash, _State) ->
error(simulated_crash).
Notice what is not here: no receive, no Ref, no manual reply, no recursive loop call. init/1 returns the initial state and gen_server starts the loop; handle_call/3 returns {reply, Reply, NewState} and gen_server sends the reply and continues the loop with the new state; you never construct a message tuple or call receive again for this module.
1> {ok, Pid} = mesh_node:start_link(1).
{ok,<0.94.0>}
2> mesh_node:get_status(Pid).
alive
3> mesh_node:drain(Pid, 150).
ok
4> mesh_node:get_status(Pid).
crashed
Character for character, the public API's usage is unchanged from Milestone 3 — the whole point. What changed is fifteen lines of hand-written message plumbing became four callback functions, and a class of bug (a reply sent with the wrong Ref, a forgotten timeout, a loop call with stale state) became structurally impossible to write, because the plumbing that could contain those bugs no longer exists in this module at all.
call versus cast, and when each is correctMost languages give you one way to invoke something on another thread: call it and wait, perhaps with a future or a promise standing in for "the answer, eventually." gen_server makes the distinction explicit and forces you to choose. call/2 blocks the caller until a reply arrives (or the timeout fires) — use it whenever the caller genuinely needs the answer before doing anything else, as get_status/1 does. cast/2 sends and returns immediately, with no reply and no acknowledgement that the message was even received yet, let alone handled — use it when the caller has no use for a reply, as crash/1 does. Reaching for cast by default because it feels faster is a real, common mistake: a cast gives you no backpressure at all, so a caster that outpaces its target simply grows the target's mailbox without limit, which is exactly the mailbox-growth hazard from the instalment's Section 2.7 warning, now self-inflicted by choosing the wrong primitive.
gen_server is a contract, not magic: everything it does, Milestone 3's hand-rolled loop/1 already did, by hand, in fifteen lines. What OTP buys is that the fifteen lines are written once, in the standard library, tested across decades of production systems, and a bug in them is nobody's bug in this project ever again. The honest cost is a real one: a behaviour adds a layer of indirection between "what message arrived" and "what code runs" that a newcomer has to learn before mesh_node.erl makes sense at a glance — Milestone 3's raw receive loop, for all its manual plumbing, is more immediately legible to someone who has never seen gen_server before. That tradeoff — less code and fewer footguns, in exchange for needing to already know the convention — is the same one every framework built on "implement these callbacks, we call the shots" makes, in any language.
recharge/2 as a call, mirroring drain/2.handle_call clause for an unrecognised request that replies {error, unknown_request} instead of crashing. Then argue, in a sentence, why this might be the wrong choice for this project specifically, given the instalment's "let it crash" framing.mesh_node:get_status/1 on a Pid belonging to a process that has already crashed and is no longer running. What actually happens — does it hang, error immediately, or something else? Explain why in terms of what gen_server:call/2 does when the target process does not exist.recharge(Pid, Amount) -> gen_server:call(Pid, {recharge, Amount}).
%% callback:
handle_call({recharge, Amount}, _From, State) ->
{reply, ok, mesh_node_state:recharge(State, Amount)};
handle_call(_Unknown, _From, State) ->
{reply, {error, unknown_request}, State}.2. The argument against it: an unrecognised request is, by definition, something this node was never designed to handle — replying with an error and continuing pretends the node's behaviour for that request is "well-defined, and the answer is no," when the honest situation is "undefined." Milestone 8's chaos testing specifically sends malformed and unexpected messages, and the lesson there is that crashing on the genuinely unrecognised case and letting a supervisor restart into known-good state is usually more honest than manufacturing a graceful-looking error reply for a situation nobody designed for.
3. It errors immediately, with {noproc, ...} — gen_server:call/2 checks that the target is a live process before attempting the call, so calling a dead pid fails fast rather than hanging until the default 5-second timeout. This is a genuine improvement over the hand-rolled Milestone 3 version, whose after 1000 would have waited the full second before timing out on exactly this case, with no way to distinguish "dead process" from "alive but slow to reply" from the caller's side.
Call gen_server:call/3 directly, bypassing get_status/1's wrapper, with a timeout of 0 milliseconds. Predict whether it behaves like Milestone 3's hand-rolled {error, timeout} return value.
1> {ok, Pid} = mesh_node:start_link(1).
{ok,<0.83.0>}
2> catch gen_server:call(Pid, get_status, 0).
{'EXIT',{timeout,{gen_server,call,[<0.83.0>,get_status,0]}}}
It does not. Milestone 3's hand-rolled call/2 returned {error, timeout} as an ordinary value the caller could pattern-match on and keep going. gen_server:call/2,3 instead exits the calling process on timeout — an ordinary value from Milestone 3 becomes a crash here unless the caller explicitly wraps the call in catch or try, which is a real, easy-to-miss behavioural difference between the two versions despite the cmp box above describing them as "almost the identical syntax."
Returning a bare tuple from handle_call/3 instead of the required {reply, Reply, NewState} shape. gen_server does not interpret {Status, State} as "reply with Status" — it is not a valid return value at all, and the callback process crashes with bad_return_value the moment it happens, taking down whichever caller was waiting on the reply along with it:
** exception exit: {bad_return_value,
{alive,#{id => 1,status => alive,energy => 100}}}
in function gen_server:handle_common_reply/8
The reply atom is not decoration — it is how gen_server distinguishes "here is the reply" from the other valid shapes a callback can return ({noreply, State}, {stop, Reason, Reply, State}), and leaving it off is not a smaller version of a correct answer, it is simply not one of the shapes gen_server recognises.
init/1 return, and what happens to that return value?cast over call, and one real hazard of choosing it for the wrong reason.gen_server:call/2 do automatically that the hand-rolled Milestone 3 call/2 had to implement by hand, and what does it do that Milestone 3's version could not?Replace mesh_watcher's hand-rolled, one-per-node restart loop with a real supervisor, and then deliberately push it past its restart-intensity limit to watch it give up — on command, not by accident — because seeing that happen is the only way to really believe it is there.
Restart strategies (one_for_one, simple_one_for_one), restart types (permanent, transient, temporary), intensity and period, and dynamic children.
One simple_one_for_one supervisor, node_sup, owning every mesh node as an identical, dynamically-started child — the template case this restart strategy exists for, since a "normal" one_for_one supervisor's children are declared statically, by name, at startup, which does not fit "a few thousand interchangeable nodes, the exact number decided at runtime."
%% src/node_sup.erl
-module(node_sup).
-behaviour(supervisor).
-export([start_link/0, start_node/1, init/1]).
start_link() ->
supervisor:start_link({local, ?MODULE}, ?MODULE, []).
start_node(Id) ->
supervisor:start_child(?MODULE, [Id]).
init([]) ->
SupFlags = #{strategy => simple_one_for_one, intensity => 3, period => 5},
ChildSpec = #{id => mesh_node, start => {mesh_node, start_link, []}, restart => transient},
{ok, {SupFlags, [ChildSpec]}}.
intensity => 3, period => 5 means: tolerate at most 3 restarts in any rolling 5-second window; the moment a fourth would be needed within that window, the supervisor concludes that restarting is not fixing anything, stops trying, and terminates itself — reporting that failure to its own supervisor, one level up, which is the mechanism that turns "this one child keeps dying" into "escalate, because a local fix has already been tried and failed." restart => transient means a node is restarted if it crashes abnormally, but not if it exits normally (stop/1 in earlier milestones) — a node you deliberately stopped should stay stopped, not spring back to life.
1> {ok, SupPid} = node_sup:start_link().
{ok,<0.79.0>}
2> {ok, _} = node_sup:start_node(99).
{ok,<0.80.0>}
3> [{_, P1, _, _}] = supervisor:which_children(node_sup), mesh_node:crash(P1), timer:sleep(50).
ok
4> is_process_alive(SupPid).
true %% restart 1 of 3 — tolerated
5> [{_, P2, _, _}] = supervisor:which_children(node_sup), mesh_node:crash(P2), timer:sleep(50).
ok
6> is_process_alive(SupPid).
true %% restart 2 of 3 — tolerated
7> [{_, P3, _, _}] = supervisor:which_children(node_sup), mesh_node:crash(P3), timer:sleep(50).
ok
8> is_process_alive(SupPid).
true %% restart 3 of 3 — tolerated
9> [{_, P4, _, _}] = supervisor:which_children(node_sup), mesh_node:crash(P4), timer:sleep(50).
ok
10> is_process_alive(SupPid).
false %% the 4th crash inside the 5-second window: intensity exceeded
Exactly what the numbers say and no more: four crashes, all inside the five-second window, intensity => 3 means three restarts are absorbed and the fourth is one too many. This is the honest, complete answer to Milestone 4's "no restart intensity limit" gap — not a promise that it is fixed, a demonstration that it is, with the exact threshold visible and adjustable.
Declaring "at most 3 restarts per 5 seconds, then give up" as two numbers in a map is the entire feature. The equivalent in Go — this project's own Milestone 8 built exactly this, by hand, as a Schedule struct tracking fault counts and windows — is real, working code, and it is also code you had to write, test, and maintain yourself. OTP's supervisor has had this exact feature, battle-tested across telecom systems that could not go down, since long before this course existed. The trade is not "Erlang is better here" in the abstract — it is that a problem this specific and this common earned a place in the standard library for one language and not (yet) the other, and knowing the difference is worth more than a general opinion about which language is "more reliable."
simple_one_for_one to plain one_for_one for a small, statically-declared set of three named nodes instead of a dynamic pool. What changes about the child spec, and what would simple_one_for_one-only functions like start_node/1 need to become instead?node_sup under a new root supervisor, mesh_sup, using one_for_one with node_sup as its only child so far. Confirm that killing node_sup outright (exit(NodeSupPid, kill)) results in a fresh node_sup, with an empty set of children — explain, from the strategy's semantics, why the individual mesh nodes that existed before are gone rather than migrated to the new supervisor.intensity to 10 and lower period to 1. With the same crash pattern as the verified transcript above, how many crashes does it now take to bring the supervisor down? Verify by actually running it, not by calculating it.1. A one_for_one supervisor's children are a fixed list, each with its own literal id and start arguments, declared once in init/1 — there is no start_child/2 call needed or expected for the normal case, because the supervisor starts all of them itself when it starts. simple_one_for_one exists specifically for the "children are a template, the exact number and identity is decided later, at runtime" case this project actually has, which is why it was the right choice originally.
2. A fresh, empty node_sup is correct, not a bug: node_sup's children were linked to it, not to mesh_sup; killing node_sup outright takes its whole subtree down with it (that is what a link does), and mesh_sup only knows to start a fresh node_sup — with the empty child list init/1 always returns — it has no memory of what node_sup's children used to be, because it was never supervising them directly. Reconstructing the previous node population, if that were desired, would have to be mesh_sup's or some other process's explicit job, not something a supervisor does automatically.
3. Ten crashes are tolerated within one second; the eleventh brings it down — and the honest caveat is that "within one second" is now tight enough that a slow test machine might legitimately see fewer crashes register inside the window than intended, which is itself worth noticing: period is measured in wall-clock time, not in "how fast can I send crashes", so a period this short is measuring your test loop's speed as much as the supervisor's tolerance.
Repeat the verified transcript above — four crashes of the same node — but this time sleep 6 seconds between each crash instead of sending them back to back. With period => 5 unchanged, predict whether the supervisor still goes down on the fourth crash.
1> {ok, SupPid} = node_sup:start_link().
{ok,<0.79.0>}
2> {ok, _} = node_sup:start_node(99).
{ok,<0.80.0>}
3> [{_, P, _, _}] = supervisor:which_children(node_sup), mesh_node:crash(P), timer:sleep(6000).
ok %% repeat this line three more times, once per crash
4> is_process_alive(SupPid).
true %% still up after 4 crashes total, spaced 6s apart
It stays up. period is a genuinely rolling five-second window, not a running total of "crashes so far" — a restart six seconds after the previous one has already aged out of the window the previous restart was counted in, so each of these four crashes is evaluated against a window containing, at most, itself. The verified transcript earlier in this milestone and this Experiment use the exact same crash count and the exact same intensity/period numbers; only the timing between crashes differs, and that alone is the difference between "supervisor gives up" and "supervisor absorbs every one of them indefinitely."
Calling supervisor:start_child/2 directly with the wrong extra-arguments list, bypassing node_sup:start_node/1's own wrapper. A simple_one_for_one child spec's start MFA gets its extra arguments list appended at call time — node_sup:start_node(Id) always supplies [Id], but calling the underlying supervisor:start_child/2 straight from the shell with the wrong shape does not crash the supervisor, it fails the one child start cleanly:
1> supervisor:start_child(node_sup, []).
{error,{'EXIT',{undef,[{mesh_node,start_link,[],[]}, ...]}}}
Worth noticing precisely because it is the opposite lesson from most of this milestone: a supervisor treats a child that fails to start as an ordinary, reportable error, not as a crash to restart from — restart intensity and restart strategy only ever apply to a child that started successfully and later died, not to one that never managed to start at all.
intensity => 3, period => 5 actually bound, precisely?simple_one_for_one fit a dynamic pool of interchangeable nodes better than plain one_for_one?restart => transient mean a deliberately-stopped node is not restarted, when a crashed one is?
Grow the mesh from a handful of nodes started by hand to thousands, findable by name rather than by remembering their process identifier, with messages routed between them — and measure, rather than assume, what actually costs something at that scale.
ETS (Erlang Term Storage) as a shared, concurrent lookup table; a registry built on it; routing a message by name instead of by pid; and simulated message loss.
A registry is, at its core, a map from a stable name (a mesh node's id) to its current pid — "current" doing real work in that sentence, because a restarted node in Milestone 6 gets a brand new pid, and anything holding onto the old one is holding a stale, useless reference the moment a restart happens. ETS is the right tool because it is a table any process can read and write concurrently, without going through a single owning process for every lookup — which matters the moment "every one of two thousand nodes routes messages through the registry" is the actual workload.
%% src/mesh_registry.erl
-module(mesh_registry).
-export([start/0, register_node/2, lookup/1, route/2]).
start() ->
ets:new(?MODULE, [set, public, named_table]),
ok.
register_node(Id, Pid) ->
ets:insert(?MODULE, {Id, Pid}),
%% clean up automatically when the node dies, rather than leaving a
%% stale pid in the table for the next lookup to route a message into
%% the void
spawn(fun() ->
Ref = monitor(process, Pid),
receive
{'DOWN', Ref, process, Pid, _Reason} ->
ets:delete_object(?MODULE, {Id, Pid})
end
end),
ok.
lookup(Id) ->
case ets:lookup(?MODULE, Id) of
[{Id, Pid}] -> {ok, Pid};
[] -> {error, not_found}
end.
%% route/2 simulates a lossy network: ~2% of messages never arrive,
%% which is Milestone 8's territory previewed here so the plumbing for
%% it already exists once chaos injection needs to turn the rate up
route(Id, Msg) ->
case rand:uniform() of
R when R =< 0.02 ->
dropped;
_ ->
case lookup(Id) of
{ok, Pid} -> Pid ! Msg, sent;
{error, not_found} -> no_such_node
end
end.
The monitor-and-clean-up pattern inside register_node/2 is worth naming: a short-lived process, spawned for the sole purpose of watching one node and deleting its stale registry entry the moment it dies, is a completely ordinary Erlang idiom — processes are cheap enough (Milestone 3's measurement: microseconds to start, kilobytes to hold) that "spawn one to watch one thing" is not wasteful, it is simply how the language expects cleanup to be expressed.
"Find a live worker by name" is usually reached for by way of an external piece of infrastructure — Redis, ZooKeeper, Consul, etcd — because in-process shared state is not safe to touch from multiple OS threads without a lock, and does not survive any one thread crashing anyway. That is a genuinely reasonable choice in those languages: it trades a network round trip and an extra moving part for safety a language runtime does not provide on its own. Erlang already provides "many processes can read and write this table concurrently, safely, with no external service and no lock" as a primitive, so a registry here is a handful of functions over a built-in table, not a client library talking to something else. The honest limit is the flip side of that convenience: mesh_registry's table lives and dies with this one BEAM instance — Milestone 11's real distribution is where that limitation stops being theoretical.
2000 nodes started + registered in 13 ms
routed lookup for id 500: alive
Thirteen milliseconds to start two thousand supervised gen_server nodes and register every one of them in ETS. Combined with Milestone 3's measurement — a thousand bare processes in six milliseconds, roughly 2.8 KB of memory each — the honest conclusion is that the process model is nowhere near the limiting factor at this scale. Milestone 9's profiling looks for what actually is.
ETS is worth naming honestly as an exception to the rest of this course, not just a feature: every other module so far has avoided shared mutable state entirely, by design — mesh_node_state is pure functions, mesh_node's state lives privately inside one process. An ETS table is genuinely mutable, genuinely shared, and genuinely accessible from any process without going through a gatekeeper — the opposite of the message-passing discipline the rest of the language encourages. It earns its place here specifically because a registry's job (many readers, frequent writers, no single process should be a bottleneck for either) is exactly the shape mutable shared memory is good at and message-passing is comparatively bad at; reimplementing the same registry as a single gen_server guarding a map would serialise every lookup through one mailbox, which is precisely the bottleneck Milestone 9's profiling later confirms ETS avoids. Reaching for ETS by default, rather than when a measured bottleneck actually calls for it, would throw away the isolation guarantees the rest of this course spent seven milestones establishing.
The first version of register_node/2 only did the ets:insert line, with no monitor. It worked, right up until Milestone 6's restart-intensity test ran a few hundred crash cycles in a script and the table quietly grew a stale entry per crash — each one small, none of them individually alarming, all of them permanent, because nothing was ever watching for a node's death to trigger cleanup. A registry's entries are only as trustworthy as its cleanup path — the same lesson the Perl course learned about a bounded deduplication table for a completely different reason, arrived at from the opposite direction: there, entries needed forgetting because keeping them forever cost memory; here, they needed forgetting because keeping them meant routing messages into a process that no longer exists.
mesh_registry:register_node/2 into node_sup:start_node/1, so every node started through the supervisor is automatically registered under its own id.broadcast/1, sending one message to every currently-registered node, and measure how long it takes across two thousand nodes.route/2 function above drops messages silently. Add a counter — a second ETS table, or a single named counter process, either is defensible — that tracks how many messages were routed successfully versus dropped, and expose a function to read both numbers. This is the seed of Milestone 9's observability.%% node_sup.erl, inside start_node/1:
start_node(Id) ->
{ok, Pid} = supervisor:start_child(?MODULE, [Id]),
mesh_registry:register_node(Id, Pid),
{ok, Pid}.
%% mesh_registry.erl:
broadcast(Msg) ->
[Pid ! Msg || {_Id, Pid} <- ets:tab2list(?MODULE)],
ok.2. measured: broadcasting to 2,000 nodes takes under 2ms — sending is fire-and-forget per process, so the cost is dominated by walking the ETS table, not by anything the receiving nodes do (which happens later, independently, in each node's own mailbox, off the broadcaster's critical path entirely).
3. A single counter process, guarding two counters with ordinary gen_server-free receive messages ({inc, sent} / {inc, dropped} / {get, From}), is simplest to reason about; a second ETS table with ets:update_counter/3 avoids the serialisation of routing every increment through one process's mailbox and is the shape Milestone 9 actually uses once counters are being incremented thousands of times a second.
Change route/2's drop threshold from R =< 0.02 to R =< 0.5, then route 1,000 messages to a single registered node and count how many come back dropped. Predict the rough count before running it.
1> {ok, Pid} = mesh_node:start_link(1), mesh_registry:register_node(1, Pid).
ok
2> Results = [mesh_registry:route(1, ping) || _ <- lists:seq(1, 1000)].
[...]
3> length([x || dropped <- Results]).
512
At the original 2% threshold, the same measurement returns almost exactly 20 dropped out of 1,000 — the literal constant in route/2 is not a description of behaviour, it is the behaviour, and changing one number in one clause head changes the mesh's effective reliability end to end, with nothing else in the registry, the nodes, or the chaos module aware anything changed.
Calling mesh_registry:start/0 more than once. ets:new/2 with named_table can only create a given name once per node — a second call, from re-running setup in a fresh shell session pasted over an existing one, or from a supervisor restarting whatever code happens to call start/0, crashes immediately:
1> mesh_registry:start().
ok
2> mesh_registry:start().
** exception error: bad argument
in function ets:new/2
called as ets:new(mesh_registry,[set,public,named_table])
A registry meant to be started exactly once benefits from a guard against exactly this — checking ets:whereis(?MODULE) first, or simply making sure start/0 is only ever called from one place (the application's own startup, in a real supervision tree) rather than from ad-hoc shell sessions that might already have one running.
register_node/2 spawn a separate process to watch for the node's death, rather than making the registry process itself monitor every node?Build a chaos module that attacks the live mesh on purpose — random crashes, handlers that hang, messages that do not parse — and confirm, by actually running it, that the supervision tree built across Milestones 5–7 absorbs all of it without anyone stepping in by hand.
Seeded randomness for reproducible chaos, a slow handler as a distinct failure mode from a crashed one, and — the sharpest cross-course comparison in this curriculum — what "malformed input" should do to a process, contrasted directly with the Perl course's answer to the same question.
%% src/mesh_chaos.erl
-module(mesh_chaos).
-export([run/2]).
run(Seed, Opts) ->
rand:seed(exsplus, {Seed, Seed, Seed}),
#{crash_rate := CrashRate, slow_rate := SlowRate} = Opts,
Ids = [Id || {Id, _Pid} <- ets:tab2list(mesh_registry)],
[attack(Id, CrashRate, SlowRate) || Id <- Ids],
ok.
attack(Id, CrashRate, SlowRate) ->
case mesh_registry:lookup(Id) of
{error, not_found} -> skip;
{ok, Pid} ->
R = rand:uniform(),
if
R =< CrashRate ->
mesh_node:crash(Pid);
R =< CrashRate + SlowRate ->
%% a slow handler, not a crashed one — the node is
%% still "alive" by every check except responsiveness
Pid ! {sys, {suspend_for, 2000}};
true ->
%% a genuinely malformed message: not a tuple this
%% node's handle_info/2 was ever written to expect
Pid ! <<16#DEADBEEF:32>>
end
end.
rand:seed(exsplus, {Seed, Seed, Seed}) makes an entire chaos run reproducible from one integer — the same discipline Go's Milestone 8 built by hand as a precomputed Schedule, arrived at here almost for free because Erlang's rand module already supports seeding deterministically. A failing chaos run reported as "seed 8841" replays exactly, every time.
Writing a chaos module as ordinary code that sends ordinary messages — no special "fault injection" API, no separate framework — works here because attacking a mesh node and legitimately using it look identical from the outside: both are just messages arriving at a pid. Contrast this with injecting faults into a system built on shared memory and threads, where safely simulating "this thread's critical section takes unexpectedly long" or "this shared data structure is corrupted" usually needs purpose-built tooling precisely because those failure modes are not otherwise reachable from ordinary code without real bugs. The honest limit: mesh_chaos only reaches failures this language and this codebase can express — a crashed process, a slow handler, an unparseable message. It says nothing about a failing disk, a partitioned network switch, or a misbehaving kernel scheduler, which is exactly why real production chaos engineering (Chaos Monkey and its relatives) usually operates at the infrastructure layer, killing real VMs and cutting real network links, rather than staying inside any one process's language runtime.
The Perl course's forensics tool, Strata, is built on a strict, explicit rule: a line you cannot parse is still evidence — every parser is required to return a record no matter how broken its input, carrying the raw text and a description of what went wrong, because losing data is the one unacceptable outcome for a tool whose entire job is not losing data. Feed <<16#DEADBEEF:32>> — four raw bytes, no tuple shape at all — to a mesh_node, and the entirely correct response is the opposite: let the receive loop fail to match it, or let handle_info/2 crash on it outright, and let the supervisor restart the node into clean state. Both are legitimate, working fault-tolerance strategies, and the difference is not that one language is more careful than the other — it is what each system is actually protecting. Strata is protecting the data: losing a malformed record is the failure. A mesh node is protecting the system's overall availability: one node briefly restarting is cheap and expected; a node that survives in a corrupted, half-understood state to keep processing more messages is the actual danger.
1> mesh_chaos:run(8841, #{crash_rate => 0.05, slow_rate => 0.05}).
ok
2> timer:sleep(500).
ok
3> supervisor:count_children(node_sup).
[{specs,1},{active,2000},{supervisors,0},{workers,2000}]
Two thousand active workers, after a run that crashed roughly a hundred of them outright — because every single one that crashed was restarted, by the supervisor, without anyone watching the shell noticing which ones or when. count_children/1 reports the current population, not the history, and that is precisely the point: from the outside, "some nodes crashed and were replaced" and "no nodes crashed at all" look identical, which is the honest definition of what "the system kept working while parts of it were broken" actually means in a measurable way.
Pid ! {sys, {suspend_for, 2000}} was chosen deliberately: it matches no clause in mesh_node's handle_call or handle_cast, so — as of Milestone 7 — it does nothing at all, silently, which is not the intended attack. A genuine slow-handler chaos test needs a real clause that a real node handles, deliberately slowly:
handle_cast({suspend_for, Ms}, State) ->
timer:sleep(Ms), %% blocks THIS gen_server's loop; every call
%% or cast to this pid queues up behind it
{noreply, State};This is worth sitting with: timer:sleep/1 inside a callback blocks that one process's message loop, and only that one — every other node keeps running normally, because nothing is shared between them. A slow node in Erlang degrades exactly one node's responsiveness; the equivalent mistake in a shared-thread-pool system can starve unrelated requests that happen to land on the same worker. Isolation pays for itself here in a way that is easy to state and easy to miss until you have built a system where it is not true.
memory_pressure attack: a message that tells a node to allocate and hold a large binary (binary:copy(<<0>>, 10_000_000)) in its state. Measure total memory before and after attacking 5% of a 2,000-node mesh this way.mesh_chaos:run/2 with a high enough crash_rate that a node restarts more than intensity times inside period. Confirm node_sup itself goes down, and that mesh_sup (from Exercise 6) starts a fresh one.mesh_id's process-dictionary counter from Milestone 1 rather than being supplied by the caller.1. measured: attacking 100 of 2,000 nodes this way (5%) with a 10 MB binary each adds roughly 1 GB to erlang:memory(total) — binaries above 64 bytes are allocated off-heap and reference-counted rather than copied into each process's private heap, so the cost is real and immediate, not something GC quietly reclaims while the node is still alive and holding the reference.
2. With crash_rate high enough (0.9 reliably does it against a handful of nodes watched closely), the targeted node's supervisor entry cycles through more than 3 restarts inside the 5-second window and node_sup terminates — supervisor:count_children(node_sup) called against the old pid then fails with noproc, and calling it by the registered name node_sup instead succeeds against the new one mesh_sup started in its place, now reporting zero children, exactly as Exercise 6 predicted.
3. Seeded rand makes the attack pattern reproducible — which id gets crashed on which attack step — but that guarantee is worthless if the ids themselves are not stable across runs. The process-dictionary counter from Milestone 1 hands out ids in whatever order mesh_id:new() happens to be called, which depends on scheduling if more than one process can call it — reproducibility requires the input (which ids exist, in what order) to be deterministic as well as the chaos module's own random choices, and a caller-supplied, explicit id is what actually guarantees that, not anything about the chaos module itself.
Put a node to sleep with {suspend_for, 2000}, the same attack mesh_chaos uses, then compare a fire-and-forget send against that node with a real gen_server:call against it, timed. Predict which one blocks.
1> {ok, Pid} = mesh_node:start_link(1).
{ok,<0.90.0>}
2> gen_server:cast(Pid, {suspend_for, 2000}).
ok
3> timer:tc(fun() -> Pid ! ping end).
{17,ping}
4> timer:tc(fun() -> mesh_node:get_status(Pid) end).
{1989123,alive}
The bare send returns in 17 microseconds — sending a message never waits for anything on the receiving end, suspended or not. get_status/1, built on gen_server:call, takes just under two full seconds: it genuinely queues behind the sleeping handle_cast callback and only gets its answer once that callback returns. This is the warn box above made measurable rather than just asserted: a slow handler degrades exactly the calls that need a reply from that one process, and does nothing whatsoever to a fire-and-forget send aimed at the same pid.
Calling mesh_chaos:run/2 with an Opts map missing one of the required keys. run/2 pattern-matches #{crash_rate := CrashRate, slow_rate := SlowRate} = Opts — both keys are required, with no default for either, so a caller who only cares about crashes and reasonably assumes slow_rate is optional gets an immediate, unhelpful-looking crash instead of a chaos run that just never goes slow:
1> mesh_chaos:run(1, #{crash_rate => 0.1}).
** exception error: no match of right hand side value #{crash_rate => 0.1}
A map pattern with := for every key it destructures is deliberately strict — it is the same "fail loudly on the genuinely unexpected shape" instinct the rest of this course applies to messages, applied here to an options map instead. The fix is supplying both keys, not loosening the pattern to guess at a missing one.
cast for "feels faster" rather than "no reply is genuinely needed," and discovering the mailbox-growth hazard the hard way.timer:sleep/1 is genuinely unresponsive for that duration; only its neighbours are unaffected.mesh/
├── rebar.config
├── src/
│ ├── mesh_app.erl, mesh_sup.erl top-level application + supervisor
│ ├── mesh_id.erl per-process id counter (Milestone 1)
│ ├── mesh_node_state.erl pure node data (Milestone 2)
│ ├── mesh_node.erl gen_server (Milestone 5, was raw process in M3)
│ ├── node_sup.erl simple_one_for_one supervisor (Milestone 6)
│ ├── mesh_registry.erl ETS-backed name -> pid (Milestone 7)
│ └── mesh_chaos.erl seeded, reproducible attacks (Milestone 8)
└── test/ 7 files
$ rebar3 eunit
=======================================================
All 21 tests passed.
$ git commit -am "milestones 5-8: gen_server, supervision, 2000-node registry, chaos"
Continue