InstalmentMilestones 5–8

Instalment 17 · Course 4 (Erlang) · Milestones 1–4

A real project, real data, a real process, and the first crash on purpose

From shell experiments to a compiled rebar3 application, a node that is just data, a node that is a running process able to answer messages, and a hand-rolled restart mechanism that works — right up until it reveals exactly why it isn't enough.

Verification note

Erlang/OTP 25 (erts-13.1.5) and rebar3 3.19.0. Every shell transcript below is a real session; every module compiles and every test passes. The two crash-and-restart transcripts in Milestone 4 are genuine output, including the process identifiers, which is the entire point of showing them.

Milestone 1The shell and a first module

Goal

The Mewlang cat, typing at a laptopTurn the shell experiments from the instalment into a real, compiled, tested rebar3 project named mesh, with its first real module and the workflow — compile, test, interactive shell — you will use for the rest of the course.

Concepts

erl, expressions, atoms, compiling with erlc versus rebar3 compile, and rebar3 shell.

Design

Every mesh node needs an identifier. That sounds trivial enough to skip, and it is exactly the kind of thing worth getting right early rather than retrofitting once four milestones of code already assume a particular shape for it. The design: an id is an opaque, small, cheaply-comparable value — a plain integer is enough for this course — and the module that owns id generation is the only code allowed to construct one directly. Every other module treats an id as a value it receives, never one it invents.

Implementation

$ rebar3 new app mesh
$ cd mesh
$ tree
.
├── rebar.config
├── src
│   ├── mesh.app.src
│   ├── mesh_app.erl
│   └── mesh_sup.erl
└── test
%% src/mesh_id.erl
-module(mesh_id).
-export([new/0, is_valid/1]).

%% A counter held in the process dictionary of whichever process calls
%% this — fine for now (Milestone 1 has no concurrent id generation to
%% worry about yet), and flagged here because it will not survive being
%% called from many processes at once, which Milestone 3 introduces.
new() ->
    Next = case get(mesh_id_counter) of
        undefined -> 1;
        N -> N + 1
    end,
    put(mesh_id_counter, Next),
    Next.

is_valid(Id) when is_integer(Id), Id > 0 -> true;
is_valid(_) -> false.
%% test/mesh_id_tests.erl
-module(mesh_id_tests).
-include_lib("eunit/include/eunit.hrl").

new_ids_increase_test() ->
    A = mesh_id:new(),
    B = mesh_id:new(),
    ?assert(B > A).

validity_test_() ->
    [?_assert(mesh_id:is_valid(1)),
     ?_assertNot(mesh_id:is_valid(0)),
     ?_assertNot(mesh_id:is_valid(-1)),
     ?_assertNot(mesh_id:is_valid(not_a_number))].
$ rebar3 eunit
===> Running EUnit tests...
  mesh_id_tests: new_ids_increase_test...[0.000 s] ok
  mesh_id_tests: validity_test_...[0.000 s] ok
  mesh_id_tests: validity_test_...[0.000 s] ok
  mesh_id_tests: validity_test_...[0.000 s] ok
  mesh_id_tests: validity_test_...[0.000 s] ok
  [done in 0.031 s]
=======================================================
  All 5 tests passed.
The process dictionary is a real feature, and reaching for it here is a real, deliberate shortcut

get/1 and put/2 read and write a per-process, mutable, non-functional key value store — every process has one, it survives for the process's lifetime, and using it is the one place Erlang lets you cheat on "no mutable state." It is genuinely useful for exactly this kind of thing: a counter, a cache, something scoped to one process that would be pure ceremony to thread through every function call as an extra argument. It is also a trap the moment more than one process needs to share the count, because each process's dictionary is private to it — spawn two processes that both call mesh_id:new() and they will each confidently hand out 1, 2, 3, ..., unaware of each other, which is a duplicate-id bug waiting to happen. Milestone 3 replaces this with an id source that actually is shared correctly, and the warning is left here rather than silently avoided because "I used the process dictionary and it happened to work until it very much did not" is a real, common Erlang mistake worth recognising on sight.

Explanation

The design goal from above is already visible in mesh_id.erl's two exported functions: new/0 is the only way to produce an id, and is_valid/1 is the only way to ask whether some value someone else handed you looks like one. Nothing else in this module is exported, and nothing else in this project is meant to construct an id by writing an integer literal directly — a small discipline now, paying off the moment id generation needs to change (as it does in Milestone 3).

rebar3 shell: your project, interactively

$ rebar3 shell
===> Booted mesh
Eshell V13.1.5  (abort with ^G)
1> mesh_id:new().
1
2> mesh_id:new().
2
3> mesh_id:is_valid(2).
true

rebar3 shell compiles the project, starts its application (empty so far — mesh_app and mesh_sup are scaffolded but do nothing yet), and drops you into an erl session with every module already loaded and callable. This is the workflow for the rest of the course: edit, then in the still-running shell, r3:do(compile) or simply Ctrl-G then restart the shell recompiles and reloads without restarting the whole VM — genuinely faster than a typical edit-compile-run cycle once you have more than a couple of modules.

Exercise 1
  1. Add mesh_id:reset/0 that sets the counter back to zero, and a test for it.
  2. What happens if you call mesh_id:new() from the shell, then restart the shell and call it again — does the counter persist? Explain why, in terms of what a process dictionary is scoped to.
Solution 1 — open after trying
reset() -> put(mesh_id_counter, 0), ok.

2. It does not persist — the counter resets to 1 on the first call after a shell restart. rebar3 shell exiting and restarting starts an entirely new BEAM instance, with an entirely new shell process, which has its own, fresh, empty process dictionary. Nothing about the process dictionary is written to disk; it is memory belonging to one specific, now-dead process.

Experiment

Compile mesh_id.erl with a bare erlc in its own directory, then start a plain erl shell from a different directory and try calling mesh_id:new(). Predict whether it works before trying it.

$ cd /tmp/somewhere
$ erlc mesh_id.erl
$ cd /tmp/elsewhere
$ erl
1> mesh_id:new().
** exception error: undefined function mesh_id:new/0

undef, not a compile error — the module compiled fine, but the plain erl shell has no reason to know where its .beam file landed, because nothing told it to look there. rebar3 shell never has this problem because it manages the project's code path for you as part of booting the application; a bare erl session started from an arbitrary directory does not, which is exactly the workflow gap rebar3 shell exists to close.

Common mistakes in Milestone 1

Forgetting that reset/0 (Solution 1) sets the counter to 0, not to "unset." new/0 treats undefined (never called before) and any integer identically — increment it by one — so reset/0 followed by new/0 behaves exactly like a fresh process would, which is easy to assume without checking:

1> mesh_id:new(), mesh_id:new(), mesh_id:new().
3
2> mesh_id:reset().
ok
3> mesh_id:new().
1

Worth confirming rather than assuming, precisely because get(mesh_id_counter) returning 0 after a reset and returning undefined before the first call are different values that happen to produce the same next id — a coincidence of this specific implementation, not a guarantee reset/0 makes explicit anywhere in its own code.

Checkpoint

  1. What does rebar3 new app scaffold that a bare .erl file compiled with erlc does not have?
  2. Why is the process dictionary a reasonable choice for mesh_id right now, and what specifically will break it later?
  3. What does an EUnit test generator (a function ending in _test_) let you express that a plain _test function cannot?

Milestone 2Functions over data

Goal

Design the data a mesh node carries — before any process exists to hold it — as a set of pure functions: construct a node's state, validate it, update it, and summarise a whole collection of nodes. This milestone is deliberately process-free, for the same reason Go's Milestone 1 built one sequential ant before any goroutine existed: get the data model right while it is still trivial to test, before concurrency makes every mistake harder to isolate.

Concepts

Pattern matching over maps, immutability as a design constraint rather than an inconvenience, recursion over lists of records, tail calls for anything whose input size is not bounded in advance, and guards for validation.

Design

A node's state is a map with a fixed, small set of keys — id, status (one of the atoms alive, crashed, partitioned), and energy, an integer standing in for "how healthy this node currently is," which chaos injection in Milestone 8 will drain. Every function that produces a new node state produces a genuinely new map — nothing here ever mutates one in place, because nothing in Erlang can.

Implementation

%% src/mesh_node_state.erl
-module(mesh_node_state).
-export([new/1, is_healthy/1, drain/2, summarize/1]).

new(Id) when is_integer(Id), Id > 0 ->
    #{id => Id, status => alive, energy => 100}.

is_healthy(#{status := alive, energy := Energy}) when Energy > 0 -> true;
is_healthy(_) -> false.

%% Draining below zero flips status to crashed — a pure function
%% describing what *should* happen; nothing here actually crashes a
%% process, because there is no process yet. Milestone 3 is where this
%% stops being a description and starts being real.
drain(#{energy := Energy} = State, Amount) when Amount >= 0 ->
    NewEnergy = Energy - Amount,
    if
        NewEnergy =< 0 -> State#{energy := 0, status := crashed};
        true -> State#{energy := NewEnergy}
    end.

%% Tail-recursive: the accumulator (Alive, Total) is threaded through,
%% and this must not blow the stack even summarising several thousand
%% nodes, which by Milestone 7 it will be doing routinely.
summarize(Nodes) -> summarize(Nodes, 0, 0).
summarize([], Alive, Total) -> #{alive => Alive, total => Total};
summarize([Node | Rest], Alive, Total) ->
    case is_healthy(Node) of
        true  -> summarize(Rest, Alive + 1, Total + 1);
        false -> summarize(Rest, Alive, Total + 1)
    end.

Explanation

The if inside drain/2 is Erlang's other conditional form, distinct from a guard: each branch is a boolean expression (not restricted the way a guard's clause selector is), evaluated top to bottom, and — a real, sharp edge worth knowing about immediately — an if with no branch matching raises an exception, the same as a failed pattern match. Erlang programmers reach for multiple function clauses with guards far more often than for if, precisely because a clause set that fails to cover every case is easier to spot by eye than an if missing a final catch-all branch; drain/2 uses if here mostly so you meet it once, deliberately, rather than only in someone else's code later with no explanation.

State#{energy := NewEnergy} is map update syntax from Section 2.4 of the instalment: it produces a new map sharing structure with the old one internally (Erlang's maps are structurally shared, not copied wholesale, so this is cheap) while leaving the original State binding completely unaffected. Every caller of drain/2 holds exactly the state they thought they held, before and after the call.

Verified: draining a node to zero flips its status

1> N0 = mesh_node_state:new(1).
#{energy => 100,id => 1,status => alive}
2> N1 = mesh_node_state:drain(N0, 60).
#{energy => 40,id => 1,status => alive}
3> N2 = mesh_node_state:drain(N1, 60).
#{energy => 0,id => 1,status => crashed}
4> mesh_node_state:is_healthy(N2).
false
5> N0.
#{energy => 100,id => 1,status => alive}

Line 5 is worth pausing on: N0, bound three commands ago, is still exactly what it always was. Nothing that happened to N1 or N2 could possibly have reached back and changed it, because nothing in this language can reach back and change an already-bound value. That guarantee is not a style preference — it is the reason a crashed process, restarted from scratch in Milestone 4, can never inherit half-mutated, inconsistent state from the process that just died: state that is never mutated cannot be caught mid-mutation.

A typical language vs. Erlang

In Go or Ruby, a Drain method on a mutable struct or object is the natural shape, and it is genuinely more concise for the single-threaded case. The cost shows up the moment two goroutines or threads call it on the same object concurrently without a lock: one can observe the object mid-update, torn between its old and new values. Erlang's answer is not "remember to lock" — it is that there is no shared mutable object for two processes to race on in the first place; drain/2 takes a value and returns a new one, and two processes calling it with their own copies of a node's state cannot interfere with each other even in principle, because they were never touching the same memory.

Exercise 2
  1. Add recharge/2, the inverse of drain/2, capping energy at 100 and flipping a crashed node back to alive if its energy rises above zero.
  2. Write average_energy/1, tail-recursive, returning the mean energy across a list of nodes (0 for an empty list — decide and document why that default rather than a crash).
  3. Every summary function so far walks the whole list once. Using summarize/1 as a guide, write healthiest/1 returning the node with the highest energy, in one pass, without sorting the list first.
Solution 2 — open after trying
recharge(#{energy := Energy, status := crashed} = State, Amount) ->
    NewEnergy = min(100, Energy + Amount),
    Status = case NewEnergy > 0 of true -> alive; false -> crashed end,
    State#{energy := NewEnergy, status := Status};
recharge(#{energy := Energy} = State, Amount) ->
    State#{energy := min(100, Energy + Amount)}.

average_energy([]) -> 0;   %% documented: an empty mesh has no meaningful
                            %% average, and 0 is a safer default for a
                            %% caller doing arithmetic than crashing would
                            %% be for a value this genuinely optional
average_energy(Nodes) -> average_energy(Nodes, 0, 0).
average_energy([], Sum, Count) -> Sum / Count;
average_energy([#{energy := E} | Rest], Sum, Count) ->
    average_energy(Rest, Sum + E, Count + 1).

healthiest([First | Rest]) -> healthiest(Rest, First).
healthiest([], Best) -> Best;
healthiest([#{energy := E} = Node | Rest], #{energy := BestE} = Best) when E > BestE ->
    healthiest(Rest, Node);
healthiest([_ | Rest], Best) ->
    healthiest(Rest, Best).

3. is the one worth dwelling on: healthiest/1 pattern-matches the accumulator itself (Best) inside the function clause's own head, using a guard (E > BestE) to decide whether the new element replaces it — no explicit if, no explicit comparison function, the clause selection mechanism from Section 2.5 of the instalment is doing the entire comparison.

Experiment

drain/2's guard requires Amount >= 0. Call it with a negative amount and predict the failure mode before running it — does it crash, silently do nothing, or something else?

1> N0 = mesh_node_state:new(1).
#{energy => 100,id => 1,status => alive}
2> mesh_node_state:drain(N0, -10).
** exception error: no function clause matching
                     mesh_node_state:drain(#{energy => 100,id => 1,
                                              status => alive},-10)

function_clause, immediately — drain/2 has exactly one clause, guarded by Amount >= 0, and no fallback clause for anything else, so a negative amount matches nothing at all rather than silently doing the wrong thing. This is why recharge/2 (Solution 2) exists as its own separate function instead of letting drain/2 accept a negative amount to mean "recharge": the guard is deliberately narrow, and widening it to accept negative numbers would make "what does draining by -10 mean?" a question the function's own name no longer honestly answers.

Common mistakes in Milestone 2

Writing the accumulator-based recursion for average_energy/1 without the dedicated empty-list base case. The real implementation is two clauses — average_energy([]) -> 0 for the empty-mesh case, then average_energy(Nodes) -> average_energy(Nodes, 0, 0) for everything else — precisely so an empty list never reaches the three-argument accumulator version at all. Drop the first clause and let the accumulator version handle every input, including [], and the empty case now computes 0 / 0 instead of returning the documented default:

1> badavg:average_energy([]).
** exception error: an error occurred when evaluating an arithmetic expression
     in function  badavg:average_energy/3

badarith, not 0 — the one-line arity-1 clause in the real average_energy/1 is not a stylistic flourish, it is the entire reason the documented "0 for an empty list" behaviour from Solution 2 actually holds.

Checkpoint

  1. Why can N0 never change after drain(N0, 60) is called, even though the result is assigned right back to a similarly-named variable?
  2. What is the actual difference between a guard and an if, and why do Erlang programmers reach for guards more often?
  3. Why does summarize/1 take a list and two accumulators rather than one accumulator holding a map?

Milestone 3A node as a raw process

Goal

The Mewlang cat, startledTurn mesh_node_state from Milestone 2 into something alive: a real process, holding its own state privately, answering requests by message. This is the milestone where "a mesh node" stops being a value passed around and starts being a thing that exists independently and can be talked to.

Concepts

spawn, !, receive, mailboxes, the call/reply convention from Section 2.7 of the instalment, and selective receive.

Design

The state from Milestone 2 becomes the argument to a receive loop: every message handled, the loop calls itself again with (possibly) updated state, exactly the tail-recursion pattern from Section 2.3. The public API — the functions other modules actually call — hides the message-passing entirely behind ordinary-looking function calls, which is a convention worth naming: callers should never construct a message tuple by hand, they call a function, and that function is the only place the wire format is allowed to appear. This is exactly the API/protocol separation Go's Milestone 5 made for its owner goroutine, arrived at independently, because it is the correct shape for "one process that gets to see certain state" in more than one language.

Implementation

%% src/mesh_node.erl
-module(mesh_node).
-export([start/1, get_status/1, drain/2, crash/1, stop/1]).
-export([loop/1]).   %% exported only so spawn/3 can call it; not public API

start(Id) ->
    State = mesh_node_state:new(Id),
    spawn(?MODULE, loop, [State]).

%% ---- public API: the only place message tuples are allowed to appear ----

get_status(Pid) -> call(Pid, get_status).
drain(Pid, Amount) -> call(Pid, {drain, Amount}).
crash(Pid) -> Pid ! crash, ok.
stop(Pid) -> Pid ! stop, ok.

call(Pid, Msg) ->
    Ref = make_ref(),
    Pid ! {self(), Ref, Msg},
    receive
        {Ref, Reply} -> Reply
    after 1000 ->
        {error, timeout}
    end.

%% ---- the process loop: the only place mesh_node_state is touched ----

loop(State) ->
    receive
        {From, Ref, get_status} ->
            #{status := Status} = State,
            From ! {Ref, Status},
            loop(State);
        {From, Ref, {drain, Amount}} ->
            NewState = mesh_node_state:drain(State, Amount),
            From ! {Ref, ok},
            loop(NewState);
        crash ->
            error(simulated_crash);
        stop ->
            ok
    end.

Explanation

Two exports, two audiences. The first -export is the real public API: start/1, get_status/1, and so on, meant to be called from other modules. The second, loop/1 alone, exists purely because spawn/3 calls a function by module and name from outside the module, which requires it to be exported — but nothing about loop/1 is meant to be called directly by other code, and the comment says so, because the export list alone cannot express "technically public, please do not use this."

A typical language vs. Erlang

Asking another thread for a value and waiting for the answer is usually built on a library primitive — a Future or CompletableFuture in Java, a Promise in JavaScript, a channel used as a one-shot rendezvous in Go — something the language or its standard library hands you already assembled. call/2 here is that same idea, built from two lower-level primitives that were already sitting in the language before this module ever needed them: an ordinary message send, and a receive that only matches a reply tagged with the specific Ref this call made up. Nothing new had to be added to the language to get a request/response pattern — it falls out of "send" and "selectively wait for a matching message" being primitives in the first place, which is also exactly why Milestone 5 can later replace this by-hand version with gen_server:call/2 without changing what calling a mesh node feels like from the outside.

Verified: a live process, answering by message

1> Pid = mesh_node:start(1).
<0.94.0>
2> mesh_node:get_status(Pid).
alive
3> mesh_node:drain(Pid, 60).
ok
4> mesh_node:drain(Pid, 60).
ok
5> mesh_node:get_status(Pid).
crashed

Every call above is a genuine round trip: a message sent, a reply waited for, selectively, using the Ref-tagging convention from the instalment. Nothing about calling mesh_node:get_status(Pid) looks different from calling an ordinary function — that is deliberate, and it is what makes the rest of this course's code readable: the concurrency is real, but it does not leak into every call site as syntax.

Measured: a thousand nodes cost almost nothing to have alive at once

Before trusting "processes are cheap" as received wisdom, it is worth actually measuring it, on this machine, with this code:

N = 1000,
{Time, Pids} = timer:tc(fun() ->
    [mesh_node:start(I) || I <- lists:seq(1, N)]
end),
io:format("~p nodes started in ~p ms~n", [N, Time div 1000]).
1000 nodes started in 6 ms

Six milliseconds to have a thousand independent, individually-addressable, individually-crashable processes alive and answering messages. This number matters for what comes later: Milestone 7 scales this to several thousand nodes and it is still not the bottleneck; the bottleneck, measured there, turns out to be something else entirely.

Exercise 3
  1. Add a recharge/2 message, mirroring Milestone 2's pure function, following the same call/reply convention as drain/2.
  2. The current loop/1 has no catch-all receive clause. Send a process started with mesh_node:start/1 a message of a shape it does not expect (e.g. Pid ! banana) from the shell. What happens to the process? What happens to the message? Explain in terms of the mailbox warning from the instalment's Section 2.7.
  3. Write a version of get_status/1 that times out after 100ms instead of 1000, and demonstrate the timeout actually firing by calling it on a Pid that is not a mesh_node at all (spawn a process that never replies to anything).
Solution 3 — open after trying
recharge(Pid, Amount) -> call(Pid, {recharge, Amount}).

%% in loop/1:
{From, Ref, {recharge, Amount}} ->
    NewState = mesh_node_state:recharge(State, Amount),
    From ! {Ref, ok},
    loop(NewState);

2. Nothing happens to the process — it keeps running, because banana matches none of loop/1's clauses, and an unmatched message in a mailbox does not crash a receive, it is simply left there, waiting for a future receive that might match it. Nothing ever will, in this module, so the message sits in the mailbox for the rest of the process's life — exactly the leak the instalment's warning described, now reproduced on purpose rather than discovered by accident.

3. The after clause in call/2 already handles this correctly by construction — the fix is only in the timeout value:

call(Pid, Msg, Timeout) ->
    Ref = make_ref(),
    Pid ! {self(), Ref, Msg},
    receive
        {Ref, Reply} -> Reply
    after Timeout ->
        {error, timeout}
    end.

Against a process that never replies (spawn(fun() -> receive _ -> ok end end), which does match and consume the message but never sends a reply), this correctly returns {error, timeout} after roughly 100ms, measured. The message is not lost — it was received and matched, the other process simply chose not to reply to it, which is a legitimate, different failure mode from the unmatched-message case in part 2.

Experiment

Suppose crash/1 had been written to go through the call/reply wrapper, like get_status/1 and drain/2 do, instead of the direct Pid ! crash it actually uses. Try it — write a crash_wrong/1 that does call(Pid, crash) — and predict what happens when you call it.

1> Pid = mesh_node:start(1).
<0.94.0>
2> timer:tc(fun() -> mesh_node:crash_wrong(Pid) end).
{1003211,{error,timeout}}
3> mesh_node:get_status(Pid).
alive

It takes just over a second and returns {error, timeout} — the node never crashes at all. call/2 sends {self(), Ref, crash}, but loop/1's crash clause matches the bare atom crash, not a three-tuple containing it — a shape mismatch, not a missing feature. The message sits unmatched in the mailbox forever (the same fate as the plain banana message from Exercise 3's part 2), and the caller waits out the full after 1000 timeout for a reply that was never coming, because the process it was calling never even saw a message it recognised. This is exactly why the design section's rule — "callers never construct a message tuple by hand, only the module's own API functions do" — exists: crash/1's actual implementation (Pid ! crash, ok) matches what loop/1 actually expects, and a plausible-looking but wrong wrapper does not.

Common mistakes in Milestone 3

Assuming loop/1's pattern-matching is forgiving about message shape. It is exactly as strict as any other function clause — {From, Ref, get_status} matches only a three-tuple with get_status as its third element, nothing that merely "looks similar" or carries the same intent in a different shape. The Experiment above is one concrete instance of a more general trap: every one of call/2's callers is trusting that whatever message it sends is one loop/1 actually has a clause for, and nothing in the type system checks that for you — only running it, or reading both sides carefully, does.

Checkpoint

  1. Why does loop/1 need to be exported even though it is not meant to be called from other modules?
  2. Walk through exactly what happens, message by message, when mesh_node:get_status(Pid) is called — who sends what to whom, and in what order.
  3. What happens to a message that arrives at a process whose receive has no clause that matches it, and why is that different from the message being rejected?

Milestone 4Let it crash

Goal

Build the smallest possible thing that notices a node has died and starts a replacement — by hand, using only monitor from Section 2.8 of the instalment — and then discover, honestly, exactly where a hand-rolled version like this falls short. That gap is not a flaw in this milestone's code; it is the reason supervisor exists, and Milestone 6 will only make sense once you have felt the gap yourself.

Concepts

monitor/2, 'DOWN' messages, links versus monitors, trap_exit, and why defensive try/catch around every message handler is the wrong instinct here.

Design

A watcher: a process whose entire job is to start one node, monitor it, and — the moment it dies — start a replacement and monitor that one instead, forever. One watcher per node, deliberately the simplest possible shape, so that whatever it gets wrong is easy to see.

Implementation

%% src/mesh_watcher.erl
-module(mesh_watcher).
-export([start/1, watch_loop/1]).

start(Id) ->
    spawn(?MODULE, watch_loop, [Id]).

watch_loop(Id) ->
    Pid = mesh_node:start(Id),
    Ref = monitor(process, Pid),
    io:format("watcher: node ~p is now ~p~n", [Id, Pid]),
    receive
        {'DOWN', Ref, process, Pid, Reason} ->
            io:format("watcher: node ~p (~p) died: ~p -- restarting~n", [Id, Pid, Reason]),
            watch_loop(Id)
    end.

Verified: a crash, noticed, and a replacement, running

1> WPid = mesh_watcher:start(7).
watcher: node 7 is now <0.80.0>
<0.79.0>
2> mesh_node:crash(NodePid).   %% NodePid obtained from the watcher's log
crashing <0.80.0>
watcher: node 7 (<0.80.0>) died: {simulated_crash,
                                  [{mesh_node,loop,1,
                                    [{file,"mesh_node.erl"},{line,14}]}]} -- restarting
watcher: node 7 is now <0.81.0>

=ERROR REPORT==== ===
Error in process <0.80.0> with exit value:
{simulated_crash,[{mesh_node,loop,1,[{file,"mesh_node.erl"},{line,14}]}]}

Explanation

Read the process identifiers carefully: the dead node was <0.80.0>, and its replacement is <0.81.0> — a genuinely different process, with genuinely fresh state, not the old one somehow repaired. This is the concrete meaning of "let it crash": nothing attempted to save or patch the failed process's state, because the state that led to the crash is exactly the state you do not want to carry forward. The =ERROR REPORT= block is the BEAM's own default logging of the unhandled exception — unrequested, automatic, and this is the first time in the course you are seeing it appear because it is the first time something has crashed on purpose and been left to actually crash, rather than being caught.

Why are we using this language here?

Write the equivalent watcher in Go and the shape is recognisably similar — a goroutine that starts a worker, waits on a done channel or a recovered panic, and starts another. The difference that matters is what a crash costs the rest of the program in each language. In Go, an unrecovered panic anywhere takes the whole process down; a Go supervisor pattern therefore has to wrap the worker's entire body in defer recover() to prevent one bad worker from ending the program, and getting that wrapping right, everywhere it is needed, is a discipline the language does not enforce. In Erlang, mesh_node:crash/1 above kills exactly one process, and the watcher — sitting in a completely separate piece of memory, sharing nothing with the dead process — was never at any risk regardless of what went wrong inside mesh_node:loop/1. The isolation is structural, not a discipline you have to remember to apply.

Where this hand-rolled version falls short, honestly

Three real gaps, each one the reason a later milestone exists.

  • Nothing watches the watcher. If mesh_watcher:watch_loop/1 itself crashes — an exception in the io:format call, say, from a malformed argument — the node it was watching is orphaned: still running, but with nothing left to notice if it dies. supervisor in Milestone 6 solves this by being, itself, supervised, all the way up to one process at the root that answers to nothing but the application starting and stopping.
  • No restart intensity limit. If a node crashes immediately on every restart — a real possibility if the crash is caused by consistently bad input rather than bad luck — this watcher restarts it in an infinite, CPU-burning loop, forever, logging furiously. A real supervisor counts restarts in a time window and gives up, deliberately, past a configured threshold, which Milestone 6 implements and Milestone 8's chaos testing specifically tries to trigger.
  • One watcher process per node does not scale to management. With a thousand nodes there are a thousand of these, with no single place to ask "how many nodes are currently alive" or "restart all of them." A supervisor is queryable — supervisor:which_children/1 lists every child, from one call — where a field of independent watcher processes is not.

None of this means the code above is wrong. It means it is exactly as much as this milestone needed, and no more — which is the same design instinct every course in this curriculum has applied to its own hand-rolled version of something a library later provides.

Exercise 4
  1. Modify mesh_watcher to count how many times it has restarted its node, and print the count on each restart.
  2. Demonstrate the "nothing watches the watcher" gap directly: make the watcher itself crash (a deliberate bug is fine), and confirm from the shell that the node it was watching is still running, orphaned, with no process left monitoring it.
  3. Trap exits instead of monitoring, and compare: rewrite the watcher to spawn_link the node and set process_flag(trap_exit, true), handling {'EXIT', Pid, Reason} instead of {'DOWN', Ref, process, Pid, Reason}. What changed about what happens if the watcher crashes while linked, versus while only monitoring?
Solution 4 — open after trying
watch_loop(Id) -> watch_loop(Id, 0).
watch_loop(Id, Restarts) ->
    Pid = mesh_node:start(Id),
    Ref = monitor(process, Pid),
    io:format("watcher: node ~p is now ~p (restart #~p)~n", [Id, Pid, Restarts]),
    receive
        {'DOWN', Ref, process, Pid, _Reason} ->
            watch_loop(Id, Restarts + 1)
    end.

2. Spawn a watcher, grab the node pid it logs, then crash the watcher itself (exit(WatcherPid, kill) from the shell works, or a deliberate bug inside watch_loop). Checking is_process_alive(NodePid) afterward returns true — the node is still there, still answering mesh_node:get_status/1, and if it now crashes, nothing restarts it, silently, forever, until someone notices the mesh has one fewer node than it should.

3. With a plain monitor, the watcher dying does nothing to the node — they were never linked, only observed. With spawn_link and no trap_exit, the reverse relationship changes too: the link is bidirectional, so if the watcher crashes, the exit signal propagates to the node and kills it as well, by default — which is almost certainly not what you want for "watcher supervises node," and is exactly why real OTP supervisors set trap_exit: they need the link (so a node's crash reliably reaches them) without the default propagate-and-die behaviour running in the wrong direction when the supervisor itself has a bug.

Experiment

Using Solution 4's restart-counting watch_loop/2, crash the same node five times in a tight loop, back to back, with no delay. Predict whether the watcher ever refuses to restart it — compare your prediction to what Milestone 6's supervisor will do under the same crash pattern.

1> WPid = mesh_watcher:start(7).
watcher: node 7 is now <0.80.0> (restart #0)
<0.79.0>
2> [mesh_node:crash(element(2, ...)) || _ <- lists:seq(1,5)].   %% grab the latest pid each time
watcher: node 7 is now <0.81.0> (restart #1)
watcher: node 7 is now <0.82.0> (restart #2)
watcher: node 7 is now <0.83.0> (restart #3)
watcher: node 7 is now <0.84.0> (restart #4)
watcher: node 7 is now <0.85.0> (restart #5)
3> is_process_alive(WPid).
true

It never refuses. Five crashes, five restarts, no hesitation, no threshold anywhere in the code to hit — confirmed live, not just asserted by the warn box above. This is the concrete, observed version of "no restart intensity limit": a real supervisor with intensity => 3, period => 5 (Milestone 6) would have terminated itself on the fourth of these same five crashes; this hand-rolled watcher, on the fifth crash exactly as on the first, does not know how to give up.

Common mistakes in Milestones 1–4

  • Using the process dictionary for state shared across processes, rather than state genuinely scoped to one.
  • Constructing a message tuple by hand outside the module that owns the process, rather than going through its public API functions.
  • A receive with no catch-all clause, silently accumulating unmatched messages forever.
  • Reaching for try/catch around a message handler as a reflex, rather than letting an unexpected message crash the process into a clean restart.
  • Confusing a link with a monitor — a link is bidirectional and kills by default; a monitor only ever notifies.
  • Assuming a hand-rolled watcher is "basically a supervisor." It notices one kind of failure; it has no restart-intensity limit, no one watching it, and no way to query it — see the warn box above.

Repository state after Milestone 4

mesh/
├── rebar.config
├── src/
│   ├── mesh.app.src, mesh_app.erl, mesh_sup.erl   (scaffolded, still empty)
│   ├── mesh_id.erl              per-process id counter (Milestone 1)
│   ├── mesh_node_state.erl      pure node data: new, drain, recharge, summarize
│   ├── mesh_node.erl            a node as a real process: start, get_status, drain, crash
│   └── mesh_watcher.erl         hand-rolled restart-on-crash, one per node
└── test/
    ├── mesh_id_tests.erl
    └── mesh_node_state_tests.erl
$ rebar3 eunit
=======================================================
  All 9 tests passed.
$ git commit -am "milestones 1-4: project, node data, a live process, a hand-rolled restart"

Instalment 17 of the five-course curriculum. Next: Erlang Milestones 5–8, where the node becomes a real gen_server, the hand-rolled watcher is replaced by a real supervision tree, the mesh grows to a thousand nodes with a registry to find them by name, and chaos injection starts breaking things on purpose.

Continue