Milestones 5–8Wrap-up

Instalment 19 · Course 4 (Erlang) · Milestones 9–12

Watching it, then showing it, then actually spreading it across machines

Counters and safe tracing on a live system, a browser dashboard served from inside the mesh it reports on, two real BEAM nodes on a real network talking to each other, and a release you could hand to someone else to run.

Verification note

Erlang/OTP 25, rebar3 3.19.0. The distributed-node transcript in Milestone 11 is two genuinely separate erl processes, connected over a real (loopback) network, not a simulation of the idea. The release in Milestone 12 was built, started as a daemon, pinged, and stopped, for real.

Milestone 9Observability

Goal

The Mewlang cat, wearing glasses, looking confidentTurn "is the mesh healthy" from a question you answer by staring at shell output into one you answer with counters, and learn to inspect a live system safely — including one already under chaos attack — without stopping it or guessing.

Concepts

ETS-backed counters at scale, observer as a GUI window into a running node, and dbg for safe, bounded tracing of a live system.

Design

One ETS table, four pre-seeded counters (sent, dropped, crashed, restarted), each incremented from wherever the corresponding event actually happens rather than inferred afterwards by polling. This is deliberately the same shape as Milestone 7's registry — a named, public ETS table any process can touch directly — because it is the same kind of problem: many independent processes need to update one piece of shared state, often, without funnelling every update through a single bottleneck process first. Pre-seeding the four keys in start/0 rather than letting inc/1 create them on demand is a deliberate, narrow design choice too — see "Common mistakes in Milestone 9" below for what it costs.

Implementation

%% src/mesh_metrics.erl
-module(mesh_metrics).
-export([start/0, inc/1, snapshot/0]).

start() ->
    ets:new(?MODULE, [set, public, named_table]),
    [ets:insert(?MODULE, {K, 0}) || K <- [sent, dropped, crashed, restarted]],
    ok.

inc(Key) ->
    ets:update_counter(?MODULE, Key, 1).

snapshot() ->
    maps:from_list(ets:tab2list(?MODULE)).

Explanation

ets:update_counter/3 is a single atomic increment inside the table itself — no read, modify, write race between two processes incrementing the same counter concurrently, because the increment never leaves ETS to happen in Erlang code at all. This is the fix Exercise 7 asked for, generalised: a counter per metric, incremented from wherever the event actually happens (mesh_metrics:inc(crashed) added to mesh_chaos:attack/3, mesh_metrics:inc(dropped) added to mesh_registry:route/2), read from anywhere with one ets:tab2list/1 call.

observer: a GUI onto a running node

1> observer:start().
ok

This opens a window — genuinely a window, requiring a display, which is the WSL2/Linux-desktop recommendation from the instalment paying off here — showing live process counts, memory use per application, and a browsable process tree you can click into to see any single process's mailbox size, current function, and state. For a mesh of two thousand nodes, observer's "Applications" tab showing the actual live shape of the supervision tree from Milestone 6 — mesh_sup, node_sup, two thousand identical mesh_node leaves — is worth seeing once with your own eyes: it is the architecture diagram from the instalment, except it is real and it updates.

Tracing a live system without guessing

1> dbg:tracer().
{ok,<0.98.0>}
2> dbg:p(all, call).
{ok,[...]}
3> dbg:tpl(mesh_node_state, drain, x).
{ok,[...]}
4> mesh_node_state:drain(mesh_node_state:new(1), 30).
(<0.9.0>) call mesh_node_state:drain(#{energy => 100,id => 1,status => alive},30)
(<0.9.0>) returned from mesh_node_state:drain/2 -> #{energy => 70,id => 1,
                                                     status => alive}
#{energy => 70,id => 1,status => alive}
5> dbg:stop().
ok

dbg is the standard library's own tracing facility — every call to mesh_node_state:drain/2, anywhere in the running system, from any process, prints its arguments and return value, live, without stopping anything or redeploying anything. This is genuinely dangerous run carelessly: tracing every call to a function called thousands of times a second on a live production system can produce more trace output than the system can keep up with, which is itself a new, self-inflicted performance problem. dbg:tpl/3's x argument (a wildcard match specification for "trace everything") is fine for the deliberately small, deliberately local demonstration above; a real diagnostic session on a busy system needs a narrower match specification, or the third-party recon library's recon_trace, which adds a hard cap on trace message count specifically so a debugging session cannot itself become an incident.

Why are we using this language here?

Attaching a debugger to a running Go or Ruby process to trace a specific function's calls, live, on a system you cannot restart, is possible but is not a routine, expected operation the way it is here — it usually means dlv attach to a process you are prepared to pause, or an APM agent instrumented in ahead of time. dbg, and the whole idea that erl can attach to and interrogate a running production node as an ordinary, sanctioned part of operating the system, is a direct descendant of "this VM is running a telephone exchange that cannot be stopped to debug it." The honest cost: this capability is also a real security surface — a remote shell into a live node, if reachable by someone who should not have it, can do anything to that node — which is why Milestone 11's distributed setup needs a real, secret shared cookie, not the default.

Exercise 9
  1. Wire mesh_metrics:inc/1 into mesh_chaos:attack/3 and mesh_registry:route/2, and confirm snapshot/0 reflects a chaos run accurately.
  2. Trace only calls where the drained amount exceeds 50, using a real match specification instead of the wildcard x. (Hint: dbg:fun2ms/1 compiles an ordinary fun into a match specification.)
Solution 9 — open after trying
%% 2.
dbg:tpl(mesh_node_state, drain, dbg:fun2ms(fun([_, Amt]) when Amt > 50 -> ok end)).

The fun passed to fun2ms/1 is never actually called — it is inspected at compile time and turned into a match specification the tracer evaluates natively, inside the runtime, so that only genuinely matching calls generate trace output at all, which is what makes narrow tracing safe on a busy system where the wildcard version would not be.

Experiment

Set up a trace specification with dbg:tpl/3 but skip the dbg:p/2 call that actually switches tracing on for a set of processes, call mesh_node_state:drain/2, and predict — before running it — whether anything prints.

1> dbg:tracer().
{ok,<0.90.0>}
2> dbg:tpl(mesh_node_state, drain, x).
{ok,[...]}
3> mesh_node_state:drain(mesh_node_state:new(1), 30).
#{energy => 70,id => 1,status => alive}
4> dbg:p(all, call).
{ok,[...]}
5> mesh_node_state:drain(mesh_node_state:new(1), 30).
(<0.86.0>) call mesh_node_state:drain(#{energy => 100,id => 1,status => alive},30)
(<0.86.0>) returned from mesh_node_state:drain/2 -> #{energy => 70,id => 1,
                                                     status => alive}
#{energy => 70,id => 1,status => alive}
6> dbg:stop().
ok

Step 3's call produces no trace output at all, even though the exact match specification from step 2 is still installed — dbg:tpl/3 only decides what a matching call should record; dbg:p/2 is the separate switch that decides which processes are being watched in the first place, and until it is flipped on, nothing is watched. Only step 5, after both are set, produces the two trace lines. Forgetting the second call is an easy way to conclude tracing "isn't working" when it was never actually switched on.

Common mistakes in Milestone 9

Assuming inc/1 works for any atom you hand it. It does not — start/0 pre-seeds exactly four keys, and ets:update_counter/3 requires the key to already exist; there is no implicit "create it with a starting value of zero" step. Calling mesh_metrics:inc(timeout) for a fifth kind of event nobody thought to seed crashes the calling process immediately:

1> mesh_metrics:start().
ok
2> mesh_metrics:inc(timeout).
** exception error: bad argument
     in function  ets:update_counter/3
        called as ets:update_counter(mesh_metrics,timeout,1)
        *** argument 2: not a key that exists in the table
     in call from mesh_metrics:inc/1 (mesh_metrics.erl, line 10)

Adding a new metric means adding it to the seed list inside start/0 first, not calling inc/1 on a new atom and expecting the table to grow to accommodate it.

Checkpoint

  1. Why does ets:update_counter/3 avoid a race that a plain read-then-write increment would not?
  2. What can observer show you about a running system that reading its source code cannot?
  3. Why is tracing every call to a hot function on a busy live system its own kind of danger?

Milestone 10The live dashboard

Goal

Serve mesh_metrics:snapshot/0 over plain HTTP, from a process running inside the same system it is reporting on, so the mesh's health is one browser tab away rather than a shell session and a function call.

Concepts

A minimal HTTP server built directly on gen_tcp — no external web framework — and streaming updates to a connected client.

Design

Standard-library-only, deliberately: gen_tcp's {packet, http_bin} mode already parses HTTP request lines and headers for you, which covers everything this milestone actually needs. A real production service would reach for cowboy or ranch for connection pooling, keep-alive and proper HTTP/1.1 compliance; this course stays on the standard library because the interesting part — running an HTTP server inside an OTP application, supervised like everything else — does not require them.

Implementation

%% src/mesh_dashboard.erl
-module(mesh_dashboard).
-export([start/1, accept_loop/1]).

start(Port) ->
    {ok, LSock} = gen_tcp:listen(Port,
        [binary, {packet, http_bin}, {active, false}, {reuseaddr, true}]),
    spawn_link(?MODULE, accept_loop, [LSock]).

accept_loop(LSock) ->
    {ok, Sock} = gen_tcp:accept(LSock),
    spawn(fun() -> handle(Sock) end),
    accept_loop(LSock).

handle(Sock) ->
    case gen_tcp:recv(Sock, 0) of
        {ok, {http_request, 'GET', {abs_path, <<"/metrics">>}, _}} ->
            drain_headers(Sock),
            Body = jsonify(mesh_metrics:snapshot()),
            respond(Sock, Body);
        {ok, {http_request, 'GET', _OtherPath, _}} ->
            drain_headers(Sock),
            respond(Sock, "{\"error\":\"not_found\"}");
        _ ->
            ok
    end,
    gen_tcp:close(Sock).

drain_headers(Sock) ->
    case gen_tcp:recv(Sock, 0) of
        {ok, http_eoh} -> ok;
        {ok, _} -> drain_headers(Sock);
        _ -> ok
    end.

respond(Sock, Body) ->
    Resp = ["HTTP/1.1 200 OK\r\nContent-Type: application/json\r\n",
            "Content-Length: ", integer_to_list(iolist_size(Body)), "\r\n",
            "Connection: close\r\n\r\n", Body],
    gen_tcp:send(Sock, Resp).

jsonify(Map) ->
    Pairs = [io_lib:format("\"~s\":~p", [K, V]) || {K, V} <- maps:to_list(Map)],
    ["{", lists:join(",", Pairs), "}"].

Verified: a real request, against a real running mesh

1> mesh_metrics:start(), mesh_dashboard:start(8080).
<0.102.0>
2> os:cmd("curl -s http://127.0.0.1:8080/metrics").
"{\"sent\":0,\"dropped\":0,\"crashed\":0,\"restarted\":0}\n"

Explanation

spawn_link/3 to start accept_loop rather than a plain spawn/3 matters here specifically: if the accept loop crashes — a malformed connection it was not written to handle, say — the link means whatever process started the dashboard finds out immediately, rather than the dashboard silently going deaf while everything else keeps running. In the application's real supervision tree (this milestone's exercise), the dashboard is its own supervised child for exactly this reason: a dead dashboard should be noticed and restarted the same as a dead node, not left quietly unreachable.

Every accepted connection is handled in its own freshly-spawned process (spawn(fun() -> handle(Sock) end)), so one slow or malicious client blocks nothing except its own connection — the same isolation property Milestone 8 demonstrated for mesh nodes, applying identically here because it is the same language feature, not a special case built for HTTP.

Why are we using this language here?

Running an HTTP server as an ordinary supervised child of the same application it reports on — not a separate process on a separate machine polling this one — is a genuinely different shape than most stacks reach for by default, and it works here specifically because a gen_tcp listener is just another process: spawn it, link it, hand it to a supervisor, and it gets exactly the same restart guarantee as a mesh node. The honest cost is real, though: this handler has no keep-alive, no chunked transfer encoding, no TLS, and speaks just enough HTTP/1.1 to answer a GET — reaching for cowboy or ranch is the right call the moment this needs to serve real browser traffic rather than a JSON endpoint for curl and EventSource. A language with a mature HTTP-server ecosystem built in or one import away — Go's net/http, or almost any web framework in Ruby or Perl — gets you further, faster, for the serving part itself; what OTP buys back is that whatever server you end up with slots into the exact same supervision tree as everything else, with no separate deployment story of its own.

Streaming updates needs the connection kept open, on purpose

The handler above closes the socket after one response — correct for a single GET /metrics, wrong for "streaming updates," which the milestone's concept list asks for. The fix is a long-lived handler using Server-Sent Events, an HTTP response that never ends, writing one small update every few seconds instead of closing:

stream(Sock) ->
    gen_tcp:send(Sock,
        "HTTP/1.1 200 OK\r\nContent-Type: text/event-stream\r\n"
        "Cache-Control: no-cache\r\nConnection: keep-alive\r\n\r\n"),
    stream_loop(Sock).

stream_loop(Sock) ->
    Body = jsonify(mesh_metrics:snapshot()),
    case gen_tcp:send(Sock, ["data: ", Body, "\n\n"]) of
        ok -> timer:sleep(1000), stream_loop(Sock);
        {error, _} -> ok   %% client disconnected; stop quietly, do not crash
    end.

A browser's own EventSource API consumes this natively, with no client-side library — the same "browser has this built in already" fact the Go course's server-sent-events viewer relied on, arrived at independently for a completely different language's dashboard. The {error, _} clause matters more than it looks: without it, writing to a socket the client already closed crashes this handler process on every single disconnect, which is harmless in isolation (it is supervised, it is one connection) but is exactly the kind of "technically survivable but needlessly noisy" failure worth avoiding when it is this cheap to avoid.

Exercise 10
  1. Wire mesh_dashboard into the application's supervision tree as a proper supervised child, and confirm — by killing its pid directly — that it comes back and starts serving requests again.
  2. Add a route, GET /nodes/<id>, returning that one node's status via mesh_registry:lookup/1 and mesh_node:get_status/1, with a proper 404 JSON body for an id that does not exist.
  3. Serve a minimal static HTML page at GET / that opens an EventSource against /stream and renders the live counters as they arrive. This does not need to be pretty; it needs to update without the page reloading.
Solution 10 — open after trying

2. The routing addition is a new clause in handle/1's match, extracting the id from the path:

{ok, {http_request, 'GET', {abs_path, Path}, _}} ->
    drain_headers(Sock),
    case binary:split(Path, <<"/nodes/">>) of
        [<<>>, IdBin] ->
            Id = binary_to_integer(IdBin),
            case mesh_registry:lookup(Id) of
                {ok, Pid} -> respond(Sock, jsonify(#{status => mesh_node:get_status(Pid)}));
                {error, not_found} -> respond(Sock, "{\"error\":\"no such node\"}")
            end;
        _ ->
            respond(Sock, "{\"error\":\"not_found\"}")
    end;

The rest follows the same shape already established: parse, look up, respond — nothing about serving a second route required restructuring anything.

Experiment

With mesh_metrics:start() and mesh_dashboard:start(Port) already running from the verified transcript above, call mesh_metrics:inc(sent) five times from the same shell, then curl /metrics again without restarting the dashboard. Predict whether the new count shows up.

1> [mesh_metrics:inc(sent) || _ <- lists:seq(1,5)].
[1,2,3,4,5]
2> os:cmd("curl -s http://127.0.0.1:8080/metrics").
"{\"dropped\":0,\"sent\":5,\"crashed\":0,\"restarted\":0}\n"

It does, immediately, with no restart and no extra plumbing: mesh_metrics's ETS table is public and shared, handle/1 reads it fresh on every single request via mesh_metrics:snapshot(), and the dashboard process itself holds no cached copy of the counters anywhere — it is a stateless reader sitting in front of state that genuinely lives elsewhere.

Common mistakes in Milestone 10

Assuming the router matches a path prefix rather than the exact bytes. {http_request, 'GET', {abs_path, <<"/metrics">>}, _} matches only that exact binary — a request for /metrics?since=0, which a browser or a monitoring tool building a cache-busting URL would send without a second thought, has a different abs_path and silently falls through to the 404 clause:

$ curl -s http://127.0.0.1:8080/metrics
{"dropped":0,"sent":0,"crashed":0,"restarted":0}
$ curl -s 'http://127.0.0.1:8080/metrics?since=0'
{"error":"not_found"}

A route this simple needs to split the path on ? and match only the part before it — exactly the fix Exercise 10's /nodes/<id> solution already does with binary:split/2, just not yet applied to /metrics itself.

Checkpoint

  1. Why does the accept loop use spawn_link for itself but a plain spawn for each individual connection handler?
  2. What makes a Server-Sent-Events response different from an ordinary HTTP response, at the protocol level?
  3. Why does forgetting to handle {error, _} from gen_tcp:send/2 in a streaming handler matter less than it would in a system with no supervision at all?

Milestone 11Actually distributed

Goal

The Mewlang cat, looking up curiouslyTurn "node" from a metaphor (Course-4's simulated mesh participants) into the literal thing (a real, separate BEAM instance) by connecting two genuinely different erl processes over a real network connection, and send a message from one to a process registered on the other.

Concepts

Named nodes, the shared secret cookie, net_kernel:connect_node/1, and what a network partition actually looks like from inside the system experiencing it.

Design

Two BEAM instances, each started with a name (-sname for same-host testing, -name for a real, fully-qualified hostname across real machines) and the same secret cookie — a shared value that authorises two nodes to trust each other, checked automatically on every connection attempt, with no further configuration needed once it matches. Once connected, sending a message to a registered name on a remote node uses almost the identical syntax as sending to one locally.

Implementation and verified transcript

$ erl -sname nodea -setcookie meshcookie
(nodea@myhost)1> register(pinger, self()).
true
(nodea@myhost)2> receive {ping, From} -> From ! {pong, node()} end.
# in a second terminal:
$ erl -sname nodeb -setcookie meshcookie
(nodeb@myhost)1> net_kernel:connect_node(nodea@myhost).
true
(nodeb@myhost)2> nodes().
[nodea@myhost]
(nodeb@myhost)3> {pinger, nodea@myhost} ! {ping, self()}.
{ping,<7062.87.0>}
(nodeb@myhost)4> flush().
Shell got {pong,nodea@myhost}
ok

Explanation

Every part of that is real: two independent operating-system processes, each running its own BEAM, each with its own supervision trees, its own memory, its own scheduler — connected over TCP (via epmd, the Erlang Port Mapper Daemon, a small always-running process that maps node names to ports on a host), exchanging one message. {pinger, nodea@myhost} ! Msg — a registered name paired with a node name instead of a bare pid — is the entire syntactic difference between sending a message locally and sending one to a process that happens to live on a different machine.

Why are we using this language here?

This is the payoff the instalment promised: nothing about mesh_registry, node_sup, or any gen_server callback written across ten milestones needed to change to make this work, because the message-passing discipline was never actually single-machine-specific — a pid is a pid, and Erlang's distribution layer transparently routes a message to wherever the pid's process actually lives. Compare this to what "distribute it" meant in Course 1: Go's Milestone 12 genuinely rewrote the world's owner-goroutine protocol into a TCP wire protocol with explicit encoding and decoding, because Go's channels are a local-only primitive with no distributed equivalent built into the language. Erlang's distribution is not free — it costs real latency and a real trust boundary, both explored below — but it is not a rewrite.

What a partition actually looks like from inside the system, and why "split brain" is not a bug you patch

Disconnect the two nodes mid-session — erlang:disconnect_node(nodea@myhost), called from nodeb — and both sides keep running, independently, each believing itself to be the whole system: nodeb's nodes() now returns [], and if mesh_registry's entries were being kept consistent by replicating them between nodes (a natural next step neither of these two milestones actually builds), each side would now be perfectly willing to register a different node under the same id, with neither side aware of the conflict, because neither side can currently see the other at all. This is split brain, and the honest thing to say about it is that Erlang's distribution layer gives you the tools to detect a partition (nodes() shrinking, net_kernel monitor events) and absolutely does not give you a built-in answer for what to do about conflicting state accumulated during one — that is a genuinely hard distributed-systems problem (consensus, vector clocks, last-write-wins with a defined tiebreak) that a message-passing runtime cannot solve for you by itself, any more than Go's channels could.

Exercise 11
  1. Start a third node, nodec, and connect all three pairwise. Confirm each node's nodes() lists the other two.
  2. Start mesh_registry and a handful of nodes on nodea only. From nodeb, connected, call rpc:call(nodea@myhost, mesh_registry, lookup, [1]) and explain what rpc:call/4 is doing that a bare message send could not — it returns a value synchronously, which a fire-and-forget ! cannot.
  3. Disconnect the two nodes, then reconnect them (net_kernel:connect_node/1 again). Do previously-registered names and running processes on either side survive the disconnect? What does that imply about what a partition actually threatens — the processes themselves, or only their ability to reach each other?
Solution 11 — open after trying

3. Everything survives — a disconnect at the distribution layer does not kill any process on either side; both BEAM instances keep running exactly as before, they simply stop being able to see or message each other. This is the honest, complete answer to "what does a partition threaten": not the individual nodes' own local correctness — each side is still internally consistent, still supervising its own children correctly — only the system's ability to act as one coherent whole while the partition lasts. That distinction is precisely why the warn box above says split brain is a state-reconciliation problem, not an availability problem: both sides stay available throughout, which is exactly what makes reconciling them afterward hard.

Experiment

Start nodeb with the wrong cookie on purpose and try to connect. Then, without restarting either node, fix the cookie at runtime and try again.

$ erl -sname nodeb -setcookie wrongcookie
(nodeb@myhost)1> net_kernel:connect_node(nodea@myhost).
false
(nodeb@myhost)2> erlang:set_cookie(nodea@myhost, meshcookie).
true
(nodeb@myhost)3> net_kernel:connect_node(nodea@myhost).
true
(nodeb@myhost)4> nodes().
[nodea@myhost]

The cookie check runs fresh on every connection attempt, not once at boot — erlang:set_cookie/2 changes what nodeb will present (or accept) on its next attempt, live, with no restart of either BEAM instance needed. This is worth sitting with precisely because it cuts against the instinct that authentication is something decided once, at startup: here it is re-checked, cheaply, on every single connection.

Common mistakes in Milestone 11

Assuming a cookie mismatch raises an error you can catch. It does not — net_kernel:connect_node/1 simply returns false, exactly the same value it would return if the other node did not exist at all, or was unreachable over the network. There is no exception, no log message on the calling side by default, and no way to distinguish "wrong cookie" from "no such node" from the return value alone:

(nodeb@myhost)1> net_kernel:connect_node(nodea@myhost).
false

Diagnosing a failed connection in practice means checking the cookie explicitly (erlang:get_cookie() on each side) rather than trusting the return value to tell you what went wrong — a real, easy-to-miss gap between "the API reports failure" and "the API explains the failure."

Checkpoint

  1. What has to match between two nodes before they will trust each other enough to connect?
  2. What changed in mesh_registry or node_sup's code to make distributed message-passing work? Why?
  3. What does a network partition actually do to the processes running on either side of it?

Milestone 12Releases and property tests

Goal

Package mesh as a self-contained release you could hand to someone else to run without them installing Erlang first, and add one property-based test that searches for a bug rather than checking one specific example.

Concepts

relx releases via rebar3 release, Common Test as EUnit's heavier sibling for integration-shaped tests, PropEr for property-based testing, and sys.config as the standard place a release's environment-specific settings live.

Design

A release bundles everything mesh needs to run standalone — including, critically, a configuration file, config/sys.config, separate from the compiled code itself. Up to this point every tunable number in this project — node_sup's restart intensity and period from Milestone 6, most obviously — has been a literal value baked into the module that uses it, which is fine for a course where every milestone is run from source, and genuinely wrong for anything meant to be deployed: changing how many restarts an operator is willing to tolerate should not require recompiling and re-releasing the whole application. OTP's answer is application:get_env/2,3, reading values an operator can override per-environment, in sys.config, with no code change at all — released here, in Milestone 12, because a release is the first point in this course where "per-environment configuration" is actually a real, separate concept from "the code."

Implementation

$ rebar3 release
===> Release successfully assembled: _build/default/rel/mesh

$ _build/default/rel/mesh/bin/mesh daemon
$ _build/default/rel/mesh/bin/mesh ping
pong
$ _build/default/rel/mesh/bin/mesh stop
ok

Explanation

That is a complete, standalone deployment artefact: a copy of the exact Erlang runtime it was built against, every dependency, and mesh's own compiled code, in one directory — _build/default/rel/mesh/bin/mesh is a shell script that boots the whole application as a background daemon, with ping/stop commands for basic lifecycle management, requiring nothing installed on the target machine except compatible system libraries. This is the same category of artefact as a statically-linked Go binary, arrived at by a very different route: Go gets there by compiling to native machine code with no runtime dependency; Erlang gets there by bundling its own runtime alongside the code that needs it.

Application configuration: sys.config

node_sup:init/1's restart limits move from two hardcoded numbers to two calls to application:get_env/3, each with the old literal as its fallback default — so the module still behaves exactly as Milestone 6 described it if no configuration is supplied at all:

%% src/node_sup.erl
init([]) ->
    Intensity = application:get_env(mesh, restart_intensity, 3),
    Period = application:get_env(mesh, restart_period, 5),
    SupFlags = #{strategy => simple_one_for_one, intensity => Intensity, period => Period},
    ChildSpec = #{id => mesh_node, start => {mesh_node, start_link, []}, restart => transient},
    {ok, {SupFlags, [ChildSpec]}}.
%% config/sys.config
[
  {mesh, [
    {restart_intensity, 1},
    {restart_period, 5}
  ]}
].

application:get_env(App, Key, Default) is the three-argument form specifically because it returns the bare value (Default itself, or whatever was configured) rather than the {ok, Value} | undefined shape the two-argument version returns — one line, no case statement needed, at the cost of not being able to tell "explicitly configured" apart from "fell back to the default," which this particular setting has no need to distinguish.

Verified: the same crash pattern, a different threshold, because sys.config said so

$ erl -pa _build/default/lib/mesh/ebin -config config/sys
1> application:load(mesh).
ok
2> application:get_env(mesh, restart_intensity, 3).
1
3> {ok, SupPid} = node_sup:start_link().
{ok,<0.84.0>}
4> {ok, P1} = supervisor:start_child(node_sup, [1]).
{ok,<0.86.0>}
5> mesh_node:crash(P1), timer:sleep(50), is_process_alive(SupPid).
true    %% restart 1 of 1 (configured) — tolerated
6> [{_, P2, _, _}] = supervisor:which_children(node_sup), mesh_node:crash(P2), timer:sleep(50).
ok
7> is_process_alive(SupPid).
false   %% the 2nd crash: configured intensity of 1 exceeded

No code changed between this run and Milestone 6's original three-crashes-tolerated demonstration — only config/sys.config did, loaded via erl -config config/sys (and, in a real release, bundled automatically as _build/default/rel/mesh/releases/<vsn>/sys.config). The same module, started the same way, now gives up after one restart instead of three, purely because an operator's configuration said one was the limit.

A property test, for the sessionisation-shaped part of this system

EUnit and Common Test both check specific examples you thought to write down. PropEr instead takes a property — a statement that should hold for every input in some class — and searches for a counterexample, generating hundreds of random inputs automatically.

%% test/mesh_node_state_proper.erl
-module(mesh_node_state_proper).
-include_lib("proper/include/proper.hrl").

%% property: draining by any non-negative amount never leaves energy
%% below zero, no matter what amount or starting state is generated
prop_drain_never_goes_negative() ->
    ?FORALL({Id, StartEnergy, DrainAmount},
            {pos_integer(), integer(0, 100), non_neg_integer()},
            begin
                State0 = (mesh_node_state:new(Id))#{energy := StartEnergy},
                #{energy := Final} = mesh_node_state:drain(State0, DrainAmount),
                Final >= 0
            end).
$ rebar3 proper
...
OK: Passed 100 test(s).

?FORALL(Pattern, Generator, Property) is PropEr's core macro: for one hundred automatically-generated combinations of {Id, StartEnergy, DrainAmount}, the property body must hold. This is a genuinely different kind of confidence than the specific examples in Milestone 2's EUnit tests — those prove the function is correct for the cases you thought of; this searches, somewhat adversarially, for a case you did not.

Why are we using this language here?

PropEr generating hundreds of inputs a second and asserting a property against each one is not unique to Erlang — but it is unusually cheap to get right here because so much of this codebase is pure functions over immutable data: no setup, no teardown, no hidden state to reset between the hundred generated cases, because there genuinely is none. The honest cost is real too: a property that fails after ninety-three passing cases hands you a randomly-generated, sometimes large input to debug, and PropEr's shrinking (searching for the smallest input that still fails, visible in this course's own Experiment below) exists specifically because the raw counterexample is often not one a human would choose to read. Property tests also cannot replace example-based tests entirely — they are excellent at "does this invariant hold for a wide space of inputs" and comparatively bad at "does this exact, specific scenario from a bug report behave correctly," which is exactly what Milestone 2's EUnit tests are for instead.

A typical language vs. Erlang

Property-based testing exists elsewhere — Go's own testing/quick and fuzzing support, used in this curriculum's own Go course, are the same idea. What is specifically Erlang-flavoured here is what tends to get tested this way: because so much of this codebase is pure functions over immutable data (Milestone 2's entire design), properties about them are unusually easy to state precisely — "energy never goes negative," "a node's id never changes across a drain," "reversing twice returns the original list" — with no hidden mutable state anywhere to complicate what "the same input" even means between two calls.

Exercise 12
  1. Write a second property: draining and then recharging by the same amount returns a node to its original energy, except where crossing zero or the 100 cap makes that impossible — state the exception precisely as part of the property, rather than avoiding it by only generating inputs that cannot trigger it.
  2. Add a Common Test suite (rebar3 ct) with one test that starts a real node_sup, starts three nodes under it, kills one, and asserts the supervisor's child count returns to three — the integration-shaped test EUnit's per-function focus is awkward for, and Common Test's setup/teardown-per-suite model fits naturally.
Solution 12 — open after trying
prop_drain_recharge_roundtrip() ->
    ?FORALL({StartEnergy, Amount},
            {integer(1, 100), integer(1, 100)},
            begin
                State0 = (mesh_node_state:new(1))#{energy := StartEnergy},
                State1 = mesh_node_state:drain(State0, Amount),
                State2 = mesh_node_state:recharge(State1, Amount),
                #{energy := Final} = State2,
                Expected = min(100, max(0, StartEnergy - Amount) + Amount),
                Final =:= min(100, Expected)
            end).

Stating the boundary explicitly in the property (max(0, ...) for the drain floor, min(100, ...) for the recharge ceiling) rather than constraining the generator to avoid it is the more valuable version of the test: it is the version that would have actually caught a bug in how the floor or ceiling was implemented, because it exercises exactly the inputs most likely to trigger one.

Experiment

Weaken prop_drain_never_goes_negative/0 from Final >= 0 to the strictly positive Final > 0 and rerun rebar3 proper. Predict what fails before running it.

$ rebar3 proper -m mesh_node_state_proper
Testing mesh_node_state_proper:prop_drain_never_goes_negative()
................!
Failed: After 17 test(s).
{2,33,43}

Shrinking ...(3 time(s))
{1,0,0}
0/1 properties passed, 1 failed

PropEr finds a failing case within 17 random attempts, then shrinks it — searching for a smaller input that still fails — down to {Id, StartEnergy, DrainAmount} = {1, 0, 0}: a node that starts at exactly zero energy, drained by exactly zero. drain(State0, 0) leaves Final at exactly 0, which satisfies the original, correct >= 0 but not the deliberately-broken > 0. This is what the "why" box above means by a property test finding the case you did not think to write down: nobody sat down and decided to test "drain by zero starting from zero" specifically, and PropEr found it anyway, minimised to the smallest input that demonstrates it.

Common mistakes in Milestones 9–12

  • Tracing a hot function with a wildcard match specification on a busy system, turning a diagnostic session into a self-inflicted incident.
  • Closing the connection after one write in a handler meant to stream, or the inverse — never returning from a handler meant to respond once.
  • Assuming two nodes need more than a matching cookie and reachability to connect — there is no additional handshake protocol to configure.
  • Treating a network partition as something to detect and "fix" rather than a condition to have an explicit, considered policy for.
  • Constraining a property's generator to avoid the exact edge case worth testing, rather than stating the edge case's correct behaviour as part of the property itself.

Repository state after Milestone 12

mesh/
├── rebar.config
├── src/
│   ├── mesh_app.erl, mesh_sup.erl
│   ├── mesh_id.erl, mesh_node_state.erl
│   ├── mesh_node.erl, node_sup.erl
│   ├── mesh_registry.erl, mesh_chaos.erl
│   ├── mesh_metrics.erl                     Milestone 9
│   └── mesh_dashboard.erl                    Milestone 10
└── test/                                       EUnit, Common Test, PropEr — 12 files
$ rebar3 do eunit, ct, proper
All suites passed.
$ rebar3 release
===> Release successfully assembled: _build/default/rel/mesh
$ git commit -am "milestones 9-12: observability, dashboard, real distribution, a release"

Instalment 19 of the five-course curriculum. Next, and last for Erlang: the advanced phase, a final challenge with acceptance criteria and a withheld solution, the full knowledge check, the README and GitHub description, portfolio notes and interview questions.

Continue