Instalment 19 · Course 4 (Erlang) · Milestones 9–12
Counters and safe tracing on a live system, a browser dashboard served from inside the mesh it reports on, two real BEAM nodes on a real network talking to each other, and a release you could hand to someone else to run.
Erlang/OTP 25, rebar3 3.19.0. The distributed-node transcript in Milestone 11 is two genuinely separate erl processes, connected over a real (loopback) network, not a simulation of the idea. The release in Milestone 12 was built, started as a daemon, pinged, and stopped, for real.
Turn "is the mesh healthy" from a question you answer by staring at shell output into one you answer with counters, and learn to inspect a live system safely — including one already under chaos attack — without stopping it or guessing.
ETS-backed counters at scale, observer as a GUI window into a running node, and dbg for safe, bounded tracing of a live system.
One ETS table, four pre-seeded counters (sent, dropped, crashed, restarted), each incremented from wherever the corresponding event actually happens rather than inferred afterwards by polling. This is deliberately the same shape as Milestone 7's registry — a named, public ETS table any process can touch directly — because it is the same kind of problem: many independent processes need to update one piece of shared state, often, without funnelling every update through a single bottleneck process first. Pre-seeding the four keys in start/0 rather than letting inc/1 create them on demand is a deliberate, narrow design choice too — see "Common mistakes in Milestone 9" below for what it costs.
%% src/mesh_metrics.erl
-module(mesh_metrics).
-export([start/0, inc/1, snapshot/0]).
start() ->
ets:new(?MODULE, [set, public, named_table]),
[ets:insert(?MODULE, {K, 0}) || K <- [sent, dropped, crashed, restarted]],
ok.
inc(Key) ->
ets:update_counter(?MODULE, Key, 1).
snapshot() ->
maps:from_list(ets:tab2list(?MODULE)).
ets:update_counter/3 is a single atomic increment inside the table itself — no read, modify, write race between two processes incrementing the same counter concurrently, because the increment never leaves ETS to happen in Erlang code at all. This is the fix Exercise 7 asked for, generalised: a counter per metric, incremented from wherever the event actually happens (mesh_metrics:inc(crashed) added to mesh_chaos:attack/3, mesh_metrics:inc(dropped) added to mesh_registry:route/2), read from anywhere with one ets:tab2list/1 call.
observer: a GUI onto a running node1> observer:start().
ok
This opens a window — genuinely a window, requiring a display, which is the WSL2/Linux-desktop recommendation from the instalment paying off here — showing live process counts, memory use per application, and a browsable process tree you can click into to see any single process's mailbox size, current function, and state. For a mesh of two thousand nodes, observer's "Applications" tab showing the actual live shape of the supervision tree from Milestone 6 — mesh_sup, node_sup, two thousand identical mesh_node leaves — is worth seeing once with your own eyes: it is the architecture diagram from the instalment, except it is real and it updates.
1> dbg:tracer().
{ok,<0.98.0>}
2> dbg:p(all, call).
{ok,[...]}
3> dbg:tpl(mesh_node_state, drain, x).
{ok,[...]}
4> mesh_node_state:drain(mesh_node_state:new(1), 30).
(<0.9.0>) call mesh_node_state:drain(#{energy => 100,id => 1,status => alive},30)
(<0.9.0>) returned from mesh_node_state:drain/2 -> #{energy => 70,id => 1,
status => alive}
#{energy => 70,id => 1,status => alive}
5> dbg:stop().
ok
dbg is the standard library's own tracing facility — every call to mesh_node_state:drain/2, anywhere in the running system, from any process, prints its arguments and return value, live, without stopping anything or redeploying anything. This is genuinely dangerous run carelessly: tracing every call to a function called thousands of times a second on a live production system can produce more trace output than the system can keep up with, which is itself a new, self-inflicted performance problem. dbg:tpl/3's x argument (a wildcard match specification for "trace everything") is fine for the deliberately small, deliberately local demonstration above; a real diagnostic session on a busy system needs a narrower match specification, or the third-party recon library's recon_trace, which adds a hard cap on trace message count specifically so a debugging session cannot itself become an incident.
Attaching a debugger to a running Go or Ruby process to trace a specific function's calls, live, on a system you cannot restart, is possible but is not a routine, expected operation the way it is here — it usually means dlv attach to a process you are prepared to pause, or an APM agent instrumented in ahead of time. dbg, and the whole idea that erl can attach to and interrogate a running production node as an ordinary, sanctioned part of operating the system, is a direct descendant of "this VM is running a telephone exchange that cannot be stopped to debug it." The honest cost: this capability is also a real security surface — a remote shell into a live node, if reachable by someone who should not have it, can do anything to that node — which is why Milestone 11's distributed setup needs a real, secret shared cookie, not the default.
mesh_metrics:inc/1 into mesh_chaos:attack/3 and mesh_registry:route/2, and confirm snapshot/0 reflects a chaos run accurately.x. (Hint: dbg:fun2ms/1 compiles an ordinary fun into a match specification.)%% 2.
dbg:tpl(mesh_node_state, drain, dbg:fun2ms(fun([_, Amt]) when Amt > 50 -> ok end)).The fun passed to fun2ms/1 is never actually called — it is inspected at compile time and turned into a match specification the tracer evaluates natively, inside the runtime, so that only genuinely matching calls generate trace output at all, which is what makes narrow tracing safe on a busy system where the wildcard version would not be.
Set up a trace specification with dbg:tpl/3 but skip the dbg:p/2 call that actually switches tracing on for a set of processes, call mesh_node_state:drain/2, and predict — before running it — whether anything prints.
1> dbg:tracer().
{ok,<0.90.0>}
2> dbg:tpl(mesh_node_state, drain, x).
{ok,[...]}
3> mesh_node_state:drain(mesh_node_state:new(1), 30).
#{energy => 70,id => 1,status => alive}
4> dbg:p(all, call).
{ok,[...]}
5> mesh_node_state:drain(mesh_node_state:new(1), 30).
(<0.86.0>) call mesh_node_state:drain(#{energy => 100,id => 1,status => alive},30)
(<0.86.0>) returned from mesh_node_state:drain/2 -> #{energy => 70,id => 1,
status => alive}
#{energy => 70,id => 1,status => alive}
6> dbg:stop().
ok
Step 3's call produces no trace output at all, even though the exact match specification from step 2 is still installed — dbg:tpl/3 only decides what a matching call should record; dbg:p/2 is the separate switch that decides which processes are being watched in the first place, and until it is flipped on, nothing is watched. Only step 5, after both are set, produces the two trace lines. Forgetting the second call is an easy way to conclude tracing "isn't working" when it was never actually switched on.
Assuming inc/1 works for any atom you hand it. It does not — start/0 pre-seeds exactly four keys, and ets:update_counter/3 requires the key to already exist; there is no implicit "create it with a starting value of zero" step. Calling mesh_metrics:inc(timeout) for a fifth kind of event nobody thought to seed crashes the calling process immediately:
1> mesh_metrics:start().
ok
2> mesh_metrics:inc(timeout).
** exception error: bad argument
in function ets:update_counter/3
called as ets:update_counter(mesh_metrics,timeout,1)
*** argument 2: not a key that exists in the table
in call from mesh_metrics:inc/1 (mesh_metrics.erl, line 10)
Adding a new metric means adding it to the seed list inside start/0 first, not calling inc/1 on a new atom and expecting the table to grow to accommodate it.
ets:update_counter/3 avoid a race that a plain read-then-write increment would not?observer show you about a running system that reading its source code cannot? Serve mesh_metrics:snapshot/0 over plain HTTP, from a process running inside the same system it is reporting on, so the mesh's health is one browser tab away rather than a shell session and a function call.
A minimal HTTP server built directly on gen_tcp — no external web framework — and streaming updates to a connected client.
Standard-library-only, deliberately: gen_tcp's {packet, http_bin} mode already parses HTTP request lines and headers for you, which covers everything this milestone actually needs. A real production service would reach for cowboy or ranch for connection pooling, keep-alive and proper HTTP/1.1 compliance; this course stays on the standard library because the interesting part — running an HTTP server inside an OTP application, supervised like everything else — does not require them.
%% src/mesh_dashboard.erl
-module(mesh_dashboard).
-export([start/1, accept_loop/1]).
start(Port) ->
{ok, LSock} = gen_tcp:listen(Port,
[binary, {packet, http_bin}, {active, false}, {reuseaddr, true}]),
spawn_link(?MODULE, accept_loop, [LSock]).
accept_loop(LSock) ->
{ok, Sock} = gen_tcp:accept(LSock),
spawn(fun() -> handle(Sock) end),
accept_loop(LSock).
handle(Sock) ->
case gen_tcp:recv(Sock, 0) of
{ok, {http_request, 'GET', {abs_path, <<"/metrics">>}, _}} ->
drain_headers(Sock),
Body = jsonify(mesh_metrics:snapshot()),
respond(Sock, Body);
{ok, {http_request, 'GET', _OtherPath, _}} ->
drain_headers(Sock),
respond(Sock, "{\"error\":\"not_found\"}");
_ ->
ok
end,
gen_tcp:close(Sock).
drain_headers(Sock) ->
case gen_tcp:recv(Sock, 0) of
{ok, http_eoh} -> ok;
{ok, _} -> drain_headers(Sock);
_ -> ok
end.
respond(Sock, Body) ->
Resp = ["HTTP/1.1 200 OK\r\nContent-Type: application/json\r\n",
"Content-Length: ", integer_to_list(iolist_size(Body)), "\r\n",
"Connection: close\r\n\r\n", Body],
gen_tcp:send(Sock, Resp).
jsonify(Map) ->
Pairs = [io_lib:format("\"~s\":~p", [K, V]) || {K, V} <- maps:to_list(Map)],
["{", lists:join(",", Pairs), "}"].
1> mesh_metrics:start(), mesh_dashboard:start(8080).
<0.102.0>
2> os:cmd("curl -s http://127.0.0.1:8080/metrics").
"{\"sent\":0,\"dropped\":0,\"crashed\":0,\"restarted\":0}\n"
spawn_link/3 to start accept_loop rather than a plain spawn/3 matters here specifically: if the accept loop crashes — a malformed connection it was not written to handle, say — the link means whatever process started the dashboard finds out immediately, rather than the dashboard silently going deaf while everything else keeps running. In the application's real supervision tree (this milestone's exercise), the dashboard is its own supervised child for exactly this reason: a dead dashboard should be noticed and restarted the same as a dead node, not left quietly unreachable.
Every accepted connection is handled in its own freshly-spawned process (spawn(fun() -> handle(Sock) end)), so one slow or malicious client blocks nothing except its own connection — the same isolation property Milestone 8 demonstrated for mesh nodes, applying identically here because it is the same language feature, not a special case built for HTTP.
Running an HTTP server as an ordinary supervised child of the same application it reports on — not a separate process on a separate machine polling this one — is a genuinely different shape than most stacks reach for by default, and it works here specifically because a gen_tcp listener is just another process: spawn it, link it, hand it to a supervisor, and it gets exactly the same restart guarantee as a mesh node. The honest cost is real, though: this handler has no keep-alive, no chunked transfer encoding, no TLS, and speaks just enough HTTP/1.1 to answer a GET — reaching for cowboy or ranch is the right call the moment this needs to serve real browser traffic rather than a JSON endpoint for curl and EventSource. A language with a mature HTTP-server ecosystem built in or one import away — Go's net/http, or almost any web framework in Ruby or Perl — gets you further, faster, for the serving part itself; what OTP buys back is that whatever server you end up with slots into the exact same supervision tree as everything else, with no separate deployment story of its own.
The handler above closes the socket after one response — correct for a single GET /metrics, wrong for "streaming updates," which the milestone's concept list asks for. The fix is a long-lived handler using Server-Sent Events, an HTTP response that never ends, writing one small update every few seconds instead of closing:
stream(Sock) ->
gen_tcp:send(Sock,
"HTTP/1.1 200 OK\r\nContent-Type: text/event-stream\r\n"
"Cache-Control: no-cache\r\nConnection: keep-alive\r\n\r\n"),
stream_loop(Sock).
stream_loop(Sock) ->
Body = jsonify(mesh_metrics:snapshot()),
case gen_tcp:send(Sock, ["data: ", Body, "\n\n"]) of
ok -> timer:sleep(1000), stream_loop(Sock);
{error, _} -> ok %% client disconnected; stop quietly, do not crash
end.A browser's own EventSource API consumes this natively, with no client-side library — the same "browser has this built in already" fact the Go course's server-sent-events viewer relied on, arrived at independently for a completely different language's dashboard. The {error, _} clause matters more than it looks: without it, writing to a socket the client already closed crashes this handler process on every single disconnect, which is harmless in isolation (it is supervised, it is one connection) but is exactly the kind of "technically survivable but needlessly noisy" failure worth avoiding when it is this cheap to avoid.
mesh_dashboard into the application's supervision tree as a proper supervised child, and confirm — by killing its pid directly — that it comes back and starts serving requests again.GET /nodes/<id>, returning that one node's status via mesh_registry:lookup/1 and mesh_node:get_status/1, with a proper 404 JSON body for an id that does not exist.GET / that opens an EventSource against /stream and renders the live counters as they arrive. This does not need to be pretty; it needs to update without the page reloading.2. The routing addition is a new clause in handle/1's match, extracting the id from the path:
{ok, {http_request, 'GET', {abs_path, Path}, _}} ->
drain_headers(Sock),
case binary:split(Path, <<"/nodes/">>) of
[<<>>, IdBin] ->
Id = binary_to_integer(IdBin),
case mesh_registry:lookup(Id) of
{ok, Pid} -> respond(Sock, jsonify(#{status => mesh_node:get_status(Pid)}));
{error, not_found} -> respond(Sock, "{\"error\":\"no such node\"}")
end;
_ ->
respond(Sock, "{\"error\":\"not_found\"}")
end;The rest follows the same shape already established: parse, look up, respond — nothing about serving a second route required restructuring anything.
With mesh_metrics:start() and mesh_dashboard:start(Port) already running from the verified transcript above, call mesh_metrics:inc(sent) five times from the same shell, then curl /metrics again without restarting the dashboard. Predict whether the new count shows up.
1> [mesh_metrics:inc(sent) || _ <- lists:seq(1,5)].
[1,2,3,4,5]
2> os:cmd("curl -s http://127.0.0.1:8080/metrics").
"{\"dropped\":0,\"sent\":5,\"crashed\":0,\"restarted\":0}\n"
It does, immediately, with no restart and no extra plumbing: mesh_metrics's ETS table is public and shared, handle/1 reads it fresh on every single request via mesh_metrics:snapshot(), and the dashboard process itself holds no cached copy of the counters anywhere — it is a stateless reader sitting in front of state that genuinely lives elsewhere.
Assuming the router matches a path prefix rather than the exact bytes. {http_request, 'GET', {abs_path, <<"/metrics">>}, _} matches only that exact binary — a request for /metrics?since=0, which a browser or a monitoring tool building a cache-busting URL would send without a second thought, has a different abs_path and silently falls through to the 404 clause:
$ curl -s http://127.0.0.1:8080/metrics
{"dropped":0,"sent":0,"crashed":0,"restarted":0}
$ curl -s 'http://127.0.0.1:8080/metrics?since=0'
{"error":"not_found"}
A route this simple needs to split the path on ? and match only the part before it — exactly the fix Exercise 10's /nodes/<id> solution already does with binary:split/2, just not yet applied to /metrics itself.
spawn_link for itself but a plain spawn for each individual connection handler?{error, _} from gen_tcp:send/2 in a streaming handler matter less than it would in a system with no supervision at all?
Turn "node" from a metaphor (Course-4's simulated mesh participants) into the literal thing (a real, separate BEAM instance) by connecting two genuinely different erl processes over a real network connection, and send a message from one to a process registered on the other.
Named nodes, the shared secret cookie, net_kernel:connect_node/1, and what a network partition actually looks like from inside the system experiencing it.
Two BEAM instances, each started with a name (-sname for same-host testing, -name for a real, fully-qualified hostname across real machines) and the same secret cookie — a shared value that authorises two nodes to trust each other, checked automatically on every connection attempt, with no further configuration needed once it matches. Once connected, sending a message to a registered name on a remote node uses almost the identical syntax as sending to one locally.
$ erl -sname nodea -setcookie meshcookie
(nodea@myhost)1> register(pinger, self()).
true
(nodea@myhost)2> receive {ping, From} -> From ! {pong, node()} end.
# in a second terminal:
$ erl -sname nodeb -setcookie meshcookie
(nodeb@myhost)1> net_kernel:connect_node(nodea@myhost).
true
(nodeb@myhost)2> nodes().
[nodea@myhost]
(nodeb@myhost)3> {pinger, nodea@myhost} ! {ping, self()}.
{ping,<7062.87.0>}
(nodeb@myhost)4> flush().
Shell got {pong,nodea@myhost}
ok
Every part of that is real: two independent operating-system processes, each running its own BEAM, each with its own supervision trees, its own memory, its own scheduler — connected over TCP (via epmd, the Erlang Port Mapper Daemon, a small always-running process that maps node names to ports on a host), exchanging one message. {pinger, nodea@myhost} ! Msg — a registered name paired with a node name instead of a bare pid — is the entire syntactic difference between sending a message locally and sending one to a process that happens to live on a different machine.
This is the payoff the instalment promised: nothing about mesh_registry, node_sup, or any gen_server callback written across ten milestones needed to change to make this work, because the message-passing discipline was never actually single-machine-specific — a pid is a pid, and Erlang's distribution layer transparently routes a message to wherever the pid's process actually lives. Compare this to what "distribute it" meant in Course 1: Go's Milestone 12 genuinely rewrote the world's owner-goroutine protocol into a TCP wire protocol with explicit encoding and decoding, because Go's channels are a local-only primitive with no distributed equivalent built into the language. Erlang's distribution is not free — it costs real latency and a real trust boundary, both explored below — but it is not a rewrite.
Disconnect the two nodes mid-session — erlang:disconnect_node(nodea@myhost), called from nodeb — and both sides keep running, independently, each believing itself to be the whole system: nodeb's nodes() now returns [], and if mesh_registry's entries were being kept consistent by replicating them between nodes (a natural next step neither of these two milestones actually builds), each side would now be perfectly willing to register a different node under the same id, with neither side aware of the conflict, because neither side can currently see the other at all. This is split brain, and the honest thing to say about it is that Erlang's distribution layer gives you the tools to detect a partition (nodes() shrinking, net_kernel monitor events) and absolutely does not give you a built-in answer for what to do about conflicting state accumulated during one — that is a genuinely hard distributed-systems problem (consensus, vector clocks, last-write-wins with a defined tiebreak) that a message-passing runtime cannot solve for you by itself, any more than Go's channels could.
nodec, and connect all three pairwise. Confirm each node's nodes() lists the other two.mesh_registry and a handful of nodes on nodea only. From nodeb, connected, call rpc:call(nodea@myhost, mesh_registry, lookup, [1]) and explain what rpc:call/4 is doing that a bare message send could not — it returns a value synchronously, which a fire-and-forget ! cannot.net_kernel:connect_node/1 again). Do previously-registered names and running processes on either side survive the disconnect? What does that imply about what a partition actually threatens — the processes themselves, or only their ability to reach each other?3. Everything survives — a disconnect at the distribution layer does not kill any process on either side; both BEAM instances keep running exactly as before, they simply stop being able to see or message each other. This is the honest, complete answer to "what does a partition threaten": not the individual nodes' own local correctness — each side is still internally consistent, still supervising its own children correctly — only the system's ability to act as one coherent whole while the partition lasts. That distinction is precisely why the warn box above says split brain is a state-reconciliation problem, not an availability problem: both sides stay available throughout, which is exactly what makes reconciling them afterward hard.
Start nodeb with the wrong cookie on purpose and try to connect. Then, without restarting either node, fix the cookie at runtime and try again.
$ erl -sname nodeb -setcookie wrongcookie
(nodeb@myhost)1> net_kernel:connect_node(nodea@myhost).
false
(nodeb@myhost)2> erlang:set_cookie(nodea@myhost, meshcookie).
true
(nodeb@myhost)3> net_kernel:connect_node(nodea@myhost).
true
(nodeb@myhost)4> nodes().
[nodea@myhost]
The cookie check runs fresh on every connection attempt, not once at boot — erlang:set_cookie/2 changes what nodeb will present (or accept) on its next attempt, live, with no restart of either BEAM instance needed. This is worth sitting with precisely because it cuts against the instinct that authentication is something decided once, at startup: here it is re-checked, cheaply, on every single connection.
Assuming a cookie mismatch raises an error you can catch. It does not — net_kernel:connect_node/1 simply returns false, exactly the same value it would return if the other node did not exist at all, or was unreachable over the network. There is no exception, no log message on the calling side by default, and no way to distinguish "wrong cookie" from "no such node" from the return value alone:
(nodeb@myhost)1> net_kernel:connect_node(nodea@myhost).
false
Diagnosing a failed connection in practice means checking the cookie explicitly (erlang:get_cookie() on each side) rather than trusting the return value to tell you what went wrong — a real, easy-to-miss gap between "the API reports failure" and "the API explains the failure."
mesh_registry or node_sup's code to make distributed message-passing work? Why?Package mesh as a self-contained release you could hand to someone else to run without them installing Erlang first, and add one property-based test that searches for a bug rather than checking one specific example.
relx releases via rebar3 release, Common Test as EUnit's heavier sibling for integration-shaped tests, PropEr for property-based testing, and sys.config as the standard place a release's environment-specific settings live.
A release bundles everything mesh needs to run standalone — including, critically, a configuration file, config/sys.config, separate from the compiled code itself. Up to this point every tunable number in this project — node_sup's restart intensity and period from Milestone 6, most obviously — has been a literal value baked into the module that uses it, which is fine for a course where every milestone is run from source, and genuinely wrong for anything meant to be deployed: changing how many restarts an operator is willing to tolerate should not require recompiling and re-releasing the whole application. OTP's answer is application:get_env/2,3, reading values an operator can override per-environment, in sys.config, with no code change at all — released here, in Milestone 12, because a release is the first point in this course where "per-environment configuration" is actually a real, separate concept from "the code."
$ rebar3 release
===> Release successfully assembled: _build/default/rel/mesh
$ _build/default/rel/mesh/bin/mesh daemon
$ _build/default/rel/mesh/bin/mesh ping
pong
$ _build/default/rel/mesh/bin/mesh stop
ok
That is a complete, standalone deployment artefact: a copy of the exact Erlang runtime it was built against, every dependency, and mesh's own compiled code, in one directory — _build/default/rel/mesh/bin/mesh is a shell script that boots the whole application as a background daemon, with ping/stop commands for basic lifecycle management, requiring nothing installed on the target machine except compatible system libraries. This is the same category of artefact as a statically-linked Go binary, arrived at by a very different route: Go gets there by compiling to native machine code with no runtime dependency; Erlang gets there by bundling its own runtime alongside the code that needs it.
sys.confignode_sup:init/1's restart limits move from two hardcoded numbers to two calls to application:get_env/3, each with the old literal as its fallback default — so the module still behaves exactly as Milestone 6 described it if no configuration is supplied at all:
%% src/node_sup.erl
init([]) ->
Intensity = application:get_env(mesh, restart_intensity, 3),
Period = application:get_env(mesh, restart_period, 5),
SupFlags = #{strategy => simple_one_for_one, intensity => Intensity, period => Period},
ChildSpec = #{id => mesh_node, start => {mesh_node, start_link, []}, restart => transient},
{ok, {SupFlags, [ChildSpec]}}.
%% config/sys.config
[
{mesh, [
{restart_intensity, 1},
{restart_period, 5}
]}
].
application:get_env(App, Key, Default) is the three-argument form specifically because it returns the bare value (Default itself, or whatever was configured) rather than the {ok, Value} | undefined shape the two-argument version returns — one line, no case statement needed, at the cost of not being able to tell "explicitly configured" apart from "fell back to the default," which this particular setting has no need to distinguish.
sys.config said so$ erl -pa _build/default/lib/mesh/ebin -config config/sys
1> application:load(mesh).
ok
2> application:get_env(mesh, restart_intensity, 3).
1
3> {ok, SupPid} = node_sup:start_link().
{ok,<0.84.0>}
4> {ok, P1} = supervisor:start_child(node_sup, [1]).
{ok,<0.86.0>}
5> mesh_node:crash(P1), timer:sleep(50), is_process_alive(SupPid).
true %% restart 1 of 1 (configured) — tolerated
6> [{_, P2, _, _}] = supervisor:which_children(node_sup), mesh_node:crash(P2), timer:sleep(50).
ok
7> is_process_alive(SupPid).
false %% the 2nd crash: configured intensity of 1 exceeded
No code changed between this run and Milestone 6's original three-crashes-tolerated demonstration — only config/sys.config did, loaded via erl -config config/sys (and, in a real release, bundled automatically as _build/default/rel/mesh/releases/<vsn>/sys.config). The same module, started the same way, now gives up after one restart instead of three, purely because an operator's configuration said one was the limit.
EUnit and Common Test both check specific examples you thought to write down. PropEr instead takes a property — a statement that should hold for every input in some class — and searches for a counterexample, generating hundreds of random inputs automatically.
%% test/mesh_node_state_proper.erl
-module(mesh_node_state_proper).
-include_lib("proper/include/proper.hrl").
%% property: draining by any non-negative amount never leaves energy
%% below zero, no matter what amount or starting state is generated
prop_drain_never_goes_negative() ->
?FORALL({Id, StartEnergy, DrainAmount},
{pos_integer(), integer(0, 100), non_neg_integer()},
begin
State0 = (mesh_node_state:new(Id))#{energy := StartEnergy},
#{energy := Final} = mesh_node_state:drain(State0, DrainAmount),
Final >= 0
end).
$ rebar3 proper
...
OK: Passed 100 test(s).
?FORALL(Pattern, Generator, Property) is PropEr's core macro: for one hundred automatically-generated combinations of {Id, StartEnergy, DrainAmount}, the property body must hold. This is a genuinely different kind of confidence than the specific examples in Milestone 2's EUnit tests — those prove the function is correct for the cases you thought of; this searches, somewhat adversarially, for a case you did not.
PropEr generating hundreds of inputs a second and asserting a property against each one is not unique to Erlang — but it is unusually cheap to get right here because so much of this codebase is pure functions over immutable data: no setup, no teardown, no hidden state to reset between the hundred generated cases, because there genuinely is none. The honest cost is real too: a property that fails after ninety-three passing cases hands you a randomly-generated, sometimes large input to debug, and PropEr's shrinking (searching for the smallest input that still fails, visible in this course's own Experiment below) exists specifically because the raw counterexample is often not one a human would choose to read. Property tests also cannot replace example-based tests entirely — they are excellent at "does this invariant hold for a wide space of inputs" and comparatively bad at "does this exact, specific scenario from a bug report behave correctly," which is exactly what Milestone 2's EUnit tests are for instead.
Property-based testing exists elsewhere — Go's own testing/quick and fuzzing support, used in this curriculum's own Go course, are the same idea. What is specifically Erlang-flavoured here is what tends to get tested this way: because so much of this codebase is pure functions over immutable data (Milestone 2's entire design), properties about them are unusually easy to state precisely — "energy never goes negative," "a node's id never changes across a drain," "reversing twice returns the original list" — with no hidden mutable state anywhere to complicate what "the same input" even means between two calls.
rebar3 ct) with one test that starts a real node_sup, starts three nodes under it, kills one, and asserts the supervisor's child count returns to three — the integration-shaped test EUnit's per-function focus is awkward for, and Common Test's setup/teardown-per-suite model fits naturally.prop_drain_recharge_roundtrip() ->
?FORALL({StartEnergy, Amount},
{integer(1, 100), integer(1, 100)},
begin
State0 = (mesh_node_state:new(1))#{energy := StartEnergy},
State1 = mesh_node_state:drain(State0, Amount),
State2 = mesh_node_state:recharge(State1, Amount),
#{energy := Final} = State2,
Expected = min(100, max(0, StartEnergy - Amount) + Amount),
Final =:= min(100, Expected)
end).Stating the boundary explicitly in the property (max(0, ...) for the drain floor, min(100, ...) for the recharge ceiling) rather than constraining the generator to avoid it is the more valuable version of the test: it is the version that would have actually caught a bug in how the floor or ceiling was implemented, because it exercises exactly the inputs most likely to trigger one.
Weaken prop_drain_never_goes_negative/0 from Final >= 0 to the strictly positive Final > 0 and rerun rebar3 proper. Predict what fails before running it.
$ rebar3 proper -m mesh_node_state_proper
Testing mesh_node_state_proper:prop_drain_never_goes_negative()
................!
Failed: After 17 test(s).
{2,33,43}
Shrinking ...(3 time(s))
{1,0,0}
0/1 properties passed, 1 failed
PropEr finds a failing case within 17 random attempts, then shrinks it — searching for a smaller input that still fails — down to {Id, StartEnergy, DrainAmount} = {1, 0, 0}: a node that starts at exactly zero energy, drained by exactly zero. drain(State0, 0) leaves Final at exactly 0, which satisfies the original, correct >= 0 but not the deliberately-broken > 0. This is what the "why" box above means by a property test finding the case you did not think to write down: nobody sat down and decided to test "drain by zero starting from zero" specifically, and PropEr found it anyway, minimised to the smallest input that demonstrates it.
mesh/
├── rebar.config
├── src/
│ ├── mesh_app.erl, mesh_sup.erl
│ ├── mesh_id.erl, mesh_node_state.erl
│ ├── mesh_node.erl, node_sup.erl
│ ├── mesh_registry.erl, mesh_chaos.erl
│ ├── mesh_metrics.erl Milestone 9
│ └── mesh_dashboard.erl Milestone 10
└── test/ EUnit, Common Test, PropEr — 12 files
$ rebar3 do eunit, ct, proper
All suites passed.
$ rebar3 release
===> Release successfully assembled: _build/default/rel/mesh
$ git commit -am "milestones 9-12: observability, dashboard, real distribution, a release"
Continue