Instalment 3 · Course 1 (Go) · Milestones 5–8
The mutex disappears. Pheromone trails make the colony measurably smarter. Shutdown becomes something you can prove. Then we start breaking things on purpose and watch the supervisor put them back.
All code was compiled, vetted, race-tested and executed before it reached the page, and the numbers quoted are measurements from those runs, not estimates. Where my single-core sandbox distorts a result, I say so.
Delete the mutex. One goroutine owns the world; ants send it messages and get answers back. Along the way the Behaviour interface changes shape in a way that turns out to matter enormously later.
Ownership as a design principle, request/response over channels, per-client reply channels, buffered channels as leak prevention, the "ask the owner" pattern for reading state, and value types as messages.
The failure of Milestone 4 was structural: many goroutines reaching into one mutable object. There are only three ways out of that, and it is worth seeing all three before picking one.
| Approach | Idea | Why not (here) |
|---|---|---|
| Finer locks | One mutex per cell or region | 262,144 locks, a mandatory global lock ordering, and a deadlock the first time a feature needs two regions at once |
| Lock-free | Atomics and compare-and-swap on cells | Works for counters, becomes research-grade difficulty the moment an operation touches two cells |
| Single owner | One goroutine holds the data; everyone else sends messages | This one. Serialises writes by construction, needs no lock discipline, and survives becoming a network protocol |
The shape we are moving to:
ant 1 ──┐ ┌──────────────────────┐
ant 2 ──┤ │ serve() goroutine │
ant 3 ──┼──► reqs chan request ──► │ │
... │ │ OWNS: world, food, │
ant N ──┘ │ pheromone, counters │
└──────────┬───────────┘
each ant has its own │
reply channel (buffered 1) ◄────────────────────┘
response sent to r.reply
Three decisions inside that picture deserve argument.
The reply channel belongs to the ant, not to the request. The obvious version allocates make(chan response) per request, which is one allocation and one garbage-collected object per ant per tick: at 50,000 ants that is millions of channels per second. Instead each ant creates one buffered channel at birth and reuses it forever. The buffer matters too, and not for speed: if the ant gives up waiting (a timeout, a cancelled context) and the owner later sends a reply, an unbuffered channel would block the owner permanently, wedging the entire simulation. A buffer of one means the owner always completes its send and moves on.
Messages carry values, never pointers. If a request contained *Ant, then the owner and the ant goroutine would both hold a pointer to the same struct, and we would be exactly back where Milestone 4 started, only with extra steps. Sending pos world.Position and carrying bool by value means the message is the shared state, and it is copied. This costs a few bytes and buys the property that the whole protocol could be serialised and sent over a socket without redesign. Milestone 12 does precisely that.
The ant perceives rather than reads. Decide currently takes *world.World and reads whatever it likes, which cannot work when the world lives in another goroutine. So the interface changes:
before: Decide(a *Ant, w *world.World, rng *rand.Rand) Action
after: Decide(a *Ant, s Sense, rng *rand.Rand) Action
A Sense is a small pointer-free snapshot of what one ant can perceive from where it stands. This is a better model of an ant anyway (real ants have no global map) and it makes every behaviour trivially testable with a literal. Concurrency pushed us toward a design we would have wanted regardless, which happens more often than you would expect.
package sim
import "github.com/yourname/antfarm/internal/world"
// Sense is everything one ant can perceive on one tick: its own position,
// the nest it remembers, and what is in the eight cells around it. It is a
// pointer-free value, so handing it to another goroutine is safe by
// construction.
type Sense struct {
Pos world.Position
Nest world.Position
Food int // food on the ant's own cell
Adjacent [8]int // food in each of the eight neighbours
Pher [8]float64 // pheromone in each of the eight neighbours
}
// senseAt builds a Sense by reading the world. Only code that owns the
// world may call it.
func senseAt(w *world.World, pos world.Position) Sense {
s := Sense{Pos: pos, Nest: w.Nest, Food: w.Food.At(pos)}
for i, d := range directions[:8] {
n := pos.Add(d)
s.Adjacent[i] = w.Food.At(n)
s.Pher[i] = w.Pher.At(n)
}
return s
}
[8]int is an array, not a slice, and that distinction is the entire point: an array is stored inline in the struct, so copying a Sense copies the data. A []int would copy a pointer, and two goroutines would share the elements. When designing a message type, arrays are safe and slices are not. (The Pher field is for Milestone 6; it costs nothing to add now.)
type reqKind int
const (
reqSense reqKind = iota
reqAct
reqStats
reqEvaporate
reqLost
)
// request is a message from anyone to the world owner. Every field is a
// value: no pointers cross the channel, so nothing is shared.
type request struct {
kind reqKind
id int
pos world.Position
carrying bool
act Action
reply chan response
}
// response is the owner's answer, also pointer-free.
type response struct {
sense Sense
pos world.Position
ok bool
err error
stats Stats
}
One request type with a kind tag rather than five separate channels. Both designs are idiomatic; the tagged union is easier to extend and keeps ordering between different request kinds, which will matter when evaporation competes with ant traffic. The cost is a switch and some unused fields, which is Go's usual price for not having sum types.
// Engine runs the colony as one owner goroutine plus one goroutine per ant.
// No mutex appears anywhere in this file.
type Engine struct {
cfg Config
b Behaviour
reqs chan request
// nest never changes after construction, so every goroutine may read it.
nest world.Position
// Owned exclusively by the serve goroutine. Nothing else may touch
// these, which is why they need no synchronisation at all.
world *world.World
rng *rand.Rand
tick int
delivered int
carrying int
failedPickups int
lost int
// Shared counters, written from many goroutines.
restarts atomic.Int64
dropped atomic.Int64
timeouts atomic.Int64
}
// serve is the only goroutine allowed to read or write the world.
func (e *Engine) serve(ctx context.Context) {
for {
select {
case <-ctx.Done():
return
case r := <-e.reqs:
e.handle(r)
}
}
}
func (e *Engine) handle(r request) {
switch r.kind {
case reqSense:
r.reply <- response{sense: senseAt(e.world, r.pos)}
case reqAct:
r.reply <- e.applyAction(r)
case reqStats:
r.reply <- response{stats: e.statsNow()}
case reqEvaporate:
e.world.Pher.Evaporate(e.cfg.Evaporation)
e.tick++
case reqLost:
e.lost++
e.carrying--
}
}
Read the comments on the struct fields as if they were enforced, because they are the only enforcement there is. Go cannot express "this field may only be touched by that goroutine", so the convention is a comment plus discipline plus go test -race in CI. Grouping the owner-only fields together, with one comment covering the block, makes a violation visible in review.
nest is duplicated out of the world deliberately. It never changes after construction, so it is safe for anyone to read, and having it available outside the owner avoids a message round trip in the supervisor later. Immutable-after-construction is a third category alongside "owned" and "synchronised", and it is worth naming in comments when you use it.
func (e *Engine) applyAction(r request) response {
switch r.act.Kind {
case ActMove:
np := e.world.Clamp(r.pos.Add(r.act.Dir))
if r.carrying {
e.world.Pher.AddAt(np, e.cfg.Deposit, e.cfg.PherMax)
}
return response{pos: np, ok: true}
case ActPickUp:
if err := e.world.TakeFood(r.pos); err != nil {
e.failedPickups++
return response{pos: r.pos, err: err}
}
e.carrying++
return response{pos: r.pos, ok: true}
case ActDrop:
if r.pos != e.world.Nest {
return response{pos: r.pos, err: fmt.Errorf("ant %d drop away from nest: %w", r.id, world.ErrNotCarrying)}
}
e.world.Deliver()
e.carrying--
return response{pos: r.pos, ok: true}
}
return response{pos: r.pos, ok: true}
}
Note what the owner does not do: it never writes to an Ant. It computes the consequence and returns it. The ant updates itself. This keeps a clean rule that is easy to check by reading: the owner writes the world, each ant writes itself, and nothing writes both.
The owner also validates. An ant that asks to drop food away from the nest gets an error rather than a delivery. A behaviour is untrusted input, in the same way a client request is untrusted input to a server, and for the same reason: in Milestone 8 those requests start arriving corrupted.
// antClient is one ant's private connection to the owner: its own reply
// channel and its own random source, touched by no other goroutine.
type antClient struct {
e *Engine
ant *Ant
reply chan response
rng *rand.Rand
}
func (c *antClient) roundTrip(ctx context.Context, r request) (response, error) {
// A reply from a request we gave up on may still be sitting here.
// Throw it away before asking again.
select {
case <-c.reply:
default:
}
r.reply = c.reply
select {
case c.e.reqs <- r:
case <-ctx.Done():
return response{}, ctx.Err()
}
timer := time.NewTimer(c.e.cfg.RequestTimeout)
defer timer.Stop()
select {
case resp := <-c.reply:
return resp, nil
case <-timer.C:
c.e.timeouts.Add(1)
return response{}, ErrTimeout
case <-ctx.Done():
return response{}, ctx.Err()
}
}
This function is twenty lines and every one of them is load-bearing.
select { case <-c.reply: default: } is a non-blocking receive: take a value if one is waiting, otherwise carry on immediately. Without it, a late reply to a timed-out request would be read as the answer to the next request, and the ant would act on stale information. This is a real distributed-systems bug in miniature, and it appears the moment you add timeouts to any request/response protocol.select with ctx.Done(). A plain c.e.reqs <- r blocks when the queue is full, and a blocked send ignores cancellation, so shutdown would hang until the owner drained. Every potentially-blocking channel operation in a long-lived goroutine should be paired with cancellation. time.NewTimer with defer timer.Stop(), not time.After. time.After allocates a timer that stays alive until it fires, even if you stopped caring. In a loop running millions of times per second that is a slow memory leak. This is one of the most common performance defects in real Go services.// runAnt is the whole life of one ant: sense, decide, act, repeat.
func (e *Engine) runAnt(ctx context.Context, a *Ant) {
c := &antClient{
e: e,
ant: a,
reply: make(chan response, 1),
rng: rand.New(rand.NewPCG(e.cfg.Seed, uint64(a.ID)+a.Generation<<32)),
}
for {
if ctx.Err() != nil {
return
}
sensed, err := c.roundTrip(ctx, request{kind: reqSense, id: a.ID, pos: a.Pos})
if err != nil {
if errors.Is(err, ErrTimeout) {
continue // ask again
}
return // context cancelled
}
act := e.b.Decide(a, sensed.sense, c.rng)
done, err := c.roundTrip(ctx, request{
kind: reqAct, id: a.ID, pos: a.Pos, carrying: a.Carrying, act: act,
})
if err != nil {
if errors.Is(err, ErrTimeout) {
continue
}
return
}
c.applyLocally(act, done)
}
}
// applyLocally updates the ant's own state from the owner's answer. The ant
// goroutine is the only writer of these fields.
func (c *antClient) applyLocally(act Action, r response) {
switch act.Kind {
case ActMove:
if r.ok {
c.ant.Pos = r.pos
c.ant.Energy--
}
case ActPickUp:
if r.ok {
c.ant.Carrying = true
}
case ActDrop:
if r.ok {
c.ant.Carrying = false
}
}
}
Sense, decide, act. The decide step is the only part that runs outside the owner, and it is the only part that runs in parallel across ants. Remember that when we discuss performance in a moment.
// Stats asks the owner for a snapshot. Safe to call from any goroutine.
func (e *Engine) Stats(ctx context.Context) (Stats, error) {
reply := make(chan response, 1)
select {
case e.reqs <- request{kind: reqStats, reply: reply}:
case <-ctx.Done():
return Stats{}, ctx.Err()
}
timer := time.NewTimer(e.cfg.RequestTimeout)
defer timer.Stop()
select {
case r := <-reply:
return r.stats, nil
case <-timer.C:
return Stats{}, ErrTimeout
case <-ctx.Done():
return Stats{}, ctx.Err()
}
}
Compare with Milestone 4's solution, where Stats took a lock and we had to invent the statsLocked convention to avoid deadlocking against ourselves. Here the question "what if somebody calls Stats from inside the owner?" cannot arise, because the owner never calls Stats; it calls statsNow, which is an ordinary function on data it owns. Reentrancy stops being a hazard when there is no lock to re-enter.
Here a fresh channel per call is fine, because Stats is called a few times a second, not a few million. Reuse is an optimisation, and optimisations belong where the traffic is.
$ go test -race -run TestEngine -count=1 ./internal/sim/
ok github.com/yourname/antfarm/internal/sim 1.965s
Correct, race-free, and no mutex in the package. Now the honest part.
Every world access in the colony passes through one goroutine, so world operations happen strictly one at a time, exactly as they did behind the single mutex. If you were expecting linear speedup from replacing the lock, you will not get it, and anyone claiming otherwise is describing a different program.
What actually changed:
Decide now runs in parallel across all ants, off the owner's goroutine. In this project that computation is small; in a simulation where agents do real work (pathfinding, learning, inference) it is most of the cost, and this structure parallelises all of it.reqs chan request with a capacity is a visible, measurable, tunable buffer. Under a mutex the queue exists too, inside the runtime, where you cannot see its depth or apply a policy when it fills. Milestone 11 turns that visibility into backpressure.If raw throughput on one machine were the only goal, the right answer is neither locks nor a single owner: it is sharding, several owners each holding a region of the grid, which Milestone 11 builds. The single owner is the design you should reach for first because it is obviously correct, and shard it only when measurement says you must.
Shared state + locks Single owner + messages
──────────────────── ───────────────────────
any goroutine may touch data one goroutine may touch data
correctness = lock discipline correctness = structure
invisible queueing in the runtime an explicit channel you can measure
deadlock is a real risk deadlock needs a request cycle
reentrancy hazards none: no lock to re-enter
cannot express "message lost" one line at the owner
does not survive a network already is a network protocolGo supports both and takes no side. The proverb "share memory by communicating" is advice, not a rule, and the standard library uses mutexes heavily where they fit. What Go provides that most languages do not is that both options are first-class and cheap enough to choose between on the merits.
Add a reqTeleport request that moves an ant to an arbitrary position, and expose it as func (e *Engine) Teleport(ctx context.Context, id int, to world.Position) error so an external caller (the future dashboard) can poke the simulation.
Requirements: reject out-of-bounds targets with a useful error; the ant's own goroutine must end up agreeing with the owner about where it is; no new mutex; and a test that calls Teleport concurrently with a running colony and passes under -race.
The interesting part is the second requirement. The owner does not write to ants, and the ant is asleep in its own loop. Think about who can change a.Pos and how the news reaches them.
The trap is to have Teleport write a.Pos directly. It compiles, and it is a data race against the ant's goroutine, which writes the same field in applyLocally.
The clean answer is a mailbox per ant: the owner records a pending command, and the ant collects it on its next sense round trip. One extra field on response, no new synchronisation:
// in Engine, owner-only state:
// pending map[int]world.Position // ant ID -> forced position
func (e *Engine) handle(r request) {
switch r.kind {
case reqSense:
resp := response{sense: senseAt(e.world, r.pos)}
if to, ok := e.pending[r.id]; ok {
delete(e.pending, r.id)
resp.teleport, resp.hasTeleport = to, true
}
r.reply <- resp
case reqTeleport:
if !e.world.InBounds(r.pos) {
r.reply <- response{err: &world.OutOfBoundsError{Pos: r.pos, W: e.cfg.Width, H: e.cfg.Height}}
return
}
e.pending[r.id] = r.pos
r.reply <- response{ok: true}
}
}
// in runAnt, right after a successful sense:
if sensed.hasTeleport {
a.Pos = sensed.teleport
continue // re-sense from the new position before deciding
}
Three things to take from this:
pending map is owner-only state, so a map is perfectly safe here. The thing that made maps dangerous in Milestone 4 was sharing, not maps.continue matters. Deciding on a Sense gathered at the old position after teleporting would make the ant act on a view of somewhere it no longer is. Stale reads are the recurring theme of this milestone.Set QueueSize to 1 and run 2,000 ants. Then set it to 65536. Watch delivered-per-second and the timeout counter. A tiny queue means senders block constantly, which is slow but self-limiting; a huge queue means nobody ever blocks, latency grows without bound, and ants time out on replies to requests that are still sitting in line. Neither extreme is good, and the middle is not obvious. That tension is what Milestone 11 is about.
An unbuffered reply channel. The owner blocks forever on r.reply <- resp if the ant has stopped listening. The whole simulation freezes and the stack dump shows every goroutine blocked on the same channel. Buffer of one, always.ant *Ant in the request struct, stop.reqs to signal shutdown. Senders panic with "send on closed channel", and there are thousands of them. Only close a channel when there is exactly one sender, and here there are many. Cancel a context instead.Stats from inside handle. The owner would send itself a request and wait for a reply it can only produce by returning. Instant, permanent, self-inflicted deadlock. This is the message-passing equivalent of non-reentrant locks, and the same discipline applies: internal code calls the internal function, not the public one.Sense use [8]int rather than []int?Ant.Pos, and how would you catch a violation?time.NewTimer plus Stop instead of time.After in roundTrip? The single-owner pattern is not a Go idea — it is the actor model, and you can write it in Erlang, Elixir or Akka too. What Go contributes is that it is cheap here: 1,000 ant goroutines plus one owner goroutine plus 1,000 buffered channels cost a few megabytes and start in microseconds. The same shape in a language with OS-thread-weight concurrency would force you toward a thread pool and a work queue long before you reached this ant count, which is a different, more complicated program.
The honest cost: nothing in the language stops you from breaking the single-owner rule. The comment on the struct fields is the entire enforcement mechanism, and it is enforced by discipline plus go test -race, not by the compiler. Erlang and Rust each take this further in their own way — Erlang by making every process's memory genuinely private, Rust by making a smuggled pointer a compile error rather than a comment. Go trusts you, which is fast to write and easy to get wrong once a codebase has more than one author.
Ants that find food lay a trail on the way home; searching ants are attracted to trails; trails evaporate. The colony gets measurably better at foraging, and we gain a second background goroutine with its own timing.
time.Ticker and periodic work, non-blocking sends as load shedding, weighted random choice, float grids, and why every positive feedback loop needs a cap and a decay.
Ant colony optimisation in three rules:
1 + bias × pheromone. Rule 3 is not decoration. Without evaporation the grid saturates: every cell eventually reaches the cap, all weights become equal, and the trail carries no information. Worse, an early lucky path to an exhausted food source stays attractive forever. Evaporation is what makes the trail a memory of recent success rather than an accumulation of all history. Rule 2's 1 + is the same idea from the other side: it guarantees an unmarked cell still has a chance, so the colony keeps exploring instead of collapsing onto the first path it finds. Positive feedback plus a decay plus a floor on exploration, and this is exactly the structure of many reinforcement algorithms.
Where does evaporation run? It must mutate the grid, and only the owner may do that, so it becomes a third kind of client: a goroutine holding a ticker that sends reqEvaporate on a schedule.
ants ────────►┐
│ reqs chan ──► serve() (owns world)
evaporator ──►┤
Stats caller ►┘
A typical periodic job — a cron entry, a Quartz scheduler, a setInterval — assumes its work will run to completion every time it fires, and treats a run that is still busy when the next tick arrives as an operational problem to fix (usually by making the job faster or the interval longer). The scheduler and the work are separate concerns, and a missed tick is unusual enough to page someone.
The idiomatic Go version here treats a missed tick as the normal case under load: evaporate tries a non-blocking send and, if the owner's queue is full, simply skips that cycle. No error, no retry, no alert. This works because a time.Ticker paired with select and a default case turns "is anyone listening right now" into a two-line question you can ask on every tick, which most schedulers do not expose at all. The cost is that this policy is only correct because evaporation is idempotent and its value decays with delay; a periodic job with side effects that must not be skipped (say, a billing cycle) needs the traditional "queue it and guarantee it runs" discipline instead.
// FloatGrid is a grid of float64 values, used for pheromone concentration.
type FloatGrid struct {
W, H int
cells []float64
}
// AddAt deposits delta at p, capping the cell so a single trail cannot
// dominate the whole grid.
func (g *FloatGrid) AddAt(p Position, delta, max float64) {
if !g.InBounds(p) {
return
}
i := p.Y*g.W + p.X
g.cells[i] += delta
if g.cells[i] > max {
g.cells[i] = max
}
}
// Evaporate multiplies every cell by factor (0 < factor < 1) and zeroes
// values too small to matter, which keeps the grid sparse in practice.
func (g *FloatGrid) Evaporate(factor float64) {
for i, v := range g.cells {
v *= factor
if v < 1e-4 {
v = 0
}
g.cells[i] = v
}
}
The 1e-4 floor is worth a sentence. Multiplying by 0.95 forever produces ever-smaller denormal floats, which are both meaningless and, on some hardware, dramatically slower to compute with. Snapping to zero keeps the arithmetic fast and the data honest. Whenever you write an exponential decay, decide where it stops.
Evaporate is O(width × height), so on a 512×512 grid it touches 262,144 cells every time it runs. That is a real cost sitting on the owner's goroutine, blocking every ant behind it. Remember this when you profile in Milestone 11; the fix is either a coarser schedule or a lazy scheme that decays a cell when it is read, using the timestamp of its last update.
// evaporate asks the owner to decay the pheromone grid on a fixed schedule.
func (e *Engine) evaporate(ctx context.Context) {
ticker := time.NewTicker(e.cfg.EvaporateEvery)
defer ticker.Stop()
for {
select {
case <-ctx.Done():
return
case <-ticker.C:
select {
case e.reqs <- request{kind: reqEvaporate}:
case <-ctx.Done():
return
default:
// The owner is saturated. Skipping one evaporation is
// better than stalling the clock behind the ants.
}
}
}
}
time.NewTicker fires repeatedly; time.NewTimer fires once. A ticker that is never stopped keeps firing after you stop reading it, which pins the goroutine and leaks. defer ticker.Stop() is not optional. select has three cases and the default is the interesting one. A default makes the send non-blocking: if the owner's queue is full, we skip this evaporation entirely rather than queueing behind ten thousand ant requests. That is load shedding, and it is the correct policy for periodic work whose value decays: a late evaporation is worth less than no evaporation, because by the time it runs the schedule has already slipped. // TrailFollower is Forager plus pheromones: when searching it biases its
// steps toward neighbouring cells that other ants have marked.
type TrailFollower struct {
// Bias is how strongly pheromone attracts. 0 is a random walk.
Bias float64
}
func (t TrailFollower) Decide(a *Ant, s Sense, rng *rand.Rand) Action {
if a.Carrying {
if s.Pos == s.Nest {
return Action{Kind: ActDrop}
}
return Action{Kind: ActMove, Dir: stepToward(s.Pos, s.Nest)}
}
if s.Food > 0 {
return Action{Kind: ActPickUp}
}
if i, ok := bestAdjacentFood(s); ok {
return Action{Kind: ActMove, Dir: directions[i]}
}
return Action{Kind: ActMove, Dir: t.weightedStep(s, rng)}
}
// weightedStep picks a neighbour with probability proportional to
// 1 + Bias*pheromone, so an unmarked grid still produces a random walk.
func (t TrailFollower) weightedStep(s Sense, rng *rand.Rand) world.Position {
var weights [8]float64
total := 0.0
for i, p := range s.Pher {
w := 1 + t.Bias*p
weights[i] = w
total += w
}
r := rng.Float64() * total
for i, w := range weights {
r -= w
if r <= 0 {
return directions[i]
}
}
return directions[rng.IntN(8)] // unreachable except for float rounding
}
Weighted random selection by subtracting from a uniform sample in [0, total) is worth knowing: it needs no sorting, no allocation, and one pass. The final return after the loop looks like dead code and is not: floating-point rounding can leave r a hair above zero after subtracting every weight. Reviewers will ask about that line, so the comment is part of the code.
TrailFollower has a value receiver and one field, so it is copied into the interface rather than shared. Every ant reads the same Bias and nothing writes it. Compare with the Scout from Milestone 3, whose per-ant maps would have been a race: the difference is that this strategy keeps no per-ant state at all.
Deposition lives at the owner, in applyAction, the three lines you already saw:
case ActMove:
np := e.world.Clamp(r.pos.Add(r.act.Dir))
if r.carrying {
e.world.Pher.AddAt(np, e.cfg.Deposit, e.cfg.PherMax)
}
return response{pos: np, ok: true}
Same seed, same world, same wall-clock budget, 300 ants on a 64×64 grid with six food sources of 200 units each:
forager: delivered 323, pheromone left 1693
trail : delivered 404, pheromone left 397
A 25% improvement in delivered food, measured rather than asserted. Two details in those numbers are more interesting than the headline.
The forager run has more pheromone left at the end (1693 against 397) even though it ignores pheromone entirely. Deposition happens at the owner for any carrying ant, so both runs lay trails; only one reads them. The forager's trails accumulate because its ants take longer random-walk journeys and are more spread out, while the trail follower's ants concentrate on short paths that evaporation keeps trimmed. A metric moving in the direction you did not predict is usually the most informative thing on the screen.
Second: this is a measurement of one seed on one machine over half a second, which is not a result, it is an anecdote. Before claiming pheromones help, you would run several seeds and compare distributions. That is Exercise 6.
Turn the anecdote into evidence, then find where the benefit lives.
Bias over 0, 1, 2, 4, 8, 16 and plot delivered food. There is an optimum, and beyond it the colony gets worse. Explain why in terms of exploration and exploitation.Evaporation over 0.99, 0.95, 0.8, 0.5 with Bias fixed. Explain both ends of the resulting curve.func TestTrailBeatsRandomWalk(t *testing.T) {
const runs = 10
wins, forageTotal, trailTotal := 0, 0, 0
for seed := range uint64(runs) {
cfg := Config{Width: 64, Height: 64, Ants: 300,
FoodSources: 6, FoodPerSource: 200, Seed: seed}
a := runFor(t, cfg, Forager{}, 300*time.Millisecond)
b := runFor(t, cfg, TrailFollower{Bias: 4}, 300*time.Millisecond)
forageTotal += a.Delivered
trailTotal += b.Delivered
if b.Delivered > a.Delivered {
wins++
}
}
t.Logf("forager mean %d, trail mean %d, trail won %d/%d",
forageTotal/runs, trailTotal/runs, wins, runs)
if wins < runs*7/10 {
t.Errorf("trail following won only %d of %d seeds", wins, runs)
}
}
Mean versus majority. A mean can be carried by one spectacular run: if the trail follower occasionally finds a rich source early and delivers ten times as much, the average looks wonderful while it loses most individual races. Counting wins asks "is this reliably better?", which is usually the question you care about. Reporting both, as the t.Logf does, costs nothing.
Note also that this test is time-based, so it measures throughput on whatever machine runs it, and it will be flaky on a loaded CI box. A more rigorous version fixes the number of ticks rather than the wall clock, which is a good argument for keeping the sequential Sim from Milestone 3 around as a measurement harness. Deterministic experiments belong in a deterministic runner.
The bias sweep. You should see improvement up to roughly 4–8, then decline. Low bias is a random walk that ignores information. High bias makes the first trail overwhelmingly attractive, so every ant follows it, deposits more on it, and makes it more attractive still: the colony locks onto one path and stops exploring, and when that food source runs out it keeps walking the trail to an empty cell until evaporation erases it. This is the exploration/exploitation trade-off, and it is the same failure mode as a recommender system that only ever shows you what you already clicked.
The evaporation sweep. At 0.99 trails persist far too long, so stale paths to exhausted sources keep attracting ants. At 0.5 a trail is gone within a few ticks and cannot be reinforced fast enough to guide anyone, so you have paid for pheromones and got a random walk. The useful range is narrow and depends on how long a round trip takes, which is a hint that the parameter should really be expressed relative to journey time rather than as an absolute per-tick factor.
Set PherMax to a huge number (say 1e9) and run with Bias: 4. The first trail to form saturates, and because weights are proportional to concentration, the colony funnels every ant onto one path within seconds. Watch delivered food collapse. Then set PherMax to 1 and watch the trail carry almost no signal. Every positive feedback loop in a simulation needs a ceiling, and finding it experimentally teaches more than choosing it analytically.
time.NewTicker without defer Stop() leaks the ticker and its goroutine. The leak test in Milestone 7 catches it.default, evaporation queues behind ant traffic, the schedule slips under load, and trails stop decaying exactly when the colony is busiest.1 + in the weight. With weights equal to raw pheromone, a fresh grid has total weight zero, every weight is zero, and the loop falls through to the fallback on every single step. Symptom: the colony behaves randomly and no amount of bias tuning changes anything.weightedStep add 1 to every weight?default case in the evaporator's inner select do, and what policy does it encode?Evaporate is O(cells) and runs on the owner's goroutine. What does that imply at 512×512 with a 20 ms period?TrailFollower safe to share across 50,000 ants when Scout from Milestone 3 was not?Nothing about pheromones needed Go specifically — FloatGrid.Evaporate and weightedStep are ordinary numeric code you could write identically in Python or Java. What Go earns its keep on here is the plumbing around that code: a time.Ticker plus a non-blocking select gives the evaporator goroutine a load-shedding policy in four lines, and because it is just another client of the same owner, it needed no new synchronisation primitive at all, only a new message kind.
The honest verdict: this milestone's actual algorithm (weighted random choice with a decay term) is the interesting part, and Go contributed nothing to that part beyond what any language with closures and a decent standard library would. Where a language choice would matter more is at much larger scale — millions of cells decaying continuously — where you would want either a lazy, timestamp-based decay (no periodic sweep at all) or a language with SIMD-friendly array operations built in, neither of which this milestone needed yet.
Ctrl-C stops the colony cleanly: every ant finishes its current round trip, the final statistics are collected, background goroutines exit, and the process leaves nothing running. A test proves there are no leaks.
Context trees and derived cancellation, signal.NotifyContext, shutdown ordering, the WaitGroup misuse that bites concurrent supervisors, and goroutine leak detection.
"Just cancel everything" is wrong, and the reason is worth spelling out. The last thing Run does is collect final statistics, which requires sending a request to the owner and getting a reply. If the owner was cancelled along with everyone else, that request hangs until it times out and you get empty results. So shutdown has an order, and the order is the reverse of the dependency graph:
ctx cancelled (Ctrl-C, timeout, or the caller's choice)
│
▼
1. supervisor stops no new ants will be started
│
▼
2. ants drain and exit each finishes its current round trip
│
▼
3. final Stats collected the owner is still alive to answer
│
▼
4. evaporator stops nothing left that needs a clock
│
▼
5. owner stops last to go, because everyone needed it
This is why Run builds three contexts rather than passing one everywhere. The ants get the caller's context, so cancelling it stops them. The evaporator and the owner get their own contexts derived from context.Background(), so they outlive the ants and stop only when Run says so. A context is a cancellation scope, and "everything cancels at once" is a design decision, not a default you have to accept.
A typical shutdown path in many services is a single flag or a single cancellation signal that everything listens to, on the theory that "stop" should mean stop, everywhere, at once. It is simple to reason about and it is exactly wrong the moment cleanup itself needs something that "everything" includes — here, the final Stats call needs the owner, and the owner would already be gone.
Go's context package makes the idiomatic fix cheap: contexts form a tree, so Run derives three separate cancellation scopes (ants, evaporator, owner) instead of sharing one, and shuts them down in the order the dependency graph actually requires. The cost is that you now have to think about that order explicitly — three context.WithCancel calls and a comment explaining why, instead of one flag everyone reads. For a program with no cleanup dependencies, one shared context would have been simpler and just as correct; the extra structure here is earned by "the last thing we do needs something we are shutting down."
// Run starts the colony and blocks until ctx is cancelled, then shuts
// everything down in order and returns the final snapshot.
func (e *Engine) Run(ctx context.Context) (Stats, error) {
serverCtx, stopServer := context.WithCancel(context.Background())
var serverWG sync.WaitGroup
serverWG.Add(1)
go func() {
defer serverWG.Done()
e.serve(serverCtx)
}()
evapCtx, stopEvap := context.WithCancel(context.Background())
var evapWG sync.WaitGroup
evapWG.Add(1)
go func() {
defer evapWG.Done()
e.evaporate(evapCtx)
}()
e.ants = make([]*Ant, 0, e.cfg.Ants)
for i := range e.cfg.Ants {
e.ants = append(e.ants, &Ant{ID: i, Pos: e.nest, Energy: e.cfg.StartEnergy})
}
deaths := make(chan *Ant, e.cfg.Ants)
var antWG sync.WaitGroup
start := func(a *Ant) {
antWG.Add(1)
go func() {
defer antWG.Done()
e.superviseAnt(ctx, a, deaths)
}()
}
for _, a := range e.ants {
start(a)
}
// The supervisor is the only goroutine that may add more ants, so we
// wait for it to finish before waiting on the ant group.
supervisorDone := make(chan struct{})
go func() {
defer close(supervisorDone)
for {
select {
case <-ctx.Done():
return
case a := <-deaths:
e.recycle(ctx, a)
start(a)
}
}
}()
<-ctx.Done()
<-supervisorDone
antWG.Wait()
stats, err := e.Stats(context.Background())
stopEvap()
evapWG.Wait()
stopServer()
serverWG.Wait()
return stats, err
}
<-ctx.Done()
<-supervisorDone
antWG.Wait()
Those look redundant and are not. sync.WaitGroup has a documented rule that is easy to violate: calls to Add must not happen concurrently with Wait when the counter could be zero. Our supervisor calls antWG.Add(1) every time it restarts an ant. If Run called antWG.Wait() while the supervisor was still alive, and every ant happened to be dead at that instant, Wait could return while the supervisor was starting a fresh one, and we would shut down the owner from under a running ant. Waiting for the supervisor to exit first guarantees no further Add can occur. The Go runtime detects some violations with panic: sync: WaitGroup misuse: Add called concurrently with Wait, but not all of them, and a race here is a shutdown hang that reproduces once a week.
Then e.Stats(context.Background()) deliberately uses a fresh context. Passing the caller's cancelled ctx would return context.Canceled immediately and give you an empty snapshot. Reaching for Background during cleanup is a common and correct pattern; in production code you would normally give it a short timeout so cleanup cannot hang forever.
// Ctrl-C cancels the context; a second one kills the process outright.
ctx, stop := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM)
defer stop()
ctx, cancel := context.WithTimeout(ctx, duration)
defer cancel()
signal.NotifyContext returns a context that cancels when one of the listed signals arrives. The nesting matters: the timeout context is derived from the signal context, so the run ends on whichever comes first, a signal or the deadline. That is context composition doing real work in four lines.
defer stop() restores the default signal handling, which is what makes a second Ctrl-C kill the process immediately. That is the behaviour users expect: the first interrupt asks nicely, the second insists. A program that swallows every interrupt while it "gracefully" shuts down is a program people learn to kill -9.
func TestNoGoroutineLeak(t *testing.T) {
before := countGoroutines()
cfg := Config{Width: 16, Height: 16, Ants: 100, FoodSources: 5, Seed: 3}
runFor(t, cfg, Forager{}, 150*time.Millisecond)
// Goroutines exit asynchronously; give them a moment before judging.
for range 50 {
if countGoroutines() <= before+1 {
return
}
time.Sleep(10 * time.Millisecond)
}
t.Errorf("goroutines leaked: before %d, after %d", before, countGoroutines())
}
=== RUN TestNoGoroutineLeak
--- PASS: TestNoGoroutineLeak (0.15s)
The retry loop is the part people get wrong. runtime.NumGoroutine() is a snapshot, and a goroutine that has been told to stop takes an unspecified moment to actually stop, so a single comparison immediately after Run returns is flaky in both directions. Polling with a deadline converts "eventually true" into a test that passes quickly when correct and fails reliably when not. The +1 tolerance covers the testing framework's own bookkeeping.
A goroutine leak is Go's version of a memory leak, and it is worse in one respect: a leaked goroutine holds everything it references, so one leaked ant keeps its world, its channels and its buffers alive forever. In a long-running service the symptom is memory growth that no heap profile explains until you look at the goroutine profile. Write this test early, run it in CI, and reach for go.uber.org/goleak when you want the industrial version, which names the leaked stacks instead of just counting.
errgroupProduction Go frequently uses golang.org/x/sync/errgroup, which combines a WaitGroup, error propagation and context cancellation:
g, ctx := errgroup.WithContext(ctx)
g.Go(func() error { return e.serve(ctx) })
g.Go(func() error { return e.evaporate(ctx) })
err := g.Wait() // first non-nil error, and ctx is cancelled for the restIt is genuinely good and you should know it. I have kept this project dependency-free so far, partly to avoid hiding behaviour behind a library before you have implemented it yourself, and partly because errgroup's "cancel everyone on the first error" semantics are the opposite of what this shutdown sequence needs. Add it with go get golang.org/x/sync when you want it; it is the one external dependency I would not argue about.
Add a shutdown deadline. If the ants have not all exited within cfg.ShutdownGrace (default two seconds), Run should log how many are still running, collect statistics anyway, stop the owner, and return a non-nil error that the caller can distinguish from a normal stop. Requirements: no goroutine may be left blocked on a channel send; errors.Is must work against your new sentinel; the leak test must still pass in the happy path; and add a test that forces the slow path by setting a large Chaos.SlowFor.
Hint: antWG.Wait() blocks and has no timeout. You cannot add one to a WaitGroup, so you need to convert "the group finished" into something select can wait on.
var ErrShutdownTimeout = errors.New("ants did not stop before the deadline")
// waitChan converts a WaitGroup into a channel that closes when the group
// reaches zero, so it can be used in a select.
func waitChan(wg *sync.WaitGroup) <-chan struct{} {
done := make(chan struct{})
go func() {
wg.Wait()
close(done)
}()
return done
}
<-ctx.Done()
<-supervisorDone
var shutdownErr error
select {
case <-waitChan(&antWG):
// clean stop
case <-time.After(e.cfg.ShutdownGrace):
shutdownErr = fmt.Errorf("%w after %v", ErrShutdownTimeout, e.cfg.ShutdownGrace)
log.Printf("antfarm: %v; collecting stats anyway", shutdownErr)
}
stats, err := e.Stats(context.Background())
if shutdownErr != nil {
err = shutdownErr
}
stopEvap()
evapWG.Wait()
stopServer()
serverWG.Wait()
return stats, err
The waitChan helper is the standard answer to "I need a WaitGroup in a select". Note that it starts a goroutine which survives until the group drains, so on the timeout path it outlives Run. That is acceptable here (it will exit when the stragglers do) but it is exactly the kind of thing your leak test will notice, which is why the exercise asks you to force the slow path deliberately.
The deeper lesson: you cannot force a goroutine to stop. There is no goroutine.Kill() in Go, by design. If an ant is asleep inside time.Sleep(10 * time.Second), nothing in the language will wake it, and the best you can do is stop waiting and record that you did. Every graceful shutdown in Go is therefore cooperative, with a deadline and a report, and "how do I kill a stuck worker?" has exactly one answer: you do not, you design so that workers check for cancellation often enough that they never get stuck. Erlang's answer to this question is exit(Pid, kill), which is unconditional and immediate, and Course 4 will show what becomes possible when that primitive exists.
Comment out defer ticker.Stop() in the evaporator and run the leak test. Then restore it and remove the <-supervisorDone line, and run the whole suite with -race -count=20. The first is a reliable failure; the second is an intermittent one, which is a good demonstration of why concurrency bugs need repeated runs rather than a single green tick.
WaitGroup.Add racing with Wait. Symptom: rare shutdown hangs or a runtime panic about WaitGroup misuse.os.Exit in a signal handler. Deferred functions do not run, buffered output is lost, and the graceful path you wrote never executes.WaitGroup whose goroutines are blocked on a full channel nobody is draining. Classic shutdown deadlock: stop the consumer before the producers and everything wedges.defer runs on panic in another goroutine. It runs in the panicking goroutine only. An unrecovered panic anywhere kills the whole process, deferred cleanup elsewhere included.NumGoroutine immediately. Poll with a deadline.Background rather than the caller's context? antWG.Wait() runs while the supervisor is still alive?Stats use a fresh context during shutdown?defer stop() after signal.NotifyContext buy the user?time.Sleep(30 * time.Second) when you cancel. What are your options?context.Context is the part of Go that makes this milestone almost pleasant: cancellation that composes (a timeout derived from a signal-triggered cancellation, in two lines), propagates automatically through every function that accepts a ctx parameter, and is checked cooperatively rather than requiring a separate kill mechanism per goroutine kind. sync.WaitGroup and signal.NotifyContext round out a shutdown story that needs zero third-party dependencies.
The honest cost: every one of these primitives is cooperative, which is the same fact as "you cannot kill a goroutine" wearing different clothes. context gives you a clean way to ask every goroutine to stop; it gives you nothing if a goroutine does not check. Compare Erlang, where a supervisor can unconditionally terminate a process regardless of what it is doing, or a language with structured concurrency and cancellation points enforced by the runtime. Go's answer is disciplined convention, not a guarantee, and this milestone's entire second half exists because that distinction matters in practice, not just in theory.
Break the colony on purpose and keep it running. Ants panic at random and a supervisor restarts them. The owner drops messages at random and clients time out and retry. Ants stall at random and the queue absorbs it. Every failure becomes a number you can watch.
panic and recover across goroutine boundaries, supervision built by hand, restart state and generations, accounting for work lost in a crash, at-least-once versus at-most-once delivery, and the way timeouts interact with throughput.
Three failures, each modelling something real:
| Injected | Models | Handled by |
|---|---|---|
crash=p | A bug, an assertion failure, a nil dereference in one worker | recover + supervisor restart |
drop=p | A lost packet, a dropped RPC, a full queue upstream | client timeout and retry |
slow=p | A GC pause, a slow disk, a noisy neighbour | queueing, and eventually backpressure |
Go's failure model has one fact you must design around: an unrecovered panic in any goroutine terminates the entire process. Not just that goroutine. There is no supervisor in the runtime, no isolation between goroutines, and no way for one goroutine to catch another's panic. recover works only inside a deferred function in the panicking goroutine itself.
So supervision has to be built, and the shape is forced by that constraint:
start(a) ──► goroutine ──► superviseAnt(ctx, a, deaths)
│
│ defer func(){ recover() } ← must be here
▼
runAnt(...) ← panics here
│
panic unwinds to the deferred func
│
▼
deaths <- a (a channel of dead ants)
│
▼
supervisor: recycle(a); start(a)
The supervisor is a goroutine reading a channel of casualties. That is the entire mechanism, and it is about fifteen lines. Keep the number in mind when Course 4 shows you what Erlang gives you for free.
// Chaos describes the failures we inject on purpose.
type Chaos struct {
CrashProb float64 // per ant, per tick
DropProb float64 // per message, at the owner
SlowProb float64 // per ant, per tick
SlowFor time.Duration // how long a slow tick sleeps
}
In the ant loop, before each cycle:
if e.cfg.Chaos.CrashProb > 0 && c.rng.Float64() < e.cfg.Chaos.CrashProb {
panic(fmt.Sprintf("ant %d: simulated crash at %v", a.ID, a.Pos))
}
if e.cfg.Chaos.SlowProb > 0 && c.rng.Float64() < e.cfg.Chaos.SlowProb {
time.Sleep(e.cfg.Chaos.SlowFor)
}
And at the top of the owner's handle:
// Chaos: pretend the message never arrived. No reply is sent, so the
// caller will time out. This is a dropped packet, simulated.
if e.cfg.Chaos.DropProb > 0 && e.rng.Float64() < e.cfg.Chaos.DropProb {
e.dropped.Add(1)
return
}
Chaos uses each goroutine's own random source, so it inherits the same seeding discipline as everything else and a chaotic run is as reproducible as a chaotic run can be. The probability checks are guarded by > 0 so that a zero-chaos configuration does not pay for a random number on every tick, which matters in the hot loop.
// superviseAnt runs one ant and turns a panic into a restart request.
func (e *Engine) superviseAnt(ctx context.Context, a *Ant, deaths chan<- *Ant) {
defer func() {
r := recover()
if r == nil {
return
}
e.restarts.Add(1)
select {
case deaths <- a:
case <-ctx.Done():
}
}()
e.runAnt(ctx, a)
}
recover() returns nil when there was no panic, which is how the deferred function distinguishes a crash from a normal return. Calling recover() outside a deferred function does nothing at all, silently, and that is a common bug.deaths is inside a select with ctx.Done(). During shutdown the supervisor has already exited, so an unguarded send would block forever and leak the goroutine on exactly the path where you are least likely to look.deaths chan<- *Ant is a send-only channel in the signature. The compiler now guarantees this function cannot read from it. Direction annotations are free documentation that the compiler checks.debug.Stack(), because a recovered panic with no stack trace is a bug you will never find.It is tempting to wrap every goroutine in a recover and call the system robust. What you have actually built is a system that continues running in a state its author never considered, with a corrupted invariant and no report. The Go community's position, which I think is right, is that recover is for two things: a process boundary where a crash must not take down unrelated work (an HTTP handler, a plugin, one ant), and turning a panic into an error at a library's public edge. Everywhere else, let it crash and fix the bug.
The version here qualifies, barely, because an ant is a genuine isolation boundary and because we count and report every restart. Note what it still cannot do: if the owner goroutine panics, the process dies and nothing survives. That asymmetry (one shared component whose failure is fatal) is inherent to the single-owner design, and it is a real cost to weigh against everything the design bought us in Milestone 5.
// recycle resets a crashed ant before it restarts. A crash loses whatever
// the ant was carrying, and the colony has to account for that.
func (e *Engine) recycle(ctx context.Context, a *Ant) {
if a.Carrying {
select {
case e.reqs <- request{kind: reqLost}:
case <-ctx.Done():
}
}
a.Carrying = false
a.Pos = e.nest
a.Energy = e.cfg.StartEnergy
a.Generation++
}
This function is the most interesting fifteen lines in the milestone, for a reason that is not obvious.
Who owns the ant here? The rule since Milestone 5 has been "the ant's own goroutine writes its fields, nobody else". Yet recycle runs on the supervisor's goroutine and writes four of them. That is legal, and it is legal for a specific reason: the ant's goroutine is dead. Ownership transferred at the moment the panic unwound, and it transfers back when start(a) launches the replacement. The send on deaths and the later go statement are both synchronisation events, so the Go memory model guarantees the supervisor sees the dead ant's final state and the new goroutine sees the supervisor's writes. Ownership that moves between goroutines, handed over by channel operations, is the mental model that makes this kind of code tractable, and it is the same idea Rust encodes in its type system.
The crash destroys a unit of food. The owner already counted that unit in carrying when the pickup succeeded. If we simply restarted the ant, the ledger would claim a unit is in transit that no ant is carrying, and the conservation test would drift. So the supervisor tells the owner: one unit lost. The test then checks
delivered + on the ground + carried + lost == the food we started withand passes under crash chaos. In-flight work disappears when a worker dies, and if your metrics do not model that, they will quietly lie to you. That lesson costs real companies real money, and here it costs one integer.
Generation increments so a restarted ant does not replay the same random sequence (the seed is uint64(a.ID) + a.Generation<<32). Without it, an ant that crashed would repeat its exact previous life, including the crash, at the same point. Deterministic chaos that always kills the same ant in the same place is not chaos.
$ go run ./cmd/antfarm -ants 2000 -grid 96x96 -food 8 -food-per-source 400 \
-duration 3s -every 1s -chaos crash=0.0005,drop=0.001
colony: 2000 ants, 96x96 grid, 8 food sources of 400, behaviour trail, chaos {CrashProb:0.0005 DropProb:0.001 SlowProb:0 SlowFor:0s}
tick 44 ants 2000 carrying 32 delivered 504 food left 2663 lost 1 restarts 186 dropped 735 timeouts 574 pher 850
tick 93 ants 2000 carrying 17 delivered 787 food left 2394 lost 2 restarts 372 dropped 1513 timeouts 1372 pher 284
final: tick 143 ants 2000 carrying 4 delivered 860 food left 2332 lost 4 restarts 571 dropped 2298 timeouts 2155 pher 1272
Real output. Two thousand ants, 571 crashes and restarts in three seconds, 2,298 dropped messages, and the colony delivered 860 units of food while all of that was happening. Four units were lost to crashes and the books balance.
Look at dropped 2298 against timeouts 2155. They should be nearly equal, since every dropped message costs its sender a timeout, and the small gap is messages dropped near the end whose senders were cancelled before their timers fired. When two counters that should track each other stop tracking each other, you have found something. Deliberately instrumenting both sides of a mechanism, rather than one, is what makes that possible.
Here are two runs of the same configuration, differing only in drop probability, over the same wall-clock budget:
drop = 0.02 : dropped 1601, timeouts 1400, delivered 327
drop = 0.25 : dropped 1608, timeouts 1400, delivered 13
Twelve times the message loss, but the same number of dropped messages, and delivered food collapsed by 96%. Read that again, because it is not intuitive.
The explanation: with a 50 ms timeout, an ant that loses a message spends 50 ms doing nothing, which is hundreds of times longer than a successful round trip. At 2% loss the colony still spends most of its time working. At 25% loss, one message in four costs 50 ms, so almost all the time goes into waiting, total traffic falls through the floor, and the absolute number of dropped messages stays flat because there are so few messages left to drop.
The timeout, not the loss rate, is what destroyed throughput. This is a genuine and very common production failure: a service degrades slightly, clients time out, retries multiply the load, and the system collapses far faster than the original fault would suggest. The mitigations are all things this code does not do yet, and they are worth naming: exponential backoff so retries spread out, jitter so clients do not synchronise, a retry budget so a client gives up, and a circuit breaker so a failing dependency stops being called at all. Milestone 11 adds backpressure, which addresses the same disease from the other end.
This milestone is where Go's limits show, and it would be dishonest to pretend otherwise.
What Go did well: injecting chaos cost about twenty lines, the supervisor is small enough to read in one sitting, and because ants are goroutines rather than threads, 571 restarts in three seconds is unremarkable. Restarting 571 OS threads in three seconds would be a noticeable event.
What Go made hard, all stemming from one design choice, that goroutines are not isolated:
Erlang answers all four at the language and runtime level: processes have isolated heaps so one cannot corrupt another, exit(Pid, kill) is unconditional, and supervisors come with restart strategies, intensity limits and escalation as configuration rather than code. If fault tolerance were the whole problem, Erlang would be the better tool, and Course 4 is going to make that case with the same project in miniature. Go's counter-argument is everything else: static typing, raw speed, a single deployable binary, and a profiler that will save you in Milestone 11. Most systems need some of both, which is why "which language is better" is the wrong question and "what is this language's failure model" is the right one.
Our supervisor restarts an ant instantly, forever. That is the one restart policy OTP explicitly warns against. Implement a better one.
Requirements. Add exponential backoff with jitter (first restart after ~10 ms, doubling, capped at one second). Add a restart intensity limit: if more than MaxRestarts happen within RestartWindow, stop restarting, record the colony as degraded, and let Run return an error identifying it. Expose Degraded bool and Alive int on Stats.
Constraints. No mutex. Backoff must not block the supervisor, since one sleeping ant must not delay restarting the other 49,999. Shutdown must still be prompt: a pending backoff must not keep Run waiting.
Acceptance. A test with CrashProb: 0.5 ends degraded rather than burning CPU; a test with CrashProb: 0.001 does not end degraded; the leak test still passes; and the whole suite is clean under -race -count=5.
Hints. Where does the sleep go so that it does not block the supervisor? What does the restart counter need to be, given that only the supervisor touches it? How does a sleeping restart learn that shutdown has begun?
var ErrColonyDegraded = errors.New("restart intensity exceeded; colony degraded")
// restartTracker is touched only by the supervisor goroutine, so it needs
// no synchronisation.
type restartTracker struct {
times []time.Time
window time.Duration
max int
}
func (r *restartTracker) record(now time.Time) (overloaded bool) {
cutoff := now.Add(-r.window)
kept := r.times[:0] // reuse the backing array: no allocation
for _, t := range r.times {
if t.After(cutoff) {
kept = append(kept, t)
}
}
r.times = append(kept, now)
return len(r.times) > r.max
}
tracker := &restartTracker{window: e.cfg.RestartWindow, max: e.cfg.MaxRestarts}
degraded := make(chan struct{})
go func() {
defer close(supervisorDone)
for {
select {
case <-ctx.Done():
return
case a := <-deaths:
if tracker.record(time.Now()) {
e.degraded.Store(true)
close(degraded) // tell Run; ants already running continue
return
}
e.recycle(ctx, a)
delay := backoff(a.Generation, e.cfg.RestartBackoff, time.Second, e.superRNG)
// The sleep happens in the restarting goroutine, not here,
// so one slow restart cannot delay the others.
antWG.Add(1)
go func(a *Ant) {
defer antWG.Done()
select {
case <-time.After(delay):
case <-ctx.Done():
return // do not start an ant during shutdown
}
e.superviseAnt(ctx, a, deaths)
}(a)
}
}
}()
func backoff(generation uint64, base, max time.Duration, rng *rand.Rand) time.Duration {
d := base << min(generation, 20) // doubling, saturating
if d > max || d <= 0 {
d = max
}
// Full jitter: anywhere in [0, d). Prevents restart storms from
// synchronising after a correlated failure.
return time.Duration(rng.Int64N(int64(d) + 1))
}
Five things worth extracting from this solution.
time.Sleep in the supervisor's loop would serialise every restart behind every backoff, so one crash-looping ant would delay the recovery of all the others. Whenever a supervisor waits, ask what it is not doing while it waits.antWG.Add(1) happens in the supervisor, before the go. Doing it inside the new goroutine would let Run's Wait return between the two, which is the same hazard the supervisorDone channel exists to prevent.restarts and dropped, which are written by many. Choosing between "owned" and "atomic" by asking who writes it is the whole discipline.kept := r.times[:0] is the standard Go filter-in-place idiom: reslice to zero length, keep the capacity, append what survives. It reuses the backing array and allocates nothing, and it works because append writes over elements the loop has already read. Compare the total with OTP's {one_for_one, MaxRestarts, Window}, which is one line of configuration and also handles escalation to a parent supervisor, ordered restarts of dependent children, and shutdown protocols per child. We have written maybe seventy lines for a strictly weaker version. That is not an argument against Go; it is an argument for knowing what you are reimplementing.
Run with -chaos slow=0.01,slowfor=200ms and 5,000 ants, then watch the timeout counter. Slow ants are not failing, they are just late, yet they consume queue slots and reply channels the entire time. Now raise slowfor to two seconds while keeping RequestTimeout at 200 ms and watch the system spend its life timing out. The interaction between a client's patience and a server's latency is where most production incidents actually live, and you can explore the whole space here with two flags.
recover() not inside a deferred function. Returns nil, does nothing, and the panic proceeds to kill the process. It must be defer func(){ recover() }(), not a bare call.debug.Stack() or you will never diagnose a real crash, only simulated ones.Carrying is still true will try to drop food it does not have and generate a stream of errors. deaths. Leaks a goroutine on every crash that happens during shutdown.rand. Then "it broke once" is all you will ever know.recover live in the same goroutine that panics, and what happens if you forget one?a.Pos and a.Carrying, which the ant's own goroutine normally owns. Why is that not a data race?Generation prevent?deaths chan<- *Ant written with a direction in the signature?antfarm/
├── go.mod
├── cmd/
│ └── antfarm/
│ └── main.go flags, chaos parsing, signals, reporting
└── internal/
├── sim/
│ ├── ant.go Ant, Action, ActionKind, Generation
│ ├── behaviour.go Behaviour, Forager, TrailFollower, directions
│ ├── sense.go Sense, senseAt
│ ├── engine.go requests, serve, antClient, supervisor, Run
│ ├── sim.go Config, Chaos, Stats, the sequential Sim
│ ├── concurrent.go milestone 4's mutex versions, kept for comparison
│ ├── engine_test.go conservation, stats-under-load, leaks, chaos
│ ├── sim_test.go decide tables, determinism, benchmarks
│ └── concurrent_test.go race reproduction, locked benchmarks
└── world/
├── grid.go Position, Grid
├── pheromone.go FloatGrid, deposit, evaporate
├── world.go World, food, nest, errors
└── grid_test.go bounds, food errors
$ gofmt -l . && go vet ./... && go test -race ./...
ok github.com/yourname/antfarm/internal/sim 1.965s
ok github.com/yourname/antfarm/internal/world 0.003s
$ git commit -am "milestone 8: chaos injection and a hand-rolled supervisor"
Worth pausing on: engine.go is now the heart of the project and contains no mutex at all. The only synchronisation primitives in it are channels, a context, three atomic counters for cross-goroutine metrics, and one WaitGroup per lifecycle group. Every piece of mutable state has exactly one owner, and the comments say who. That is what the Milestone 4 detour bought.