Milestones 1–4Milestones 9–12

Instalment 3 · Course 1 (Go) · Milestones 5–8

One goroutine owns the world, and everyone else asks it politely

The mutex disappears. Pheromone trails make the colony measurably smarter. Shutdown becomes something you can prove. Then we start breaking things on purpose and watch the supervisor put them back.

Everything here was run

All code was compiled, vetted, race-tested and executed before it reached the page, and the numbers quoted are measurements from those runs, not estimates. Where my single-core sandbox distorts a result, I say so.

Milestone 5The world becomes a goroutine

Goal

Delete the mutex. One goroutine owns the world; ants send it messages and get answers back. Along the way the Behaviour interface changes shape in a way that turns out to matter enormously later.

Concepts

Ownership as a design principle, request/response over channels, per-client reply channels, buffered channels as leak prevention, the "ask the owner" pattern for reading state, and value types as messages.

Design

The Mewlang cat, glancing over curiouslyThe failure of Milestone 4 was structural: many goroutines reaching into one mutable object. There are only three ways out of that, and it is worth seeing all three before picking one.

ApproachIdeaWhy not (here)
Finer locksOne mutex per cell or region262,144 locks, a mandatory global lock ordering, and a deadlock the first time a feature needs two regions at once
Lock-freeAtomics and compare-and-swap on cellsWorks for counters, becomes research-grade difficulty the moment an operation touches two cells
Single ownerOne goroutine holds the data; everyone else sends messagesThis one. Serialises writes by construction, needs no lock discipline, and survives becoming a network protocol

The shape we are moving to:

   ant 1 ──┐                              ┌──────────────────────┐
   ant 2 ──┤                              │   serve() goroutine  │
   ant 3 ──┼──►  reqs chan request  ──►   │                      │
    ...    │                              │  OWNS: world, food,  │
   ant N ──┘                              │  pheromone, counters │
                                          └──────────┬───────────┘
   each ant has its own                              │
   reply channel (buffered 1)   ◄────────────────────┘
                                   response sent to r.reply

Three decisions inside that picture deserve argument.

The reply channel belongs to the ant, not to the request. The obvious version allocates make(chan response) per request, which is one allocation and one garbage-collected object per ant per tick: at 50,000 ants that is millions of channels per second. Instead each ant creates one buffered channel at birth and reuses it forever. The buffer matters too, and not for speed: if the ant gives up waiting (a timeout, a cancelled context) and the owner later sends a reply, an unbuffered channel would block the owner permanently, wedging the entire simulation. A buffer of one means the owner always completes its send and moves on.

Messages carry values, never pointers. If a request contained *Ant, then the owner and the ant goroutine would both hold a pointer to the same struct, and we would be exactly back where Milestone 4 started, only with extra steps. Sending pos world.Position and carrying bool by value means the message is the shared state, and it is copied. This costs a few bytes and buys the property that the whole protocol could be serialised and sent over a socket without redesign. Milestone 12 does precisely that.

The ant perceives rather than reads. Decide currently takes *world.World and reads whatever it likes, which cannot work when the world lives in another goroutine. So the interface changes:

before:  Decide(a *Ant, w *world.World, rng *rand.Rand) Action
after:   Decide(a *Ant, s Sense, rng *rand.Rand) Action

A Sense is a small pointer-free snapshot of what one ant can perceive from where it stands. This is a better model of an ant anyway (real ants have no global map) and it makes every behaviour trivially testable with a literal. Concurrency pushed us toward a design we would have wanted regardless, which happens more often than you would expect.

Implementation

internal/sim/sense.go

package sim

import "github.com/yourname/antfarm/internal/world"

// Sense is everything one ant can perceive on one tick: its own position,
// the nest it remembers, and what is in the eight cells around it. It is a
// pointer-free value, so handing it to another goroutine is safe by
// construction.
type Sense struct {
	Pos      world.Position
	Nest     world.Position
	Food     int        // food on the ant's own cell
	Adjacent [8]int     // food in each of the eight neighbours
	Pher     [8]float64 // pheromone in each of the eight neighbours
}

// senseAt builds a Sense by reading the world. Only code that owns the
// world may call it.
func senseAt(w *world.World, pos world.Position) Sense {
	s := Sense{Pos: pos, Nest: w.Nest, Food: w.Food.At(pos)}
	for i, d := range directions[:8] {
		n := pos.Add(d)
		s.Adjacent[i] = w.Food.At(n)
		s.Pher[i] = w.Pher.At(n)
	}
	return s
}

[8]int is an array, not a slice, and that distinction is the entire point: an array is stored inline in the struct, so copying a Sense copies the data. A []int would copy a pointer, and two goroutines would share the elements. When designing a message type, arrays are safe and slices are not. (The Pher field is for Milestone 6; it costs nothing to add now.)

internal/sim/engine.go — the messages

type reqKind int

const (
	reqSense reqKind = iota
	reqAct
	reqStats
	reqEvaporate
	reqLost
)

// request is a message from anyone to the world owner. Every field is a
// value: no pointers cross the channel, so nothing is shared.
type request struct {
	kind     reqKind
	id       int
	pos      world.Position
	carrying bool
	act      Action
	reply    chan response
}

// response is the owner's answer, also pointer-free.
type response struct {
	sense Sense
	pos   world.Position
	ok    bool
	err   error
	stats Stats
}

One request type with a kind tag rather than five separate channels. Both designs are idiomatic; the tagged union is easier to extend and keeps ordering between different request kinds, which will matter when evaporation competes with ant traffic. The cost is a switch and some unused fields, which is Go's usual price for not having sum types.

The engine and its owner loop

// Engine runs the colony as one owner goroutine plus one goroutine per ant.
// No mutex appears anywhere in this file.
type Engine struct {
	cfg  Config
	b    Behaviour
	reqs chan request

	// nest never changes after construction, so every goroutine may read it.
	nest world.Position

	// Owned exclusively by the serve goroutine. Nothing else may touch
	// these, which is why they need no synchronisation at all.
	world         *world.World
	rng           *rand.Rand
	tick          int
	delivered     int
	carrying      int
	failedPickups int
	lost          int

	// Shared counters, written from many goroutines.
	restarts atomic.Int64
	dropped  atomic.Int64
	timeouts atomic.Int64
}

// serve is the only goroutine allowed to read or write the world.
func (e *Engine) serve(ctx context.Context) {
	for {
		select {
		case <-ctx.Done():
			return
		case r := <-e.reqs:
			e.handle(r)
		}
	}
}

func (e *Engine) handle(r request) {
	switch r.kind {
	case reqSense:
		r.reply <- response{sense: senseAt(e.world, r.pos)}
	case reqAct:
		r.reply <- e.applyAction(r)
	case reqStats:
		r.reply <- response{stats: e.statsNow()}
	case reqEvaporate:
		e.world.Pher.Evaporate(e.cfg.Evaporation)
		e.tick++
	case reqLost:
		e.lost++
		e.carrying--
	}
}

Read the comments on the struct fields as if they were enforced, because they are the only enforcement there is. Go cannot express "this field may only be touched by that goroutine", so the convention is a comment plus discipline plus go test -race in CI. Grouping the owner-only fields together, with one comment covering the block, makes a violation visible in review.

nest is duplicated out of the world deliberately. It never changes after construction, so it is safe for anyone to read, and having it available outside the owner avoids a message round trip in the supervisor later. Immutable-after-construction is a third category alongside "owned" and "synchronised", and it is worth naming in comments when you use it.

Applying an action

func (e *Engine) applyAction(r request) response {
	switch r.act.Kind {
	case ActMove:
		np := e.world.Clamp(r.pos.Add(r.act.Dir))
		if r.carrying {
			e.world.Pher.AddAt(np, e.cfg.Deposit, e.cfg.PherMax)
		}
		return response{pos: np, ok: true}

	case ActPickUp:
		if err := e.world.TakeFood(r.pos); err != nil {
			e.failedPickups++
			return response{pos: r.pos, err: err}
		}
		e.carrying++
		return response{pos: r.pos, ok: true}

	case ActDrop:
		if r.pos != e.world.Nest {
			return response{pos: r.pos, err: fmt.Errorf("ant %d drop away from nest: %w", r.id, world.ErrNotCarrying)}
		}
		e.world.Deliver()
		e.carrying--
		return response{pos: r.pos, ok: true}
	}
	return response{pos: r.pos, ok: true}
}

Note what the owner does not do: it never writes to an Ant. It computes the consequence and returns it. The ant updates itself. This keeps a clean rule that is easy to check by reading: the owner writes the world, each ant writes itself, and nothing writes both.

The owner also validates. An ant that asks to drop food away from the nest gets an error rather than a delivery. A behaviour is untrusted input, in the same way a client request is untrusted input to a server, and for the same reason: in Milestone 8 those requests start arriving corrupted.

The ant side

// antClient is one ant's private connection to the owner: its own reply
// channel and its own random source, touched by no other goroutine.
type antClient struct {
	e     *Engine
	ant   *Ant
	reply chan response
	rng   *rand.Rand
}

func (c *antClient) roundTrip(ctx context.Context, r request) (response, error) {
	// A reply from a request we gave up on may still be sitting here.
	// Throw it away before asking again.
	select {
	case <-c.reply:
	default:
	}

	r.reply = c.reply

	select {
	case c.e.reqs <- r:
	case <-ctx.Done():
		return response{}, ctx.Err()
	}

	timer := time.NewTimer(c.e.cfg.RequestTimeout)
	defer timer.Stop()

	select {
	case resp := <-c.reply:
		return resp, nil
	case <-timer.C:
		c.e.timeouts.Add(1)
		return response{}, ErrTimeout
	case <-ctx.Done():
		return response{}, ctx.Err()
	}
}

Explanation

This function is twenty lines and every one of them is load-bearing.

The ant's life

// runAnt is the whole life of one ant: sense, decide, act, repeat.
func (e *Engine) runAnt(ctx context.Context, a *Ant) {
	c := &antClient{
		e:     e,
		ant:   a,
		reply: make(chan response, 1),
		rng:   rand.New(rand.NewPCG(e.cfg.Seed, uint64(a.ID)+a.Generation<<32)),
	}

	for {
		if ctx.Err() != nil {
			return
		}

		sensed, err := c.roundTrip(ctx, request{kind: reqSense, id: a.ID, pos: a.Pos})
		if err != nil {
			if errors.Is(err, ErrTimeout) {
				continue // ask again
			}
			return // context cancelled
		}

		act := e.b.Decide(a, sensed.sense, c.rng)

		done, err := c.roundTrip(ctx, request{
			kind: reqAct, id: a.ID, pos: a.Pos, carrying: a.Carrying, act: act,
		})
		if err != nil {
			if errors.Is(err, ErrTimeout) {
				continue
			}
			return
		}
		c.applyLocally(act, done)
	}
}

// applyLocally updates the ant's own state from the owner's answer. The ant
// goroutine is the only writer of these fields.
func (c *antClient) applyLocally(act Action, r response) {
	switch act.Kind {
	case ActMove:
		if r.ok {
			c.ant.Pos = r.pos
			c.ant.Energy--
		}
	case ActPickUp:
		if r.ok {
			c.ant.Carrying = true
		}
	case ActDrop:
		if r.ok {
			c.ant.Carrying = false
		}
	}
}

Sense, decide, act. The decide step is the only part that runs outside the owner, and it is the only part that runs in parallel across ants. Remember that when we discuss performance in a moment.

Reading state: ask the owner

// Stats asks the owner for a snapshot. Safe to call from any goroutine.
func (e *Engine) Stats(ctx context.Context) (Stats, error) {
	reply := make(chan response, 1)
	select {
	case e.reqs <- request{kind: reqStats, reply: reply}:
	case <-ctx.Done():
		return Stats{}, ctx.Err()
	}
	timer := time.NewTimer(e.cfg.RequestTimeout)
	defer timer.Stop()
	select {
	case r := <-reply:
		return r.stats, nil
	case <-timer.C:
		return Stats{}, ErrTimeout
	case <-ctx.Done():
		return Stats{}, ctx.Err()
	}
}

Compare with Milestone 4's solution, where Stats took a lock and we had to invent the statsLocked convention to avoid deadlocking against ourselves. Here the question "what if somebody calls Stats from inside the owner?" cannot arise, because the owner never calls Stats; it calls statsNow, which is an ordinary function on data it owns. Reentrancy stops being a hazard when there is no lock to re-enter.

Here a fresh channel per call is fine, because Stats is called a few times a second, not a few million. Reuse is an optimisation, and optimisations belong where the traffic is.

Does it work, and is it faster?

$ go test -race -run TestEngine -count=1 ./internal/sim/
ok  	github.com/yourname/antfarm/internal/sim	1.965s

Correct, race-free, and no mutex in the package. Now the honest part.

The single owner is still a serialisation point

Every world access in the colony passes through one goroutine, so world operations happen strictly one at a time, exactly as they did behind the single mutex. If you were expecting linear speedup from replacing the lock, you will not get it, and anyone claiming otherwise is describing a different program.

What actually changed:

  • Decide now runs in parallel across all ants, off the owner's goroutine. In this project that computation is small; in a simulation where agents do real work (pathfinding, learning, inference) it is most of the cost, and this structure parallelises all of it.
  • The queue is now explicit. reqs chan request with a capacity is a visible, measurable, tunable buffer. Under a mutex the queue exists too, inside the runtime, where you cannot see its depth or apply a policy when it fills. Milestone 11 turns that visibility into backpressure.
  • Failure became expressible. "Drop this message", "answer this one slowly", "stop answering" are one line each at the owner. There is no way to express a dropped message in a mutex.
  • The protocol is now a protocol. Values in, values out, no shared memory. Milestone 12 moves the owner to another process and the ant code does not change.

If raw throughput on one machine were the only goal, the right answer is neither locks nor a single owner: it is sharding, several owners each holding a region of the grid, which Milestone 11 builds. The single owner is the design you should reach for first because it is obviously correct, and shard it only when measurement says you must.

Shared state vs message passing
Shared state + locks              Single owner + messages
────────────────────              ───────────────────────
any goroutine may touch data      one goroutine may touch data
correctness = lock discipline     correctness = structure
invisible queueing in the runtime an explicit channel you can measure
deadlock is a real risk           deadlock needs a request cycle
reentrancy hazards                none: no lock to re-enter
cannot express "message lost"     one line at the owner
does not survive a network        already is a network protocol

Go supports both and takes no side. The proverb "share memory by communicating" is advice, not a rule, and the standard library uses mutexes heavily where they fit. What Go provides that most languages do not is that both options are first-class and cheap enough to choose between on the merits.

Exercise 5

Add a reqTeleport request that moves an ant to an arbitrary position, and expose it as func (e *Engine) Teleport(ctx context.Context, id int, to world.Position) error so an external caller (the future dashboard) can poke the simulation.

Requirements: reject out-of-bounds targets with a useful error; the ant's own goroutine must end up agreeing with the owner about where it is; no new mutex; and a test that calls Teleport concurrently with a running colony and passes under -race.

The interesting part is the second requirement. The owner does not write to ants, and the ant is asleep in its own loop. Think about who can change a.Pos and how the news reaches them.

Solution 5 — open after trying

The trap is to have Teleport write a.Pos directly. It compiles, and it is a data race against the ant's goroutine, which writes the same field in applyLocally.

The clean answer is a mailbox per ant: the owner records a pending command, and the ant collects it on its next sense round trip. One extra field on response, no new synchronisation:

// in Engine, owner-only state:
//     pending map[int]world.Position   // ant ID -> forced position

func (e *Engine) handle(r request) {
	switch r.kind {
	case reqSense:
		resp := response{sense: senseAt(e.world, r.pos)}
		if to, ok := e.pending[r.id]; ok {
			delete(e.pending, r.id)
			resp.teleport, resp.hasTeleport = to, true
		}
		r.reply <- resp

	case reqTeleport:
		if !e.world.InBounds(r.pos) {
			r.reply <- response{err: &world.OutOfBoundsError{Pos: r.pos, W: e.cfg.Width, H: e.cfg.Height}}
			return
		}
		e.pending[r.id] = r.pos
		r.reply <- response{ok: true}
	}
}
// in runAnt, right after a successful sense:
if sensed.hasTeleport {
	a.Pos = sensed.teleport
	continue // re-sense from the new position before deciding
}

Three things to take from this:

  • The pending map is owner-only state, so a map is perfectly safe here. The thing that made maps dangerous in Milestone 4 was sharing, not maps.
  • Commands flow to the ant by riding on a reply the ant was already going to ask for. No new channel, no new goroutine, no polling. This is a standard trick in actor-style systems: piggyback control messages on the existing conversation.
  • The continue matters. Deciding on a Sense gathered at the old position after teleporting would make the ant act on a view of somewhere it no longer is. Stale reads are the recurring theme of this milestone.

Experiment

Set QueueSize to 1 and run 2,000 ants. Then set it to 65536. Watch delivered-per-second and the timeout counter. A tiny queue means senders block constantly, which is slow but self-limiting; a huge queue means nobody ever blocks, latency grows without bound, and ants time out on replies to requests that are still sitting in line. Neither extreme is good, and the middle is not obvious. That tension is what Milestone 11 is about.

Common mistakes in Milestone 5
  • The Mewlang cat, giving an unimpressed side-eyeAn unbuffered reply channel. The owner blocks forever on r.reply <- resp if the ant has stopped listening. The whole simulation freezes and the stack dump shows every goroutine blocked on the same channel. Buffer of one, always.
  • Forgetting the stale-reply drain. Symptom: after the first timeout, an ant starts behaving as if it is one step behind reality. Very hard to spot without the counter.
  • Putting a pointer in a message. Compiles, passes tests, races under load. If you catch yourself writing ant *Ant in the request struct, stop.
  • Closing reqs to signal shutdown. Senders panic with "send on closed channel", and there are thousands of them. Only close a channel when there is exactly one sender, and here there are many. Cancel a context instead.
  • Calling Stats from inside handle. The owner would send itself a request and wait for a reply it can only produce by returning. Instant, permanent, self-inflicted deadlock. This is the message-passing equivalent of non-reentrant locks, and the same discipline applies: internal code calls the internal function, not the public one.

Checkpoint

  1. Why is the reply channel buffered, and what exactly breaks if it is not?
  2. Why does Sense use [8]int rather than []int?
  3. Who is allowed to write Ant.Pos, and how would you catch a violation?
  4. The owner serialises world access just like the mutex did. Name three things we gained anyway.
  5. Why time.NewTimer plus Stop instead of time.After in roundTrip?
  6. What happens if an ant times out, then the owner replies, and the ant does not drain its channel?
Why are we using this language here?

The single-owner pattern is not a Go idea — it is the actor model, and you can write it in Erlang, Elixir or Akka too. What Go contributes is that it is cheap here: 1,000 ant goroutines plus one owner goroutine plus 1,000 buffered channels cost a few megabytes and start in microseconds. The same shape in a language with OS-thread-weight concurrency would force you toward a thread pool and a work queue long before you reached this ant count, which is a different, more complicated program.

The honest cost: nothing in the language stops you from breaking the single-owner rule. The comment on the struct fields is the entire enforcement mechanism, and it is enforced by discipline plus go test -race, not by the compiler. Erlang and Rust each take this further in their own way — Erlang by making every process's memory genuinely private, Rust by making a smuggled pointer a compile error rather than a comment. Go trusts you, which is fast to write and easy to get wrong once a codebase has more than one author.

Milestone 6Pheromones, and a clock of its own

Goal

Ants that find food lay a trail on the way home; searching ants are attracted to trails; trails evaporate. The colony gets measurably better at foraging, and we gain a second background goroutine with its own timing.

Concepts

time.Ticker and periodic work, non-blocking sends as load shedding, weighted random choice, float grids, and why every positive feedback loop needs a cap and a decay.

Design

Ant colony optimisation in three rules:

  1. An ant carrying food deposits pheromone on each cell it walks over.
  2. A searching ant picks its next cell with probability proportional to 1 + bias × pheromone.
  3. All pheromone decays by a constant factor on a fixed schedule.

Rule 3 is not decoration. Without evaporation the grid saturates: every cell eventually reaches the cap, all weights become equal, and the trail carries no information. Worse, an early lucky path to an exhausted food source stays attractive forever. Evaporation is what makes the trail a memory of recent success rather than an accumulation of all history. Rule 2's 1 + is the same idea from the other side: it guarantees an unmarked cell still has a chance, so the colony keeps exploring instead of collapsing onto the first path it finds. Positive feedback plus a decay plus a floor on exploration, and this is exactly the structure of many reinforcement algorithms.

Where does evaporation run? It must mutate the grid, and only the owner may do that, so it becomes a third kind of client: a goroutine holding a ticker that sends reqEvaporate on a schedule.

   ants ────────►┐
                 │  reqs chan  ──►  serve()  (owns world)
   evaporator ──►┤
   Stats caller ►┘
Typical approach vs Go's idiomatic approach

A typical periodic job — a cron entry, a Quartz scheduler, a setInterval — assumes its work will run to completion every time it fires, and treats a run that is still busy when the next tick arrives as an operational problem to fix (usually by making the job faster or the interval longer). The scheduler and the work are separate concerns, and a missed tick is unusual enough to page someone.

The idiomatic Go version here treats a missed tick as the normal case under load: evaporate tries a non-blocking send and, if the owner's queue is full, simply skips that cycle. No error, no retry, no alert. This works because a time.Ticker paired with select and a default case turns "is anyone listening right now" into a two-line question you can ask on every tick, which most schedulers do not expose at all. The cost is that this policy is only correct because evaporation is idempotent and its value decays with delay; a periodic job with side effects that must not be skipped (say, a billing cycle) needs the traditional "queue it and guarantee it runs" discipline instead.

Implementation

internal/world/pheromone.go

// FloatGrid is a grid of float64 values, used for pheromone concentration.
type FloatGrid struct {
	W, H  int
	cells []float64
}

// AddAt deposits delta at p, capping the cell so a single trail cannot
// dominate the whole grid.
func (g *FloatGrid) AddAt(p Position, delta, max float64) {
	if !g.InBounds(p) {
		return
	}
	i := p.Y*g.W + p.X
	g.cells[i] += delta
	if g.cells[i] > max {
		g.cells[i] = max
	}
}

// Evaporate multiplies every cell by factor (0 < factor < 1) and zeroes
// values too small to matter, which keeps the grid sparse in practice.
func (g *FloatGrid) Evaporate(factor float64) {
	for i, v := range g.cells {
		v *= factor
		if v < 1e-4 {
			v = 0
		}
		g.cells[i] = v
	}
}

The 1e-4 floor is worth a sentence. Multiplying by 0.95 forever produces ever-smaller denormal floats, which are both meaningless and, on some hardware, dramatically slower to compute with. Snapping to zero keeps the arithmetic fast and the data honest. Whenever you write an exponential decay, decide where it stops.

Evaporate is O(width × height), so on a 512×512 grid it touches 262,144 cells every time it runs. That is a real cost sitting on the owner's goroutine, blocking every ant behind it. Remember this when you profile in Milestone 11; the fix is either a coarser schedule or a lazy scheme that decays a cell when it is read, using the timestamp of its last update.

The evaporator goroutine

// evaporate asks the owner to decay the pheromone grid on a fixed schedule.
func (e *Engine) evaporate(ctx context.Context) {
	ticker := time.NewTicker(e.cfg.EvaporateEvery)
	defer ticker.Stop()

	for {
		select {
		case <-ctx.Done():
			return
		case <-ticker.C:
			select {
			case e.reqs <- request{kind: reqEvaporate}:
			case <-ctx.Done():
				return
			default:
				// The owner is saturated. Skipping one evaporation is
				// better than stalling the clock behind the ants.
			}
		}
	}
}

Explanation

The trail-following behaviour

// TrailFollower is Forager plus pheromones: when searching it biases its
// steps toward neighbouring cells that other ants have marked.
type TrailFollower struct {
	// Bias is how strongly pheromone attracts. 0 is a random walk.
	Bias float64
}

func (t TrailFollower) Decide(a *Ant, s Sense, rng *rand.Rand) Action {
	if a.Carrying {
		if s.Pos == s.Nest {
			return Action{Kind: ActDrop}
		}
		return Action{Kind: ActMove, Dir: stepToward(s.Pos, s.Nest)}
	}
	if s.Food > 0 {
		return Action{Kind: ActPickUp}
	}
	if i, ok := bestAdjacentFood(s); ok {
		return Action{Kind: ActMove, Dir: directions[i]}
	}
	return Action{Kind: ActMove, Dir: t.weightedStep(s, rng)}
}

// weightedStep picks a neighbour with probability proportional to
// 1 + Bias*pheromone, so an unmarked grid still produces a random walk.
func (t TrailFollower) weightedStep(s Sense, rng *rand.Rand) world.Position {
	var weights [8]float64
	total := 0.0
	for i, p := range s.Pher {
		w := 1 + t.Bias*p
		weights[i] = w
		total += w
	}

	r := rng.Float64() * total
	for i, w := range weights {
		r -= w
		if r <= 0 {
			return directions[i]
		}
	}
	return directions[rng.IntN(8)] // unreachable except for float rounding
}

Weighted random selection by subtracting from a uniform sample in [0, total) is worth knowing: it needs no sorting, no allocation, and one pass. The final return after the loop looks like dead code and is not: floating-point rounding can leave r a hair above zero after subtracting every weight. Reviewers will ask about that line, so the comment is part of the code.

TrailFollower has a value receiver and one field, so it is copied into the interface rather than shared. Every ant reads the same Bias and nothing writes it. Compare with the Scout from Milestone 3, whose per-ant maps would have been a race: the difference is that this strategy keeps no per-ant state at all.

Deposition lives at the owner, in applyAction, the three lines you already saw:

	case ActMove:
		np := e.world.Clamp(r.pos.Add(r.act.Dir))
		if r.carrying {
			e.world.Pher.AddAt(np, e.cfg.Deposit, e.cfg.PherMax)
		}
		return response{pos: np, ok: true}

Does it actually help?

Same seed, same world, same wall-clock budget, 300 ants on a 64×64 grid with six food sources of 200 units each:

forager: delivered 323, pheromone left 1693
trail  : delivered 404, pheromone left 397

The Mewlang cat, beaming with delightA 25% improvement in delivered food, measured rather than asserted. Two details in those numbers are more interesting than the headline.

The forager run has more pheromone left at the end (1693 against 397) even though it ignores pheromone entirely. Deposition happens at the owner for any carrying ant, so both runs lay trails; only one reads them. The forager's trails accumulate because its ants take longer random-walk journeys and are more spread out, while the trail follower's ants concentrate on short paths that evaporation keeps trimmed. A metric moving in the direction you did not predict is usually the most informative thing on the screen.

Second: this is a measurement of one seed on one machine over half a second, which is not a result, it is an anecdote. Before claiming pheromones help, you would run several seeds and compare distributions. That is Exercise 6.

Exercise 6

Turn the anecdote into evidence, then find where the benefit lives.

  1. Write a test that runs both behaviours over at least 10 seeds each and reports mean delivered food. Fail only if the trail follower is not better on a clear majority of seeds. Think about why "better on average" and "better on most seeds" are different claims and which one you want.
  2. Sweep Bias over 0, 1, 2, 4, 8, 16 and plot delivered food. There is an optimum, and beyond it the colony gets worse. Explain why in terms of exploration and exploitation.
  3. Sweep Evaporation over 0.99, 0.95, 0.8, 0.5 with Bias fixed. Explain both ends of the resulting curve.
Solution 6 — open after trying
func TestTrailBeatsRandomWalk(t *testing.T) {
	const runs = 10
	wins, forageTotal, trailTotal := 0, 0, 0

	for seed := range uint64(runs) {
		cfg := Config{Width: 64, Height: 64, Ants: 300,
			FoodSources: 6, FoodPerSource: 200, Seed: seed}

		a := runFor(t, cfg, Forager{}, 300*time.Millisecond)
		b := runFor(t, cfg, TrailFollower{Bias: 4}, 300*time.Millisecond)

		forageTotal += a.Delivered
		trailTotal += b.Delivered
		if b.Delivered > a.Delivered {
			wins++
		}
	}

	t.Logf("forager mean %d, trail mean %d, trail won %d/%d",
		forageTotal/runs, trailTotal/runs, wins, runs)

	if wins < runs*7/10 {
		t.Errorf("trail following won only %d of %d seeds", wins, runs)
	}
}

Mean versus majority. A mean can be carried by one spectacular run: if the trail follower occasionally finds a rich source early and delivers ten times as much, the average looks wonderful while it loses most individual races. Counting wins asks "is this reliably better?", which is usually the question you care about. Reporting both, as the t.Logf does, costs nothing.

Note also that this test is time-based, so it measures throughput on whatever machine runs it, and it will be flaky on a loaded CI box. A more rigorous version fixes the number of ticks rather than the wall clock, which is a good argument for keeping the sequential Sim from Milestone 3 around as a measurement harness. Deterministic experiments belong in a deterministic runner.

The bias sweep. You should see improvement up to roughly 4–8, then decline. Low bias is a random walk that ignores information. High bias makes the first trail overwhelmingly attractive, so every ant follows it, deposits more on it, and makes it more attractive still: the colony locks onto one path and stops exploring, and when that food source runs out it keeps walking the trail to an empty cell until evaporation erases it. This is the exploration/exploitation trade-off, and it is the same failure mode as a recommender system that only ever shows you what you already clicked.

The evaporation sweep. At 0.99 trails persist far too long, so stale paths to exhausted sources keep attracting ants. At 0.5 a trail is gone within a few ticks and cannot be reinforced fast enough to guide anyone, so you have paid for pheromones and got a random walk. The useful range is narrow and depends on how long a round trip takes, which is a hint that the parameter should really be expressed relative to journey time rather than as an absolute per-tick factor.

Experiment

Set PherMax to a huge number (say 1e9) and run with Bias: 4. The first trail to form saturates, and because weights are proportional to concentration, the colony funnels every ant onto one path within seconds. Watch delivered food collapse. Then set PherMax to 1 and watch the trail carry almost no signal. Every positive feedback loop in a simulation needs a ceiling, and finding it experimentally teaches more than choosing it analytically.

Common mistakes in Milestone 6
  • Ticker never stopped. time.NewTicker without defer Stop() leaks the ticker and its goroutine. The leak test in Milestone 7 catches it.
  • Blocking send from the evaporator. Without the default, evaporation queues behind ant traffic, the schedule slips under load, and trails stop decaying exactly when the colony is busiest.
  • Evaporating from the evaporator goroutine directly. It compiles. It is a data race against the owner. The whole point of the message is that the mutation happens on the owner's goroutine.
  • Forgetting the 1 + in the weight. With weights equal to raw pheromone, a fresh grid has total weight zero, every weight is zero, and the loop falls through to the fallback on every single step. Symptom: the colony behaves randomly and no amount of bias tuning changes anything.
  • Depositing on pickup instead of on the journey home. A single marked cell at the food source is not a trail. The trail is the path, and it only exists because the ant marks every step of the return.

Checkpoint

  1. Why must pheromone evaporate? Give two distinct failure modes of a grid that never decays.
  2. Why does weightedStep add 1 to every weight?
  3. What does the default case in the evaporator's inner select do, and what policy does it encode?
  4. Both behaviours deposit pheromone. Why did the run that ignores trails end with more of it?
  5. Evaporate is O(cells) and runs on the owner's goroutine. What does that imply at 512×512 with a 20 ms period?
  6. Why is TrailFollower safe to share across 50,000 ants when Scout from Milestone 3 was not?
Why are we using this language here?

Nothing about pheromones needed Go specifically — FloatGrid.Evaporate and weightedStep are ordinary numeric code you could write identically in Python or Java. What Go earns its keep on here is the plumbing around that code: a time.Ticker plus a non-blocking select gives the evaporator goroutine a load-shedding policy in four lines, and because it is just another client of the same owner, it needed no new synchronisation primitive at all, only a new message kind.

The honest verdict: this milestone's actual algorithm (weighted random choice with a decay term) is the interesting part, and Go contributed nothing to that part beyond what any language with closures and a decent standard library would. Where a language choice would matter more is at much larger scale — millions of cells decaying continuously — where you would want either a lazy, timestamp-based decay (no periodic sweep at all) or a language with SIMD-friendly array operations built in, neither of which this milestone needed yet.

Milestone 7Shutting down, provably

Goal

Ctrl-C stops the colony cleanly: every ant finishes its current round trip, the final statistics are collected, background goroutines exit, and the process leaves nothing running. A test proves there are no leaks.

Concepts

Context trees and derived cancellation, signal.NotifyContext, shutdown ordering, the WaitGroup misuse that bites concurrent supervisors, and goroutine leak detection.

Design

The Mewlang cat, in profile, thinking it over"Just cancel everything" is wrong, and the reason is worth spelling out. The last thing Run does is collect final statistics, which requires sending a request to the owner and getting a reply. If the owner was cancelled along with everyone else, that request hangs until it times out and you get empty results. So shutdown has an order, and the order is the reverse of the dependency graph:

  ctx cancelled (Ctrl-C, timeout, or the caller's choice)
        │
        ▼
  1. supervisor stops          no new ants will be started
        │
        ▼
  2. ants drain and exit       each finishes its current round trip
        │
        ▼
  3. final Stats collected     the owner is still alive to answer
        │
        ▼
  4. evaporator stops          nothing left that needs a clock
        │
        ▼
  5. owner stops               last to go, because everyone needed it

This is why Run builds three contexts rather than passing one everywhere. The ants get the caller's context, so cancelling it stops them. The evaporator and the owner get their own contexts derived from context.Background(), so they outlive the ants and stop only when Run says so. A context is a cancellation scope, and "everything cancels at once" is a design decision, not a default you have to accept.

Typical approach vs Go's idiomatic approach

A typical shutdown path in many services is a single flag or a single cancellation signal that everything listens to, on the theory that "stop" should mean stop, everywhere, at once. It is simple to reason about and it is exactly wrong the moment cleanup itself needs something that "everything" includes — here, the final Stats call needs the owner, and the owner would already be gone.

Go's context package makes the idiomatic fix cheap: contexts form a tree, so Run derives three separate cancellation scopes (ants, evaporator, owner) instead of sharing one, and shuts them down in the order the dependency graph actually requires. The cost is that you now have to think about that order explicitly — three context.WithCancel calls and a comment explaining why, instead of one flag everyone reads. For a program with no cleanup dependencies, one shared context would have been simpler and just as correct; the extra structure here is earned by "the last thing we do needs something we are shutting down."

Implementation

// Run starts the colony and blocks until ctx is cancelled, then shuts
// everything down in order and returns the final snapshot.
func (e *Engine) Run(ctx context.Context) (Stats, error) {
	serverCtx, stopServer := context.WithCancel(context.Background())
	var serverWG sync.WaitGroup
	serverWG.Add(1)
	go func() {
		defer serverWG.Done()
		e.serve(serverCtx)
	}()

	evapCtx, stopEvap := context.WithCancel(context.Background())
	var evapWG sync.WaitGroup
	evapWG.Add(1)
	go func() {
		defer evapWG.Done()
		e.evaporate(evapCtx)
	}()

	e.ants = make([]*Ant, 0, e.cfg.Ants)
	for i := range e.cfg.Ants {
		e.ants = append(e.ants, &Ant{ID: i, Pos: e.nest, Energy: e.cfg.StartEnergy})
	}

	deaths := make(chan *Ant, e.cfg.Ants)
	var antWG sync.WaitGroup

	start := func(a *Ant) {
		antWG.Add(1)
		go func() {
			defer antWG.Done()
			e.superviseAnt(ctx, a, deaths)
		}()
	}
	for _, a := range e.ants {
		start(a)
	}

	// The supervisor is the only goroutine that may add more ants, so we
	// wait for it to finish before waiting on the ant group.
	supervisorDone := make(chan struct{})
	go func() {
		defer close(supervisorDone)
		for {
			select {
			case <-ctx.Done():
				return
			case a := <-deaths:
				e.recycle(ctx, a)
				start(a)
			}
		}
	}()

	<-ctx.Done()
	<-supervisorDone
	antWG.Wait()

	stats, err := e.Stats(context.Background())

	stopEvap()
	evapWG.Wait()
	stopServer()
	serverWG.Wait()

	return stats, err
}

The three lines in the middle

	<-ctx.Done()
	<-supervisorDone
	antWG.Wait()

Those look redundant and are not. sync.WaitGroup has a documented rule that is easy to violate: calls to Add must not happen concurrently with Wait when the counter could be zero. Our supervisor calls antWG.Add(1) every time it restarts an ant. If Run called antWG.Wait() while the supervisor was still alive, and every ant happened to be dead at that instant, Wait could return while the supervisor was starting a fresh one, and we would shut down the owner from under a running ant. Waiting for the supervisor to exit first guarantees no further Add can occur. The Go runtime detects some violations with panic: sync: WaitGroup misuse: Add called concurrently with Wait, but not all of them, and a race here is a shutdown hang that reproduces once a week.

Then e.Stats(context.Background()) deliberately uses a fresh context. Passing the caller's cancelled ctx would return context.Canceled immediately and give you an empty snapshot. Reaching for Background during cleanup is a common and correct pattern; in production code you would normally give it a short timeout so cleanup cannot hang forever.

Catching Ctrl-C

	// Ctrl-C cancels the context; a second one kills the process outright.
	ctx, stop := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM)
	defer stop()
	ctx, cancel := context.WithTimeout(ctx, duration)
	defer cancel()

signal.NotifyContext returns a context that cancels when one of the listed signals arrives. The nesting matters: the timeout context is derived from the signal context, so the run ends on whichever comes first, a signal or the deadline. That is context composition doing real work in four lines.

defer stop() restores the default signal handling, which is what makes a second Ctrl-C kill the process immediately. That is the behaviour users expect: the first interrupt asks nicely, the second insists. A program that swallows every interrupt while it "gracefully" shuts down is a program people learn to kill -9.

Proving there is no leak

func TestNoGoroutineLeak(t *testing.T) {
	before := countGoroutines()

	cfg := Config{Width: 16, Height: 16, Ants: 100, FoodSources: 5, Seed: 3}
	runFor(t, cfg, Forager{}, 150*time.Millisecond)

	// Goroutines exit asynchronously; give them a moment before judging.
	for range 50 {
		if countGoroutines() <= before+1 {
			return
		}
		time.Sleep(10 * time.Millisecond)
	}
	t.Errorf("goroutines leaked: before %d, after %d", before, countGoroutines())
}
=== RUN   TestNoGoroutineLeak
--- PASS: TestNoGoroutineLeak (0.15s)

Explanation

The retry loop is the part people get wrong. runtime.NumGoroutine() is a snapshot, and a goroutine that has been told to stop takes an unspecified moment to actually stop, so a single comparison immediately after Run returns is flaky in both directions. Polling with a deadline converts "eventually true" into a test that passes quickly when correct and fails reliably when not. The +1 tolerance covers the testing framework's own bookkeeping.

A goroutine leak is Go's version of a memory leak, and it is worse in one respect: a leaked goroutine holds everything it references, so one leaked ant keeps its world, its channels and its buffers alive forever. In a long-running service the symptom is memory growth that no heap profile explains until you look at the goroutine profile. Write this test early, run it in CI, and reach for go.uber.org/goleak when you want the industrial version, which names the leaked stacks instead of just counting.

A note on errgroup

Production Go frequently uses golang.org/x/sync/errgroup, which combines a WaitGroup, error propagation and context cancellation:

g, ctx := errgroup.WithContext(ctx)
g.Go(func() error { return e.serve(ctx) })
g.Go(func() error { return e.evaporate(ctx) })
err := g.Wait()   // first non-nil error, and ctx is cancelled for the rest

It is genuinely good and you should know it. I have kept this project dependency-free so far, partly to avoid hiding behaviour behind a library before you have implemented it yourself, and partly because errgroup's "cancel everyone on the first error" semantics are the opposite of what this shutdown sequence needs. Add it with go get golang.org/x/sync when you want it; it is the one external dependency I would not argue about.

Exercise 7

Add a shutdown deadline. If the ants have not all exited within cfg.ShutdownGrace (default two seconds), Run should log how many are still running, collect statistics anyway, stop the owner, and return a non-nil error that the caller can distinguish from a normal stop. Requirements: no goroutine may be left blocked on a channel send; errors.Is must work against your new sentinel; the leak test must still pass in the happy path; and add a test that forces the slow path by setting a large Chaos.SlowFor.

Hint: antWG.Wait() blocks and has no timeout. You cannot add one to a WaitGroup, so you need to convert "the group finished" into something select can wait on.

Solution 7 — open after trying
var ErrShutdownTimeout = errors.New("ants did not stop before the deadline")

// waitChan converts a WaitGroup into a channel that closes when the group
// reaches zero, so it can be used in a select.
func waitChan(wg *sync.WaitGroup) <-chan struct{} {
	done := make(chan struct{})
	go func() {
		wg.Wait()
		close(done)
	}()
	return done
}
	<-ctx.Done()
	<-supervisorDone

	var shutdownErr error
	select {
	case <-waitChan(&antWG):
		// clean stop
	case <-time.After(e.cfg.ShutdownGrace):
		shutdownErr = fmt.Errorf("%w after %v", ErrShutdownTimeout, e.cfg.ShutdownGrace)
		log.Printf("antfarm: %v; collecting stats anyway", shutdownErr)
	}

	stats, err := e.Stats(context.Background())
	if shutdownErr != nil {
		err = shutdownErr
	}

	stopEvap()
	evapWG.Wait()
	stopServer()
	serverWG.Wait()
	return stats, err

The waitChan helper is the standard answer to "I need a WaitGroup in a select". Note that it starts a goroutine which survives until the group drains, so on the timeout path it outlives Run. That is acceptable here (it will exit when the stragglers do) but it is exactly the kind of thing your leak test will notice, which is why the exercise asks you to force the slow path deliberately.

The deeper lesson: you cannot force a goroutine to stop. There is no goroutine.Kill() in Go, by design. If an ant is asleep inside time.Sleep(10 * time.Second), nothing in the language will wake it, and the best you can do is stop waiting and record that you did. Every graceful shutdown in Go is therefore cooperative, with a deadline and a report, and "how do I kill a stuck worker?" has exactly one answer: you do not, you design so that workers check for cancellation often enough that they never get stuck. Erlang's answer to this question is exit(Pid, kill), which is unconditional and immediate, and Course 4 will show what becomes possible when that primitive exists.

Experiment

Comment out defer ticker.Stop() in the evaporator and run the leak test. Then restore it and remove the <-supervisorDone line, and run the whole suite with -race -count=20. The first is a reliable failure; the second is an intermittent one, which is a good demonstration of why concurrency bugs need repeated runs rather than a single green tick.

Common mistakes in Milestone 7
  • One context for everything. Final statistics come back empty because the owner died with the ants.
  • WaitGroup.Add racing with Wait. Symptom: rare shutdown hangs or a runtime panic about WaitGroup misuse.
  • Calling os.Exit in a signal handler. Deferred functions do not run, buffered output is lost, and the graceful path you wrote never executes.
  • Waiting on a WaitGroup whose goroutines are blocked on a full channel nobody is draining. Classic shutdown deadlock: stop the consumer before the producers and everything wedges.
  • Assuming defer runs on panic in another goroutine. It runs in the panicking goroutine only. An unrecovered panic anywhere kills the whole process, deferred cleanup elsewhere included.
  • Trusting NumGoroutine immediately. Poll with a deadline.

Checkpoint

  1. Why does the owner get a context derived from Background rather than the caller's context?
  2. What exactly goes wrong if antWG.Wait() runs while the supervisor is still alive?
  3. Why does Stats use a fresh context during shutdown?
  4. What does defer stop() after signal.NotifyContext buy the user?
  5. Why must the leak test poll instead of comparing once?
  6. An ant is inside time.Sleep(30 * time.Second) when you cancel. What are your options?
Why are we using this language here?

context.Context is the part of Go that makes this milestone almost pleasant: cancellation that composes (a timeout derived from a signal-triggered cancellation, in two lines), propagates automatically through every function that accepts a ctx parameter, and is checked cooperatively rather than requiring a separate kill mechanism per goroutine kind. sync.WaitGroup and signal.NotifyContext round out a shutdown story that needs zero third-party dependencies.

The honest cost: every one of these primitives is cooperative, which is the same fact as "you cannot kill a goroutine" wearing different clothes. context gives you a clean way to ask every goroutine to stop; it gives you nothing if a goroutine does not check. Compare Erlang, where a supervisor can unconditionally terminate a process regardless of what it is doing, or a language with structured concurrency and cancellation points enforced by the runtime. Go's answer is disciplined convention, not a guarantee, and this milestone's entire second half exists because that distinction matters in practice, not just in theory.

Milestone 8Chaos: crashes, dropped messages, slow ants

Goal

The Mewlang cat, giving a mischievous winkBreak the colony on purpose and keep it running. Ants panic at random and a supervisor restarts them. The owner drops messages at random and clients time out and retry. Ants stall at random and the queue absorbs it. Every failure becomes a number you can watch.

Concepts

panic and recover across goroutine boundaries, supervision built by hand, restart state and generations, accounting for work lost in a crash, at-least-once versus at-most-once delivery, and the way timeouts interact with throughput.

Design

Three failures, each modelling something real:

InjectedModelsHandled by
crash=pA bug, an assertion failure, a nil dereference in one workerrecover + supervisor restart
drop=pA lost packet, a dropped RPC, a full queue upstreamclient timeout and retry
slow=pA GC pause, a slow disk, a noisy neighbourqueueing, and eventually backpressure

Go's failure model has one fact you must design around: an unrecovered panic in any goroutine terminates the entire process. Not just that goroutine. There is no supervisor in the runtime, no isolation between goroutines, and no way for one goroutine to catch another's panic. recover works only inside a deferred function in the panicking goroutine itself.

So supervision has to be built, and the shape is forced by that constraint:

  start(a) ──► goroutine ──► superviseAnt(ctx, a, deaths)
                                   │
                                   │  defer func(){ recover() } ← must be here
                                   ▼
                              runAnt(...)   ← panics here
                                   │
                          panic unwinds to the deferred func
                                   │
                                   ▼
                            deaths <- a          (a channel of dead ants)
                                   │
                                   ▼
                        supervisor: recycle(a); start(a)

The supervisor is a goroutine reading a channel of casualties. That is the entire mechanism, and it is about fifteen lines. Keep the number in mind when Course 4 shows you what Erlang gives you for free.

Implementation

Injecting the failures

// Chaos describes the failures we inject on purpose.
type Chaos struct {
	CrashProb float64       // per ant, per tick
	DropProb  float64       // per message, at the owner
	SlowProb  float64       // per ant, per tick
	SlowFor   time.Duration // how long a slow tick sleeps
}

In the ant loop, before each cycle:

		if e.cfg.Chaos.CrashProb > 0 && c.rng.Float64() < e.cfg.Chaos.CrashProb {
			panic(fmt.Sprintf("ant %d: simulated crash at %v", a.ID, a.Pos))
		}
		if e.cfg.Chaos.SlowProb > 0 && c.rng.Float64() < e.cfg.Chaos.SlowProb {
			time.Sleep(e.cfg.Chaos.SlowFor)
		}

And at the top of the owner's handle:

	// Chaos: pretend the message never arrived. No reply is sent, so the
	// caller will time out. This is a dropped packet, simulated.
	if e.cfg.Chaos.DropProb > 0 && e.rng.Float64() < e.cfg.Chaos.DropProb {
		e.dropped.Add(1)
		return
	}

Chaos uses each goroutine's own random source, so it inherits the same seeding discipline as everything else and a chaotic run is as reproducible as a chaotic run can be. The probability checks are guarded by > 0 so that a zero-chaos configuration does not pay for a random number on every tick, which matters in the hot loop.

Recovering and restarting

// superviseAnt runs one ant and turns a panic into a restart request.
func (e *Engine) superviseAnt(ctx context.Context, a *Ant, deaths chan<- *Ant) {
	defer func() {
		r := recover()
		if r == nil {
			return
		}
		e.restarts.Add(1)
		select {
		case deaths <- a:
		case <-ctx.Done():
		}
	}()
	e.runAnt(ctx, a)
}

Explanation

recover is not exception handling, and using it like exception handling is a mistake

It is tempting to wrap every goroutine in a recover and call the system robust. What you have actually built is a system that continues running in a state its author never considered, with a corrupted invariant and no report. The Go community's position, which I think is right, is that recover is for two things: a process boundary where a crash must not take down unrelated work (an HTTP handler, a plugin, one ant), and turning a panic into an error at a library's public edge. Everywhere else, let it crash and fix the bug.

The version here qualifies, barely, because an ant is a genuine isolation boundary and because we count and report every restart. Note what it still cannot do: if the owner goroutine panics, the process dies and nothing survives. That asymmetry (one shared component whose failure is fatal) is inherent to the single-owner design, and it is a real cost to weigh against everything the design bought us in Milestone 5.

Accounting for lost work

// recycle resets a crashed ant before it restarts. A crash loses whatever
// the ant was carrying, and the colony has to account for that.
func (e *Engine) recycle(ctx context.Context, a *Ant) {
	if a.Carrying {
		select {
		case e.reqs <- request{kind: reqLost}:
		case <-ctx.Done():
		}
	}
	a.Carrying = false
	a.Pos = e.nest
	a.Energy = e.cfg.StartEnergy
	a.Generation++
}

This function is the most interesting fifteen lines in the milestone, for a reason that is not obvious.

Who owns the ant here? The rule since Milestone 5 has been "the ant's own goroutine writes its fields, nobody else". Yet recycle runs on the supervisor's goroutine and writes four of them. That is legal, and it is legal for a specific reason: the ant's goroutine is dead. Ownership transferred at the moment the panic unwound, and it transfers back when start(a) launches the replacement. The send on deaths and the later go statement are both synchronisation events, so the Go memory model guarantees the supervisor sees the dead ant's final state and the new goroutine sees the supervisor's writes. Ownership that moves between goroutines, handed over by channel operations, is the mental model that makes this kind of code tractable, and it is the same idea Rust encodes in its type system.

The crash destroys a unit of food. The owner already counted that unit in carrying when the pickup succeeded. If we simply restarted the ant, the ledger would claim a unit is in transit that no ant is carrying, and the conservation test would drift. So the supervisor tells the owner: one unit lost. The test then checks

delivered + on the ground + carried + lost == the food we started with

and passes under crash chaos. In-flight work disappears when a worker dies, and if your metrics do not model that, they will quietly lie to you. That lesson costs real companies real money, and here it costs one integer.

Generation increments so a restarted ant does not replay the same random sequence (the seed is uint64(a.ID) + a.Generation<<32). Without it, an ant that crashed would repeat its exact previous life, including the crash, at the same point. Deterministic chaos that always kills the same ant in the same place is not chaos.

Running it

$ go run ./cmd/antfarm -ants 2000 -grid 96x96 -food 8 -food-per-source 400 \
      -duration 3s -every 1s -chaos crash=0.0005,drop=0.001
colony: 2000 ants, 96x96 grid, 8 food sources of 400, behaviour trail, chaos {CrashProb:0.0005 DropProb:0.001 SlowProb:0 SlowFor:0s}
tick 44    ants 2000  carrying 32   delivered 504   food left 2663  lost 1    restarts 186  dropped 735   timeouts 574   pher 850
tick 93    ants 2000  carrying 17   delivered 787   food left 2394  lost 2    restarts 372  dropped 1513  timeouts 1372  pher 284
final: tick 143   ants 2000  carrying 4    delivered 860   food left 2332  lost 4    restarts 571  dropped 2298  timeouts 2155  pher 1272

Real output. Two thousand ants, 571 crashes and restarts in three seconds, 2,298 dropped messages, and the colony delivered 860 units of food while all of that was happening. Four units were lost to crashes and the books balance.

Look at dropped 2298 against timeouts 2155. They should be nearly equal, since every dropped message costs its sender a timeout, and the small gap is messages dropped near the end whose senders were cancelled before their timers fired. When two counters that should track each other stop tracking each other, you have found something. Deliberately instrumenting both sides of a mechanism, rather than one, is what makes that possible.

The timeout trap

Here are two runs of the same configuration, differing only in drop probability, over the same wall-clock budget:

drop = 0.02 : dropped 1601, timeouts 1400, delivered 327
drop = 0.25 : dropped 1608, timeouts 1400, delivered  13

Twelve times the message loss, but the same number of dropped messages, and delivered food collapsed by 96%. Read that again, because it is not intuitive.

The explanation: with a 50 ms timeout, an ant that loses a message spends 50 ms doing nothing, which is hundreds of times longer than a successful round trip. At 2% loss the colony still spends most of its time working. At 25% loss, one message in four costs 50 ms, so almost all the time goes into waiting, total traffic falls through the floor, and the absolute number of dropped messages stays flat because there are so few messages left to drop.

The timeout, not the loss rate, is what destroyed throughput. This is a genuine and very common production failure: a service degrades slightly, clients time out, retries multiply the load, and the system collapses far faster than the original fault would suggest. The mitigations are all things this code does not do yet, and they are worth naming: exponential backoff so retries spread out, jitter so clients do not synchronise, a retry budget so a client gives up, and a circuit breaker so a failing dependency stops being called at all. Milestone 11 adds backpressure, which addresses the same disease from the other end.

Why are we using this language here?

This milestone is where Go's limits show, and it would be dishonest to pretend otherwise.

What Go did well: injecting chaos cost about twenty lines, the supervisor is small enough to read in one sitting, and because ants are goroutines rather than threads, 571 restarts in three seconds is unremarkable. Restarting 571 OS threads in three seconds would be a noticeable event.

What Go made hard, all stemming from one design choice, that goroutines are not isolated:

  • A panic anywhere kills the process unless someone remembered to recover in that exact goroutine. Forget one and a stray nil dereference in an ant takes down the colony.
  • There is no way to kill a stuck goroutine. Our slow-ant chaos sleeps politely; a real stuck worker would sit there until the process exits.
  • Restart policy is entirely hand-rolled. We have no backoff, no restart intensity limit, and no escalation. An ant that panics immediately on start would spin in a restart loop at full CPU, and nothing in the language would notice.
  • State recovery is manual. We had to know that a crashed ant loses its cargo and remember to tell the owner.

Erlang answers all four at the language and runtime level: processes have isolated heaps so one cannot corrupt another, exit(Pid, kill) is unconditional, and supervisors come with restart strategies, intensity limits and escalation as configuration rather than code. If fault tolerance were the whole problem, Erlang would be the better tool, and Course 4 is going to make that case with the same project in miniature. Go's counter-argument is everything else: static typing, raw speed, a single deployable binary, and a profiler that will save you in Milestone 11. Most systems need some of both, which is why "which language is better" is the wrong question and "what is this language's failure model" is the right one.

Exercise 8

Our supervisor restarts an ant instantly, forever. That is the one restart policy OTP explicitly warns against. Implement a better one.

Requirements. Add exponential backoff with jitter (first restart after ~10 ms, doubling, capped at one second). Add a restart intensity limit: if more than MaxRestarts happen within RestartWindow, stop restarting, record the colony as degraded, and let Run return an error identifying it. Expose Degraded bool and Alive int on Stats.

Constraints. No mutex. Backoff must not block the supervisor, since one sleeping ant must not delay restarting the other 49,999. Shutdown must still be prompt: a pending backoff must not keep Run waiting.

Acceptance. A test with CrashProb: 0.5 ends degraded rather than burning CPU; a test with CrashProb: 0.001 does not end degraded; the leak test still passes; and the whole suite is clean under -race -count=5.

Hints. Where does the sleep go so that it does not block the supervisor? What does the restart counter need to be, given that only the supervisor touches it? How does a sleeping restart learn that shutdown has begun?

Solution 8 — open after trying
var ErrColonyDegraded = errors.New("restart intensity exceeded; colony degraded")

// restartTracker is touched only by the supervisor goroutine, so it needs
// no synchronisation.
type restartTracker struct {
	times  []time.Time
	window time.Duration
	max    int
}

func (r *restartTracker) record(now time.Time) (overloaded bool) {
	cutoff := now.Add(-r.window)
	kept := r.times[:0]           // reuse the backing array: no allocation
	for _, t := range r.times {
		if t.After(cutoff) {
			kept = append(kept, t)
		}
	}
	r.times = append(kept, now)
	return len(r.times) > r.max
}
	tracker := &restartTracker{window: e.cfg.RestartWindow, max: e.cfg.MaxRestarts}
	degraded := make(chan struct{})

	go func() {
		defer close(supervisorDone)
		for {
			select {
			case <-ctx.Done():
				return
			case a := <-deaths:
				if tracker.record(time.Now()) {
					e.degraded.Store(true)
					close(degraded)      // tell Run; ants already running continue
					return
				}
				e.recycle(ctx, a)
				delay := backoff(a.Generation, e.cfg.RestartBackoff, time.Second, e.superRNG)

				// The sleep happens in the restarting goroutine, not here,
				// so one slow restart cannot delay the others.
				antWG.Add(1)
				go func(a *Ant) {
					defer antWG.Done()
					select {
					case <-time.After(delay):
					case <-ctx.Done():
						return       // do not start an ant during shutdown
					}
					e.superviseAnt(ctx, a, deaths)
				}(a)
			}
		}
	}()

func backoff(generation uint64, base, max time.Duration, rng *rand.Rand) time.Duration {
	d := base << min(generation, 20)          // doubling, saturating
	if d > max || d <= 0 {
		d = max
	}
	// Full jitter: anywhere in [0, d). Prevents restart storms from
	// synchronising after a correlated failure.
	return time.Duration(rng.Int64N(int64(d) + 1))
}

Five things worth extracting from this solution.

  • The sleep moved into the restarting goroutine. Putting time.Sleep in the supervisor's loop would serialise every restart behind every backoff, so one crash-looping ant would delay the recovery of all the others. Whenever a supervisor waits, ask what it is not doing while it waits.
  • antWG.Add(1) happens in the supervisor, before the go. Doing it inside the new goroutine would let Run's Wait return between the two, which is the same hazard the supervisorDone channel exists to prevent.
  • The tracker needs no lock because exactly one goroutine touches it. Compare with the atomics used for restarts and dropped, which are written by many. Choosing between "owned" and "atomic" by asking who writes it is the whole discipline.
  • kept := r.times[:0] is the standard Go filter-in-place idiom: reslice to zero length, keep the capacity, append what survives. It reuses the backing array and allocates nothing, and it works because append writes over elements the loop has already read.
  • Full jitter, not fixed backoff. If a thousand ants crash together (which correlated failures do), fixed backoff makes all thousand retry at the same instant, and you get a thundering herd forever. Randomising the whole interval spreads them out. This is AWS's published recommendation and it is worth remembering as a default.

Compare the total with OTP's {one_for_one, MaxRestarts, Window}, which is one line of configuration and also handles escalation to a parent supervisor, ordered restarts of dependent children, and shutdown protocols per child. We have written maybe seventy lines for a strictly weaker version. That is not an argument against Go; it is an argument for knowing what you are reimplementing.

Experiment

Run with -chaos slow=0.01,slowfor=200ms and 5,000 ants, then watch the timeout counter. Slow ants are not failing, they are just late, yet they consume queue slots and reply channels the entire time. Now raise slowfor to two seconds while keeping RequestTimeout at 200 ms and watch the system spend its life timing out. The interaction between a client's patience and a server's latency is where most production incidents actually live, and you can explore the whole space here with two flags.

Common mistakes in Milestone 8
  • recover() not inside a deferred function. Returns nil, does nothing, and the panic proceeds to kill the process. It must be defer func(){ recover() }(), not a bare call.
  • Trying to recover another goroutine's panic. Impossible. Every goroutine that can panic needs its own deferred recover.
  • Discarding the panic value. Log it with debug.Stack() or you will never diagnose a real crash, only simulated ones.
  • Restarting without resetting state. An ant restarted mid-journey while Carrying is still true will try to drop food it does not have and generate a stream of errors.
  • Forgetting to account for in-flight work. The conservation test drifts and you spend an hour blaming the random number generator.
  • An unguarded send to deaths. Leaks a goroutine on every crash that happens during shutdown.
  • Making chaos non-reproducible by using the global rand. Then "it broke once" is all you will ever know.

Checkpoint

  1. Why must recover live in the same goroutine that panics, and what happens if you forget one?
  2. The supervisor writes a.Pos and a.Carrying, which the ant's own goroutine normally owns. Why is that not a data race?
  3. Why does the crashed ant's cargo need to be reported to the owner?
  4. What does Generation prevent?
  5. Loss went from 2% to 25% and delivered food fell by 96% while dropped messages stayed flat. Explain.
  6. Name three mitigations for retry-induced collapse that this code does not implement.
  7. What happens if the owner goroutine panics? What would it take to survive that?
  8. Why is deaths chan<- *Ant written with a direction in the signature?

Repository state after Milestone 8

antfarm/
├── go.mod
├── cmd/
│   └── antfarm/
│       └── main.go              flags, chaos parsing, signals, reporting
└── internal/
    ├── sim/
    │   ├── ant.go               Ant, Action, ActionKind, Generation
    │   ├── behaviour.go         Behaviour, Forager, TrailFollower, directions
    │   ├── sense.go             Sense, senseAt
    │   ├── engine.go            requests, serve, antClient, supervisor, Run
    │   ├── sim.go               Config, Chaos, Stats, the sequential Sim
    │   ├── concurrent.go        milestone 4's mutex versions, kept for comparison
    │   ├── engine_test.go       conservation, stats-under-load, leaks, chaos
    │   ├── sim_test.go          decide tables, determinism, benchmarks
    │   └── concurrent_test.go   race reproduction, locked benchmarks
    └── world/
        ├── grid.go              Position, Grid
        ├── pheromone.go         FloatGrid, deposit, evaporate
        ├── world.go             World, food, nest, errors
        └── grid_test.go         bounds, food errors
$ gofmt -l . && go vet ./... && go test -race ./...
ok  	github.com/yourname/antfarm/internal/sim	1.965s
ok  	github.com/yourname/antfarm/internal/world	0.003s
$ git commit -am "milestone 8: chaos injection and a hand-rolled supervisor"
Where the mutex went

Worth pausing on: engine.go is now the heart of the project and contains no mutex at all. The only synchronisation primitives in it are channels, a context, three atomic counters for cross-goroutine metrics, and one WaitGroup per lifecycle group. Every piece of mutable state has exactly one owner, and the comments say who. That is what the Milestone 4 detour bought.

The Mewlang cat, strolling forwardInstalment 3 of the five-course curriculum. Next: Milestones 9–12, where metrics get an HTTP endpoint and a profiler, the colony gets a live view you can watch in a browser, backpressure and sharding make it fast, and the world moves into a separate process that you can kill.

Continue