Milestones 9–12

Instalment 15 · Course 3 (Perl) · Advanced phase and finish

What an experienced Perl programmer reaches for next, and whether you can now explain any of it

Five advanced topics with working code, one substantial final challenge about tailing logs that rotate under you, a knowledge check of forty-one questions, and everything you need to put Strata on GitHub and defend it in an interview.

Part AThe advanced phase

The Mewlang cat, wearing glasses and looking confidentThe tool works: it ingests five formats, survives hostile files, extracts and normalises entities, correlates them into events in SQLite, behaves under a pipe and under Ctrl-C, fuzzes and profiles itself, and walks its own correlations as a graph. These five topics are what you would reach for next if Strata were something you had to operate rather than something you had to finish.

A1 · A plugin architecture for parsers

Strata::Parser::Registry has known every parser by name since Milestone 6. A real deployment wants to drop a new format handler into a directory without editing a core file at all. Perl's answer needs no dependency beyond the standard library:

package Strata::Parser::Loader;
use v5.36;
use File::Find;

# Auto-discover every Strata::Parser::* module under lib/, without a static
# list anywhere. A third party ships a new .pm file; nothing else changes.
sub discover ($class, $lib_dir = "lib") {
    my @found;
    find({
        wanted => sub {
            return unless /\.pm$/ && m{Strata/Parser/};
            (my $mod = $File::Find::name) =~ s{^\Q$lib_dir\E/}{};
            $mod =~ s{/}{::}g;
            $mod =~ s{\.pm$}{};
            push @found, $mod;
        },
        no_chdir => 1,
    }, "$lib_dir/Strata/Parser");

    for my $mod (@found) {
        (my $path = "$mod.pm") =~ s{::}{/}g;
        require $path;
    }
    return sort @found;
}

Verified against a two-file throwaway package tree: discover found and loaded both modules by name, in the order sort put them in, with no registry to keep in sync by hand. The trade-off is honest: a static list in Registry.pm is one file you can read top to bottom to know every format Strata understands; auto-discovery means reading the filesystem to answer the same question. For a tool meant to grow third-party parsers, that trade is usually worth it. What to measure: startup time as the parser directory grows — File::Find walking hundreds of files on every invocation is a real cost a static list does not pay, and the standard mitigation is a cached manifest invalidated by directory mtime.

A2 · Advanced parsing: when a regex stops being enough

Every parser so far matches one line at a time because every format so far is line-oriented, or was made to look that way (Milestone 6's pull-parser for XML). Nested configuration — the kind you would find describing which log sources feed Strata itself — is not: a regex has no memory of "how deep am I", which is precisely what nesting requires.

sub tokenize ($text) {
    my @tok;
    while ($text =~ /\G\s*(\{|\}|=|"[^"]*"|[\w.]+)/gc) {
        push @tok, $1;
    }
    return @tok;
}

sub parse_block ($tokens, $pos) {
    my %node;
    while ($$pos < @$tokens && $tokens->[$$pos] ne '}') {
        my $key = $tokens->[$$pos++];
        if ($tokens->[$$pos] eq '{') {
            $$pos++;
            $node{$key} = parse_block($tokens, $pos);
            die "expected '}'\n" unless ($tokens->[$$pos] // '') eq '}';
            $$pos++;
        } elsif ($tokens->[$$pos] eq '=') {
            $$pos++;
            (my $val = $tokens->[$$pos++]) =~ s/^"(.*)"$/$1/;
            $node{$key} = $val;
        } else {
            die "expected '{' or '=' after '$key'\n";
        }
    }
    return \%node;
}

\G anchors the next match to where the previous one left off rather than to the start of the string, and /gc (global, keep the position on failure) is what makes repeated /\G.../gc matches walk a string left to right without ever re-scanning what came before. This is Perl's built-in tokenizer loop, and most hand-rolled Perl lexers use exactly this idiom. parse_block is then ordinary recursive descent: a block is a sequence of key = value or key { ...nested block... } entries, and the recursion handles arbitrary nesting depth for free, which no single regex — however elaborate — can do, because a regular expression fundamentally cannot count in a way that carries across an unbounded nesting depth.

Verified on a three-level config with a nested block:

host: web-3, max_conn: 100
Why are we using this language here?

This is a genuine boundary of "text processing is Perl's strength". Regexes carried every format through Milestone 6; nesting needed an actual parser, which is ordinary code in any language once you know the recursive-descent shape. Racket, later in this curriculum, treats exactly this problem — turning text into a structure, structurally, with the grammar as a first-class thing rather than an ad hoc loop — as central enough to build a whole language-construction toolkit around it. Worth remembering this thirty-line parser when you get there.

A3 · Property-based testing for the normaliser

Milestone 8's timestamp tests check specific strings against specific epochs — necessary, and not sufficient, because it only tests the inputs someone thought to write down. A property test instead asserts something that must be true for any input in a category, and lets the computer generate the inputs:

use Time::Piece;

srand(9001);
my $failures = 0;
for (1 .. 2000) {
    my $epoch = 1_600_000_000 + int(rand(200_000_000));       # any epoch in ~2020-2026
    my $iso   = gmtime($epoch)->datetime . "Z";
    my $back  = Time::Piece->strptime($iso, "%Y-%m-%dT%H:%M:%SZ")->epoch;
    if ($back != $epoch) {
        say "MISMATCH: epoch=$epoch iso=$iso back=$back";
        $failures++;
    }
}
say "$failures failures out of 2000 random epochs";
0 failures out of 2000 random epochs

The property here is a round trip: format then parse should return what you started with, for every epoch, not just the six hand-picked ones from Milestone 8. Perl has no built-in property-testing framework as batteries-included as Haskell's QuickCheck or Erlang's PropEr; Test::LectroTest exists on CPAN and is worth knowing about, but the technique above — a seeded random loop with an assertion inside it — captures the essential idea in eight lines and needs nothing installed. Seeding with srand(9001) is what makes a failing run reproducible: rerun with the same seed and you get the exact same 2,000 epochs, in the exact same order, which is the property-testing equivalent of Milestone 8's deterministic chaos schedule.

A4 · Structured logs for a forensic audit trail

A tool whose job is reconstructing what happened should be able to say what it did, precisely, after the fact. A minimal leveled logger, dependency-free:

package Strata::Log;
use v5.36;
use POSIX qw(strftime);

my %LEVEL = (debug => 0, info => 1, warn => 2, error => 3);
my $threshold = $LEVEL{ $ENV{STRATA_LOG_LEVEL} // "info" } // 1;

sub log ($level, $msg, %fields) {
    return if $LEVEL{$level} < $threshold;
    my $ts = strftime("%Y-%m-%dT%H:%M:%SZ", gmtime);
    my $kv = join " ", map { qq{$_="$fields{$_}"} } sort keys %fields;
    say STDERR "$ts level=$level msg=\"$msg\" $kv";
}

sub info  ($msg, %f) { log("info",  $msg, %f) }
sub warn  ($msg, %f) { log("warn",  $msg, %f) }
sub error ($msg, %f) { log("error", $msg, %f) }
1;
2026-09-12T13:44:10Z level=info msg="ingest complete" files="11" records="4035"
2026-09-12T13:44:11Z level=warn msg="unterminated quote, falling back" file="users.csv" lineno="42"

Key-value fields rather than an interpolated sentence, for the same reason Go's advanced phase gave for log/slog: a log aggregator can filter on file="users.csv" without a regular expression. STDERR, not STDOUT — Milestone 10 already made STDOUT the data channel a script might pipe into jq or head; mixing log lines into it would corrupt exactly the output that milestone worked to make well-behaved.

A5 · Shipping it

FROM perl:5.38-slim AS build
WORKDIR /src
COPY cpanfile .
RUN cpanm --notest --local-lib=/deps --installdeps .
COPY . .

FROM perl:5.38-slim
COPY --from=build /deps /deps
COPY --from=build /src /app
ENV PERL5LIB=/deps/lib/perl5
WORKDIR /app
RUN useradd -r -s /usr/sbin/nologin strata
USER strata
ENTRYPOINT ["perl", "bin/strata"]

A two-stage build so the compiler and build-time dependencies DBD::SQLite and XML::LibXML needed to compile their C extensions never ship in the final image, and a non-root user, because there is no reason a tool that only reads log files and writes a SQLite database needs to run as root.

# /etc/systemd/system/strata-ingest.timer
[Timer]
OnCalendar=*-*-* *:00/15
Persistent=true

[Install]
WantedBy=timers.target

A systemd timer rather than a naive cron entry gets you two things cron does not for free: Persistent=true catches up a missed run after the machine was off, and journalctl -u strata-ingest gives you the structured logs from A4 alongside systemd's own record of exit status and duration, in one place.


Part BThe final challenge

The Mewlang cat, looking up curiouslyEverything up to here had a solution a few paragraphs later. This one does not, and it is deliberately at the edge of what you can now do. Spend real time on it before opening the last section.

Tailing logs that rotate, truncate, and disappear out from under you

Every milestone so far ingests files that already exist and stop changing while Strata reads them. Real logs do not hold still: they grow, get rotated (renamed aside, a fresh empty file created at the old path) or copy-truncated (truncated to zero length in place, same file, same inode, same path) by the logging system while your tool is running, sometimes every hour. Build strata watch DIR...: a live-ingesting mode that follows files as they change, survives both kinds of rotation with no duplicate and no lost record, and can be stopped and restarted without reprocessing anything it already saw.

Requirements

  1. strata watch DIR [DIR...] watches every regular file already in the given directories and picks up new files that appear in them later.
  2. Bytes appended to a file are ingested incrementally — read from where you left off, not by re-reading the file from the start.
  3. Handle rotate-by-rename (the file at a watched path becomes a different file, identified by a change of device+inode) by starting the new file at its own beginning.
  4. Handle copytruncate (the same file, same inode, truncated to a smaller size than you had already read) by resuming from the new end of file rather than either re-reading old bytes or crashing on a negative-length read.
  5. Every few seconds, run correlation on only the newly ingested occurrences and print any event that was created or extended since the last pass — a live feed, not only a final report.
  6. Ctrl-C stops cleanly (reusing Milestone 10's signal discipline) and remembers exactly how far into each watched file it got, in the store itself, so a restart neither reprocesses nor skips.
  7. Idle files must not burn CPU. Fifty quiet files and one active one should look, in a CPU profile, overwhelmingly like "waiting", not like "checking".

Constraints

Acceptance criteria

  1. Rename-rotation. Append 100 lines, mv the file aside, create a fresh file at the same path, append 50 more. All 150 lines ingested, none duplicated.
  2. Copytruncate. Append 100 lines, truncate the file to zero length in place (same inode), append 50 more. All 150 lines ingested, none duplicated — this is a different code path from (1) and must be tested separately.
  3. Restart. Watch, ingest some lines, Ctrl-C, restart against the same directories. No re-ingestion of already-seen bytes; lines appended while stopped are picked up on restart.
  4. Idle cost. 50 idle files plus one file appended to once a second for 60 seconds: the watcher's own CPU time over that minute stays low, measured with /usr/bin/time -v or equivalent, and does not scale with how long the run has lasted.
  5. Live feed. An event appears in the live output within one polling interval of the occurrence that completed it — not only when the process is later stopped.

Hints, in increasing order of how much they give away

The Mewlang cat, giving a playful winkAttempt it before reading on. Even a partial implementation with an honest account of what you didn't solve is worth more than the section below.

Solution — only look after trying

The detection rule

Three fields from one stat() call, compared against what you saw last time:

ObservationMeaningAction
device or inode changedrename-rotation: a different file now lives at this pathstart the new file from offset 0
same device and inode, size < last known offsetcopytruncate: the same file was truncated shorter than what you'd readresume from the new end of file, offset 0 is also defensible
same device and inode, size > last known offsetordinary growthread from the old offset to the new size
sub poll_file ($self, $path) {
    my @st = stat($path) or return $self->_forget($path);   # file gone
    my ($dev, $ino, $size) = @st[0, 1, 7];
    my $prev = $self->{state}{$path};

    if (!$prev || $prev->{dev} != $dev || $prev->{ino} != $ino) {
        $self->{state}{$path} = { dev => $dev, ino => $ino, offset => 0 };
        $self->_save_state($path);
        return $self->_read_from($path, 0);
    }

    if ($size < $prev->{offset}) {
        $prev->{offset} = 0;
        $self->_save_state($path);
        return $self->_read_from($path, 0);
    }

    return () if $size == $prev->{offset};
    return $self->_read_from($path, $prev->{offset});
}

Verified with a real sequence of operations — create, append, rename-and-recreate, append, copytruncate (in-place truncate via a reopened filehandle, same inode), append again — against exactly this detection logic: every transition was classified correctly (new, appended, rotated, appended, truncated, appended), with the copytruncate case in particular only working because the check compares against the previously read offset, not against zero — a naive "did the size shrink" check without that comparison cannot tell a copytruncate from a file that simply hasn't grown yet.

Persistence, so a restart knows what it already owes nothing

CREATE TABLE watch_state (
    path      TEXT PRIMARY KEY,
    dev       INTEGER NOT NULL,
    ino       INTEGER NOT NULL,
    offset    INTEGER NOT NULL,
    updated_at INTEGER NOT NULL
);

One row per path, upserted after every successful ingest of a chunk — and only after, never before, so that a crash mid-read leaves the recorded offset at the last point genuinely committed rather than claiming credit for bytes not yet safely in the database. On startup, watch loads this table before its first poll, so "restart" is not a special case in the code at all: it is simply the first poll of a process whose %state happens to have been pre-populated from SQLite instead of built up from scratch.

The live feed, without re-running the whole sweep

Milestone 9's sessionize re-sorts and re-groups every occurrence on every call, which is wrong here — correct, but wasteful, and wasteful in a loop that runs every few seconds forever. watch instead sessionises only affected keys:

my %touched;                                  # (type, value) pairs seen this poll
for my $occ (@new_occurrences) {
    $touched{"$occ->{type}\0$occ->{value}"} = 1;
}

for my $key (keys %touched) {
    my ($type, $value) = split "\0", $key;
    $self->{store}->resessionize_one($type, $value);   # re-sweep just this entity's occurrences
}

Re-sweeping one entity's occurrences (typically a handful to a few hundred rows, indexed by exactly the idx_occ_type_value_ts covering index Milestone 9 built) rather than the whole table is what keeps each poll's cost proportional to what actually changed, not to how much history the database has accumulated.

The instant-of-rotation gap, and the honest choice

Between the last poll and the rename, a writer can still append a few bytes to the file at its old inode that your next poll — which now looks at the new inode under that path — will never see by path alone. This implementation accepts that gap: it is bounded by the poll interval, it is the same choice tail --follow=name (the GNU default for a watched filename) makes, and closing it requires keeping the old file descriptor open past the rename and giving it one final read, which is what tail --follow=descriptor does instead, at the cost of tracking two live handles per watched path during the transition. Document the choice; do not leave it implicit.

What is still wrong with this, and you should say so in your README

  • Polling, not events. Fifty files at a few-hundred-millisecond interval is genuinely fine; five thousand files is not — every poll is now five thousand stat calls whether or not anything changed. Linux::Inotify2 turns this into "wake up only when something changed", at the cost of a dependency and of Linux-only portability.
  • The rotation gap above is real, bounded, and undocumented in most tools that have it, which is worse than documenting it here.
  • One process, one SQLite writer. This does not parallelise across files the way Milestone 12's batch ingest does, because a live tail needs to notice new bytes promptly on every watched file, and splitting that across forked workers means either sharding which files each worker owns (straightforward) or fighting over one SQLite writer from several processes at once (not straightforward, and not attempted here).
  • No back-pressure if ingestion falls behind the write rate. A sufficiently fast writer can grow the gap between "bytes on disk" and "bytes ingested" without bound; a production version would track that lag as a metric and alert on it rather than discover it only when disk space runs out.

If you can explain the difference between the rename-rotation and copytruncate cases, and why a naive "did the file shrink" check handles neither of them correctly without device+inode tracking, you understand the actual hard part of "tail -f", which is a component inside more infrastructure than you might expect: log shippers, metrics agents and container runtimes all solve a version of this exact problem.


Part CKnowledge check

C1 · Twenty conceptual questions

  1. State Perl's context rule precisely: what decides whether an expression is evaluated in list or scalar context, and name three operations that behave differently depending on which.
  2. Why does my $n = @array give you a count while my ($n) = @array gives you the first element?
  3. What does wantarray return in each of the three calling contexts, and why is "void" a context distinct from "scalar"?
  4. Explain autovivification. Why does if ($h{a}{b}{c}) create $h{a} and $h{a}{b} even though the condition is false?
  5. Why are hash keys always strings, and what follows for a hash keyed by something that looks numeric, like an IP octet or a port number?
  6. What is the difference between my and local? Give a case where local is the only correct choice.
  7. Why did for (my $i = 1; $i <= 3; $i++) { push @subs, sub { $i } } capture one shared $i across all three closures, while for my $i (1..3) { push @subs, sub { $i } } does not?
  8. Explain foreach's aliasing behaviour: why does $_ = uc $_ inside for (@list) { ... } modify @list itself?
  9. What does local $/ do, and why does setting it to undef enable "slurp mode"?
  10. Why is Perl's sort guaranteed stable since 5.8, and what would break in the sessioniser's sweep if it were not?
  11. What is a reference, in terms of what it actually stores, and how does it differ from a C pointer?
  12. Explain why Text::CSV or XML::LibXML cannot be replaced by a "good enough" regex, using a specific example from Milestone 6.
  13. What does PerlIO::encoding::fallback control, and what does Perl's :encoding layer do to an undecodable byte if you never set it?
  14. Why must alarm(0) be called on every exit path of a timeout-protected block, success and failure alike?
  15. Explain the difference between eq/== and when using the wrong one produces a bug that only shows up on some inputs.
  16. Why does wrapping ten thousand INSERTs in one transaction, rather than autocommitting each one, produce a two-orders-of-magnitude speedup in SQLite specifically?
  17. What makes a SQLite index "covering", and why does that matter more for a point lookup than for a full scan?
  18. After fork(), what exactly do the parent and child share, and what is copied? Why does that make a data race between them structurally impossible?
  19. Why does closing the unused end of a pipe, in both the parent and the child, matter for correctly detecting end-of-file?
  20. What does WITH RECURSIVE's plain UNION (rather than UNION ALL) prevent when the underlying graph contains a cycle?
Answers to C1
  1. Context is determined entirely by the syntactic position an expression appears in — assignment to an array or a list of variables imposes list context, assignment to a scalar imposes scalar context, a boolean test imposes scalar context — never by the expression's own contents. Three context-sensitive operations: assigning an array to a scalar (count), reverse (list: reverses elements; scalar: reverses a concatenated string), and a subroutine call, whose wantarray can act on it explicitly.
  2. Assigning an array to a bare scalar imposes scalar context on the array, which for an array means "how many elements". Assigning to a parenthesised list of one variable, ($n), imposes list context, so @array supplies its elements in order and $n receives the first one, discarding the rest.
  3. True in list context, false (but defined) in scalar context, undef in void context — called for its side effects with the return value discarded entirely, as in a bare ctx(); statement. Void is distinct from scalar because some functions legitimately do less work when nothing will read the result at all.
  4. Perl builds a hash's nested structure lazily, on first access, so that $h{a}{b}{c} can be written at all without every level already existing. Merely looking up $h{a}{b}{c} — even inside a boolean test that never assigns anything — still has to dereference $h{a} as a hash reference to get anywhere, and Perl autovivifies it on that dereference rather than requiring it to pre-exist. The condition being false afterwards changes nothing about the levels that were already created to evaluate it.
  5. Hash keys are stored and compared as strings unconditionally, so $h{7} and $h{"7"} are the same entry, but $h{"007"}, $h{"07"} and $h{7} are three distinct entries despite being numerically equal — a real risk for entity deduplication keyed on something that looks numeric but may carry leading zeros or varying formatting.
  6. my creates a new lexical variable scoped to its enclosing block, resolved at compile time. local saves and temporarily replaces the value of an existing (usually global or package) variable for the dynamic extent of the current block, restoring the old value on exit — used throughout this project for exactly one purpose: temporarily replacing $SIG{ALRM} or $SIG{INT} so the replacement is guaranteed to be undone when the protected block ends, however it exits.
  7. A C-style for loop has exactly one $i, declared once before the loop body runs at all; every closure created inside the loop captures that same variable, so all three see whatever value it holds after the loop finishes (4, from the last increment past the exit condition). for my $i (1..3) is Perl's foreach form, which creates a fresh lexical $i for each iteration, so each closure captures a genuinely distinct variable — the same distinction Go's for loop semantics changed to match, in Go 1.22.
  8. foreach does not copy each list element into $_; it aliases $_ directly to the element itself for the duration of that iteration. Assigning to $_ is therefore assigning to the original array slot, which is occasionally exactly what you want (an explicit in-place transform) and occasionally a bug from forgetting the aliasing exists at all.
  9. $/ is the input record separator that <$fh> reads up to. Setting it to undef removes any separator to stop at, so the next diamond read returns the entire remaining content of the filehandle as one scalar — "slurp mode" — which is exactly what a format like XML, with no meaningful line structure, needs instead of line-by-line reading.
  10. Stability means two elements that compare equal keep their original relative order after sorting, which Perl has guaranteed since 5.8 regardless of the underlying algorithm. The sessioniser relies on it implicitly: occurrences with identical (type, value, ts) keep whatever order SELECT ... ORDER BY produced, rather than being shuffled between runs, which matters for reproducible output when comparing two runs of the same ingest.
  11. A reference is a scalar value that holds the address and type of another value — an array, a hash, a scalar, a code value, or an object — and, unlike a C pointer, always knows what kind of thing it refers to and participates in reference counting, so Perl frees the referent automatically once nothing references it any more. There is no pointer arithmetic, and dereferencing a reference of the wrong kind (treating an array reference as a hash reference) is a checked runtime error, not undefined behaviour.
  12. CSV's quoting rules mean a field can legally contain a comma, a newline, or an escaped quote — "Multi\nline name" is one field spanning two physical lines, which a regex split on commas or newlines has no way to know without effectively re-implementing a CSV parser's state machine. XML has no fixed structure to match against at all: an element can span any number of lines or be minified onto one, so "match a tag with a regex" only ever works until a document is formatted differently than the regex assumed.
  13. It controls what Perl substitutes for a byte sequence that cannot be decoded under the active encoding. Left at its default, PerlIO's :encoding layer inserts a literal backslash-x escape sequence as ordinary text characters — not a recognisable error marker — so a decoding failure becomes indistinguishable from real data unless FB_DEFAULT is overridden with something like Encode::FB_DEFAULT tuned to produce U+FFFD instead.
  14. An armed alarm is a process-wide timer that fires in whatever code happens to be running when its deadline arrives, regardless of whether that code has anything to do with what set the alarm. If the protected block finishes early but the alarm is left pending, it fires later, in unrelated code, producing a die that looks like it came from nowhere; cancelling on both the success path and every failure path is what guarantees the timer never outlives the operation it was meant to bound.
  15. eq compares two values as strings; == converts both to numbers first and compares those. "07" == "7" is true (both convert to the number 7) while "07" eq "7" is false (different strings) — a bug that hides for exactly as long as every value in a dataset happens not to have a leading zero, a decimal point, or trailing whitespace that changes its numeric conversion.
  16. SQLite's default durability guarantee is that a transaction is not considered committed until its write-ahead log entry is flushed to physical storage, which is an expensive operation relative to an in-memory write. Autocommit mode makes every single statement its own transaction, paying that flush cost once per row; wrapping many inserts in one explicit transaction pays it once for the whole batch, which is why the measured difference was roughly two orders of magnitude rather than a modest constant factor.
  17. A covering index contains every column a query needs, so SQLite can answer the query by reading the index alone and never touching the underlying table rows at all. A full scan already has to visit most of the table regardless, so an index mainly reorders that work; a point lookup that would otherwise scan the entire table to find a handful of matching rows benefits enormously, because the index turns "check every row" into "walk directly to the matching branch of a B-tree".
  18. fork() gives the child a complete, independent copy of the parent's address space at the moment of the call — copy-on-write, so physically cheap until either side writes to a shared page, at which point the kernel actually duplicates just that page. After that instant neither process can observe or modify the other's memory at all; there is no shared mutable state for two forked processes to race on, because "shared" stops being true the moment either one writes.
  19. A pipe's read end reports end-of-file only once every writer-side file descriptor referring to it has been closed. If a process forks and neither parent nor child closes the copy of the end it does not use, that unused copy keeps the pipe's write end alive from the kernel's point of view even after the "real" writer has finished, so a read that should see EOF instead blocks forever waiting for a close that will never come from a descriptor nobody is using.
  20. UNION deduplicates identical rows as the recursion accumulates them, which is what stops a cycle in the underlying graph from being retraced indefinitely — each (node, hops) pair the recursion could produce is only ever added once, so the recursion has a finite number of distinct rows left to produce and necessarily terminates. UNION ALL keeps every duplicate, and a true graph cycle under UNION ALL recurses forever.

C2 · Ten code-reading questions

Predict the output of each, then check. All ten were run to confirm the answers.

// 1
my @a = (1,2,3);
my $n = @a;
my ($first) = @a;
say "$n $first";

// 2
my %h;
$h{a}{b}{c} = 1;
say exists $h{a} ? "yes" : "no";
say ref $h{a};

// 3
my $s = "9"; $s++;
my $t = "Az"; $t++;
say "$s $t";

// 4
my @list = (3,1,2);
say "@{[ sort @list ]}";
my @nums = (10, 9, 2);
say "@{[ sort @nums ]}";

// 5
sub ctx { return wantarray ? "list" : defined(wantarray) ? "scalar" : "void" }
ctx();
my $x = ctx();
my @y = ctx();
say "$x @y";

// 6
my @subs;
for my $i (1..3) { push @subs, sub { $i } }
say join ",", map { $_->() } @subs;

// 7
my @subs2;
for (my $i = 1; $i <= 3; $i++) { push @subs2, sub { $i } }
say join ",", map { $_->() } @subs2;

// 8
my @names = ("alice", "bob");
for (@names) { $_ = ucfirst $_ }
say "@names";

// 9
my @prices = (19.995, 4.999);
printf "%d %.2f\n", $_, $_ for @prices;

// 10
my @arr = (1,2,3,4);
my $count = () = @arr;
say $count;
Answers to C2
1.  3 1
2.  yes
    HASH
3.  10 Ba
4.  1 2 3
    10 2 9
5.  scalar list
6.  1,2,3
7.  4,4,4
8.  Alice Bob
9.  19 20.00
    4 5.00
10. 4
  1. Scalar-context assignment of an array gives its count (3); list-context assignment to ($first) gives its first element.
  2. Assigning to $h{a}{b}{c} autovivifies $h{a} as a hash reference along the way, so it exists and is a HASH ref, even though the code never explicitly created it.
  3. Magic string increment: "9" becomes numeric-looking and increments to "10"; "Az" increments alphabetically with carry, becoming "Ba", exactly like an odometer.
  4. Perl's default sort compares elements as strings unless told otherwise: numeric 3,1,2 sort correctly as strings to 1,2,3, but 10,9,2 sort as strings to "10","2","9" because "1" lt "2". A numeric sort needs an explicit sort { $a <=> $b } @nums.
  5. The bare ctx(); statement runs in void context and its return value is discarded entirely — never printed. $x is assigned in scalar context, @y in list context.
  6. for my $i (1..3) gives each closure its own fresh lexical, so the three closures return their own captured value, in order.
  7. The C-style for loop has one $i shared by every iteration and every closure; by the time any of the three closures run, the loop has finished and $i holds 4 (the value that failed the <= 3 test), so all three print the same number.
  8. foreach aliases $_ to the actual array element, not a copy, so assigning to $_ mutates @names in place — a common source of "why did my input array change" bugs.
  9. %d truncates toward zero rather than rounding, so 19.995 prints as 19 and 4.999 as 4; %.2f rounds correctly to two decimal places.
  10. () = @arr assigns the array to an empty list in list context and, wrapped in $count = ( ... ), that assignment expression evaluates in scalar context to the number of elements assigned — a genuine, if slightly cryptic, idiom for "count without naming a throwaway array", sometimes called the goatse operator for its shape.

C3 · Five debugging exercises

Each gives a symptom and a suspect. Diagnose before opening the answer.

  1. Symptom: a report script that builds closures per row for lazy formatting prints the same value in every row instead of each row's own value.
    my @formatters;
    for (my $i = 0; $i < @rows; $i++) {
        push @formatters, sub { format_row($rows[$i]) };
    }
  2. Symptom: a "does this record have a location?" check that should just read data is somehow creating thousands of empty hash entries, visible as memory growth and a much slower keys %index loop later in the same run.
    for my $rec (@records) {
        next unless $rec->{fields}{geo}{country};
        $index{ $rec->{fields}{geo}{country} }++;
    }
  3. Symptom: a "top talkers" report that groups by client IP shows the same address listed three separate times with different counts, instead of once with the combined count.
    my %by_ip;
    $by_ip{ $rec->field("ip_octet_str") }++ for @records;   # values like "007", "07", "7"
  4. Symptom: a normalising pass that's supposed to build a cleaned-up copy of a list instead corrupts the original list it was only meant to read.
    my @hosts = extract_hosts(@records);
    for (@hosts) { $_ = lc $_ }
    push @clean_hosts, @hosts;   # caller finds @hosts itself is now lowercased too
  5. Symptom: a nightly summary occasionally dies with Died within alarm handler or hangs entirely, only ever under heavy load, never in development.
    sub with_deadline ($code, $seconds) {
        $SIG{ALRM} = sub { die "timeout\n" };   # note: not local
        alarm($seconds);
        my $r = $code->();
        alarm(0);
        return $r;
    }
Answers to C3
  1. A C-style loop with one shared $i. Every closure captures the same variable, and by the time any formatter actually runs, the loop has finished and $i holds its final, past-the-end value — so every row formats using whatever row that index pointed at last. Fix: for my $i (0 .. $#rows), which gives each closure its own lexical, or capture the row itself rather than its index (my $row = $rows[$i]; push @formatters, sub { format_row($row) }).
  2. Autovivification through a chained read. $rec->{fields}{geo}{country} dereferences {fields} and then {geo} as hash references to reach country, and if a record has no geo key at all, that dereference creates an empty {geo} hashref on the spot, purely from being read. Across thousands of records with no location data, that is thousands of empty hashes silently added to memory, and exists $rec->{fields}{geo} checked first (without a chained autovivifying read past it) avoids creating anything, or use Data::Diver's non-autovivifying accessor.
  3. Hash keys are strings. "007", "07" and "7" are three different keys despite being numerically equal, so the counter silently fragments across however many textual spellings the source data happened to use for the same value. Fix: normalise to a canonical numeric or zero-padded form (exactly what Strata::Extract's normalise step exists to do) before ever using a value as a hash key.
  4. foreach aliases, it does not copy. for (@hosts) { $_ = lc $_ } mutates @hosts in place because $_ is each element, not a stand-in for it — so any list built from @hosts before this loop, and @hosts itself afterward, are all lowercased whether or not that was intended. Fix: build a new list explicitly, my @clean = map { lc $_ } @hosts;, which never touches the original.
  5. A non-localised signal handler, and a race under concurrent calls. Setting $SIG{ALRM} directly (not local) leaves it installed globally once this sub returns; under load, with several deadline-protected operations plausibly overlapping (or an earlier call's alarm not yet cancelled when a new one starts), a stray alarm can fire inside code that never expected it, or the handler from a finished call can still be armed when a completely different timeout fires. Fix: local $SIG{ALRM} = sub { die "timeout\n" }; inside an eval, exactly as Milestone 11's fuzz harness does, so the handler is guaranteed to be restored to whatever it was before, on every exit path.

C4 · Five implementation exercises

  1. A syslog parser (Milestone 6's own forward reference): both the classic Mon DD HH:MM:SS host tag: msg form and the RFC 5424 <PRI>1 timestamp ... form, in one parser, with a detect that does not falsely claim ordinary prose. Wire it into the Registry and add fixtures.
  2. Parallel correlate. Milestone 12 parallelised ingestion; sessionisation still runs single-threaded. Shard by entity (type, value) hash across forked workers (each sessionises a disjoint slice of entities, so there is no shared state to coordinate), and benchmark against the single-process version at 500,000 occurrences.
  3. A redaction stage using the keyed-hash pseudonymisation sketched in Milestone 8's exercises: replace extracted emails and card-shaped numbers in raw with stable tokens before a record is ever written to records.raw, so the database itself never holds the original sensitive text.
  4. An idle-safe strata watch health check: a --report-every 60s flag that logs (via Milestone A4's logger) a one-line summary — files watched, bytes ingested, events created, current lag if measurable — on a fixed interval regardless of whether anything changed, so an operator watching logs can distinguish "quietly healthy" from "silently stuck".
  5. A strata verify subcommand that re-reads every file referenced in records, re-parses each recorded line by (file, lineno), and reports any record whose stored raw no longer matches what is currently on disk at that position — detecting logs that were edited or replaced after ingestion, which is exactly the kind of tampering a forensic tool should be able to notice about its own evidence.

C5 · One substantial challenge

Distinct from the final challenge in Part B, and smaller, but not easy.

Build a differential fuzzer. For any two of the five format parsers, generate inputs that are ambiguous between formats on purpose (a CSV row whose first field looks like a JSON object; a line that is simultaneously a plausible Apache log line and plausible unstructured prose) and assert a specific property: that Strata::Parser::Registry's detect scores never produce a tie between two non-fallback parsers on the same input, or, where a genuine tie is unavoidable, that the registry's tie-breaking rule is deterministic and documented rather than dependent on hash key iteration order.

Requirements: at least 5,000 generated inputs per format pair; any discovered tie or nondeterminism must be saved as a permanent fixture and regression test, the same way Milestone 11's fuzzer turned its one found hang into t/90-fuzz.t; and the generator must be seeded and the seed logged, so a found problem is reproducible on demand. Hint: the interesting bugs are rarely in one parser's detect being wrong in isolation — they are in the comparison between two parsers' scores being close enough that which one wins depends on something that was never meant to be load-bearing, like hash iteration order in the registry's own scoring loop.

C6 · You should now be able to explain

C7 · You should now be able to implement


Part DShipping it: README, portfolio, interview

D1 · README draft

# strata

A text-forensics tool: ingest ugly, heterogeneous log data (Apache/nginx,
CSV, JSON lines, XML, arbitrary unstructured text — even gzipped, mis-encoded
or truncated), extract and normalise entities, correlate them into events,
and query the result as a graph. Built to survive input designed to break it.

## What it does

- Five pluggable formats behind one interface, chosen by confidence-scored
  content sniffing, never by trusting a file extension alone.
- Streams input of any size in bounded memory: gzip, BOMs, bad encodings,
  truncated files and multi-megabyte "lines" are all handled explicitly.
- Extracts and normalises entities (IPs, hosts, paths, timestamps) with a
  match -> validate -> normalise pipeline, not bare regex matching.
- Correlates occurrences into events in SQLite (one transaction per batch:
  ~700x the throughput of naive autocommit-per-row inserts) and links
  related events into a graph, queryable with recursive SQL.
- A real command-line tool: subcommands, --help, sysexits-style exit codes,
  clean shutdown on Ctrl-C, correct behaviour piped into `head`.
- A deadline-bounded fuzzer (found a real infinite loop in under a second)
  and a profiling session with a measured 50x fix, both checked into the
  test suite as permanent regressions.
- `strata watch`: follows growing log files live, correct across both
  rename-rotation and copytruncate, resumable after a restart.

No third-party dependencies beyond DBI, DBD::SQLite, Text::CSV and
XML::LibXML — no framework, no ORM.

## Quick start

    cpanm --installdeps .
    ./bin/strata ingest strata.db share/fixtures/*.*
    ./bin/strata correlate strata.db --gap 60
    ./bin/strata query strata.db --type ipv4 --value 10.0.5.100
    ./bin/strata watch strata.db /var/log/myapp

## Commands

    ingest DB FILES...     parse and store records + entity occurrences
    entities DB FILES...   list extracted entities, ranked
    correlate DB           build events from occurrences (sessionisation)
    query DB [filters]     look up records and events
    graph DB --from ID     walk correlated events as a graph, N hops
    watch DB DIRS...       live-ingest and live-correlate growing logs

## Architecture

    lib/Strata/
    ├── Source.pm, Extract.pm, Normalize.pm    survive and clean the input
    ├── Parser/                                 five formats, one contract
    ├── Store.pm, Correlate.pm, Graph.pm        SQLite: events and edges
    └── Log.pm                                  structured logs, audit trail

## Testing

    prove -l t/                    everything, including the fuzzer
    perl -d:NYTProf bin/strata ...  profile a real invocation

The fuzz suite (`t/90-fuzz.t`) is deadline-bounded: every case gets one
second via `alarm()`, so a hang is reported as a test failure, never as a
CI run that quietly never finishes.

## Known limitations

- `watch` polls; past a few hundred files, an inotify-based backend would
  be needed instead (sketched, not implemented).
- Correlation runs in one process; ingestion parallelises with `fork`,
  sessionisation currently does not.
- No consensus or distribution: one SQLite file, one machine.

## Licence

MIT

Three deliberate choices, matching Course 1's README: it leads with what was measured (the 700× transaction number, the 50× profiling fix, the sub-second fuzzer find) rather than adjectives; it names the fuzzer and profiler as permanent parts of the test suite, not one-off exercises; and it has a known limitations section, because a forensics tool that overstates its own reliability is a bad one.

D2 · GitHub project description

A text-forensics and log-correlation tool in Perl: five formats behind one confidence-scored interface, entity extraction and correlation in SQLite, a fuzzed and profiled parser, and a rotation-safe live tail. Built as a laboratory for Perl's text-processing and Unix-glue strengths.

Topics: perl, log-analysis, text-processing, sqlite, dbi, cli, fuzzing, forensics, data-correlation, parsing, regex.

D3 · Performance considerations

D4 · Security considerations

D5 · What to put in your portfolio

The Mewlang cat, in a thoughtful three-quarter poseDo not present this as "a log parser". Present it as what it is: a forensics pipeline built to survive hostile input at every stage, with every claim about performance and robustness backed by a number you actually measured. The narrative that makes it interesting is the sequence of real bugs found by testing for them on purpose, not the feature list.

  1. A parser that trusted well-formed input, broken by six deliberately hostile fixtures — gzip, BOM, bad encoding, truncation, binary bytes, a giant line — each with the specific fix.
  2. Two false positives in entity extraction (a version number mistaken for an IP, a protocol version mistaken for a path) that only a match-validate-normalise pipeline with context rules catches.
  3. A ~700× transaction speedup and a ~53× index speedup, both measured, both explained in terms of what SQLite is actually doing differently.
  4. A fuzzer that found a genuine infinite loop in under a second, turned immediately into a permanent regression test rather than a one-off finding.
  5. A profiling session where the "obvious" optimisation (cache the regex) measured as a no-op and the unglamorous one (a hash instead of grep) measured as 50×.
  6. The rotation-safe live tail, and the honest accounting of what it still does not solve (inotify, the rename-instant gap, multi-writer SQLite).

Keep a docs/ folder with the NYTProf HTML report, the EXPLAIN QUERY PLAN output, and one architecture diagram. A reviewer who spends ninety seconds on the repository should come away knowing you test for failure on purpose, not just that the happy path works.

D6 · Interview questions someone could ask, and what a good answer contains

QuestionWhat a strong answer includes
Walk me through the entity extraction design.Match, validate, normalise as three distinct steps; two real false positives (version numbers, protocol versions) that a bare regex could not have avoided; context rules as the actual defence.
Why SQLite and not Postgres or a message queue?One file, no server to operate, WAL mode for concurrent readers, and — the honest limit — no multi-writer story past one machine, stated as a known limitation rather than discovered by the interviewer.
Tell me about a performance bug you found.The grep-based dedupe at O(n·u), found by profiling rather than guessing, fixed to O(n) with a hash, with the 50× number and the line-level NYTProf evidence that pointed at it.
And one where your intuition was wrong?The qr// regex-caching guess that measured as a no-op, because Perl's engine already caches an interpolated pattern when the interpolated value hasn't changed — the point being that "obviously slow" and "measurably slow" are different lists, and only profiling tells them apart.
How does fault tolerance work here?SIGINT sets a flag rather than acting inside the handler; the main loop commits at a safe point and reports exactly how far it got, verified by actually sending the signal mid-run and checking row counts afterward, not by reading the code and assuming.
How did you test for concurrency bugs, given Perl doesn't have goroutines or a race detector?There is nothing to race on: fork gives each worker an independent address space, so the testing burden shifts entirely to the IPC boundary — correct pipe-end closing, correct reaping, and the deliberate one-way-only design that avoids the classic buffer deadlock.
What was the hardest bug to find?The my (@current, $key, $events) = ((), "", 0) list-assignment slurp: the error surfaced two lines away from the actual cause, and the fix required understanding exactly how Perl fills a mixed list of variables from one flat right-hand list.
How would you make this production-ready?An inotify-based watch backend past a few hundred files, decompressed-size bounds alongside the existing line-length bounds, a redaction stage before raw is ever persisted, and file-permission-based access control on the SQLite file stated explicitly rather than assumed.
When would you not use Perl for this?Go or Rust if the ingest needed true multi-core parsing sharing one in-memory index rather than a SQLite file; a JVM language if this needed to run as a long-lived service inside infrastructure that already standardised on one. Name what Perl won on here: CPAN's depth for exactly these formats, and how little code the streaming and text-handling primitives needed.
What would you do differently?Build the parser contract test in Milestone 6, when the gap was first named, rather than deferring it to Milestone 11; design the watch_state schema before writing the polling loop rather than after; and decide the rename-rotation gap's trade-off explicitly at design time instead of discovering it while writing the final challenge's write-up.

D7 · Extensions worth building


Course 3 completeWhat you built

The Mewlang cat, raising a paw in celebrationA dependency-light Perl distribution across roughly a dozen modules, five parsers behind one contract, a SQLite-backed correlation and graph engine with measured order-of-magnitude performance differences, a command-line tool that behaves correctly under a signal and inside a pipeline, a fuzzer that found a real bug in under a second, and a live log tailer that survives both ways a log file can change out from under you. More importantly: a habit of writing the hostile fixture before trusting the code that has to survive it, and of measuring a claimed speedup before writing it down.

The central question of this curriculum was what kinds of problems does this language make unusually natural to solve? Perl's answer, stated as precisely as this project allows: problems where the input is real-world messy, the shape of "correct" is "did not corrupt or lose the awkward 10% of records", and the win comes from CPAN's decades of exactly-this-format modules plus a handful of small, sharp built-in idioms — context, autovivification, foreach aliasing, alarm(), fork — that read as strange in isolation and as exactly right once you have needed them once. Not the fastest, not the most structured, not the friendliest first error message. The one where a text file nobody designed on purpose stops being a mystery in an afternoon.

Courses 4 and 5 continue the same comparison — Erlang's processes and supervision trees against Go's goroutines and Perl's fork, and Racket's macros against the recursive-descent parser built in this instalment's advanced phase. The syllabus overview describes what they cover.

The Mewlang cat, walking away in a rear viewInstalment 15 of the five-course curriculum, and the end of Course 3. Course 4 (Erlang) is next.

Next: Erlang instalment