Instalment 13 · Course 3 (Perl) · Milestones 5–8
A real distribution, parsers that sniff their own format, a reader that survives gzip and mangled bytes and three-megabyte lines, and the unglamorous normalisation without which nothing can be correlated at all.
Perl 5.38.2, with Text::CSV, XML::LibXML and the core modules installed. The suite is now 7 files and 46 tests, all passing. Four bugs are documented below, three of which I found by running the finished tool over the fixtures and staring at the output, which is the point of building fixtures designed to be nasty.
Turn a folder of scripts into something installable: a version, documentation, declared dependencies, a build file, and the two tests every Perl distribution should have before it has any others.
POD, Exporter and selective exports, cpanfile versus Makefile.PL, $VERSION, compile tests, and BAIL_OUT.
A distribution is not "the code plus some paperwork". Each new file answers one question a stranger — or you, in six months — actually asks before trusting the code enough to run it:
| Question a stranger asks | Answered by |
|---|---|
| "What is this, and how do I call it?" | POD in lib/Strata.pm, read by perldoc |
"What do I get if I use this module?" | @EXPORT_OK in Strata::Util — nothing by accident |
| "What has to be installed, and why?" | cpanfile, one comment per dependency |
| "Does it even load?" | a compile test, run first, that BAIL_OUTs if the answer is no |
That ordering is also the order to write the pieces in. A module nobody can require makes every other question moot, which is why the compile test — not the documentation, not the exports — is the first thing this milestone's t/ directory should contain, even though it is described last below because it is easiest to understand once you have seen what it is protecting.
package Strata;
use v5.36;
our $VERSION = '0.01';
=head1 NAME
Strata - forensic text archaeology for heterogeneous log data
=head1 SYNOPSIS
use Strata::Pipeline;
use Strata::Parser::Registry;
my $parser = Strata::Parser::Registry->for_file("access.log");
my $pipe = Strata::Pipeline->new(parser => $parser);
$pipe->run_file("access.log");
=head1 DESCRIPTION
Its central policy is that B<a line you cannot parse is still evidence>:
parsers never throw and never return undef, they return a record carrying
the raw text, its provenance, and a description of what went wrong.
=cut
1;
POD (Plain Old Documentation) is part of the language, not a comment convention. The parser skips from =head1 to =cut, and perldoc lib/Strata.pm renders it immediately, as does every module viewer and MetaCPAN. Two habits worth adopting: a SYNOPSIS that is copy-pasteable working code, and a DESCRIPTION that states the one design decision a reader must understand before using the module.
package Strata::Util;
use Exporter 'import';
our @EXPORT_OK = qw(commify top human_bytes truncate_str);
our %EXPORT_TAGS = (all => \@EXPORT_OK);
@EXPORT_OK means "you may ask for these"; the older @EXPORT means "you get these whether you asked or not", which pollutes your caller's namespace and occasionally clobbers their own subs. Use @EXPORT_OK always, so that use Strata::Util qw(commify) states at the call site where commify came from. The :all tag exists for the one script that genuinely wants everything.
# cpanfile: for developers and deployment
requires 'perl', '5.036';
requires 'Text::CSV', '2.00'; # RFC 4180 CSV, embedded quotes and newlines
requires 'XML::LibXML', '2.00'; # a real XML parser; regexes cannot do XML
requires 'DBD::SQLite', '1.60';
on 'test' => sub {
requires 'Test::More', '1.302';
};
on 'develop' => sub {
requires 'Devel::NYTProf'; # milestone 11
};cpanm --installdeps . reads that. The Makefile.PL repeats the runtime dependencies because that is what cpan and packaging tools read; the duplication is real and mildly annoying, and the modern answer (Dist::Zilla or Minilla) generates one from the other. For a project this size, writing both by hand is less machinery than adopting a distribution builder.
Note the comments. A dependency line should say why, because the reader's real question is "can I remove this?". "Regexes cannot do XML" answers it in four words.
my @modules = qw(
Strata Strata::Util Strata::Pattern Strata::Record
Strata::Pipeline Strata::Parser::Apache
);
use_ok($_) or BAIL_OUT("$_ does not compile; nothing else can pass") for @modules;
A compile test costs nothing and catches the most common Perl failure of all: a missing semicolon or a forgotten 1; in a module you have not run today. BAIL_OUT stops the entire test run, which is right here: if a module does not compile, three hundred subsequent failures tell you nothing new.
The second is the unit test for the helpers, and it is worth showing one assertion:
is_deeply [top(10, \%counts)], [qw(b a c d)], "asking for more than exists is fine";
is_deeply [top(2, {})], [], "an empty hash yields nothing";
Both are edge cases, and both are the kind of thing that works by accident until someone refactors. Testing the boring boundaries of a four-line function is cheap insurance for code that every report depends on.
Strata::Util's @EXPORT_OK already promises human_bytes and truncate_str — Milestone 1 wrote commify, Milestone 2 wrote top, and these two have been sitting in the export list unimplemented ever since. Write them. human_bytes($n) should turn a byte count into something like "206.1KB" (check it against Milestone 1's own fixture: 211,017 bytes should come out as 206.1KB), climbing through KB, MB, GB and TB as the number grows. truncate_str($str, $max) should return the string unchanged if it already fits within $max characters, and otherwise cut it and append "..." so that the whole result — text and ellipsis together — is exactly $max characters, never more. Write the tests before you write the functions, the way the rest of this milestone argues you should.
sub human_bytes ($bytes) {
my @units = ("B", "KB", "MB", "GB", "TB");
my $n = $bytes;
my $i = 0;
while (abs($n) >= 1024 && $i < $#units) {
$n /= 1024;
$i++;
}
return $i == 0 ? "${n}B" : sprintf("%.1f%s", $n, $units[$i]);
}
sub truncate_str ($str, $max = 40) {
return $str if length($str) <= $max;
return substr($str, 0, $max - 3) . "...";
}Two details worth noticing. human_bytes stops climbing units at $#units, the last valid index, so a value bigger than the largest unit still prints — as an oversized number of terabytes — rather than reading off the end of @units and returning undef silently in the middle of a report line. truncate_str subtracts 3 before cutting, not after: the whole point of a bounded field width is that the result never exceeds it, and appending "..." to an already-$max-character substring would make the truncated version longer than the string it replaced for inputs near the boundary, which defeats the reason to truncate at all.
$ prove -l t/05-util.t
t/05-util.t .. ok
All tests successful.Remove the compile test's BAIL_OUT — delete just the or BAIL_OUT(...) part — and deliberately break one module by deleting the trailing 1; from the end of Strata::Util, which makes require fail because a .pm file has to return a true value. Run the compile test both ways. Without BAIL_OUT, the broken module's own use_ok fails and the test file runs to completion regardless, reporting "1 of N failed". With BAIL_OUT restored, the run stops the instant the broken module is hit and never reaches anything after it. That is the right behaviour once prove -l t/ is running dozens of files: letting three hundred unrelated assertions fail because one module never loaded would bury the single line that actually explains what went wrong.
@EXPORT because it needs no qw(...) at the call site. It works right up until two modules you both use export a sub with the same name, and the caller has no way to tell which one it is running.1;. The module compiles, every sub in it is syntactically fine, and it still fails to require — a genuinely confusing error the first time you meet it, because the reported failure line is wherever the module was loaded, not wherever the missing statement should have been.cpanfile and Makefile.PL drift apart. Add a dependency to one and forget the other, and the failure shows up as "works on my machine" for whichever installer the other developer happens to use.requires 'XML::LibXML'; on its own answers that worse than nothing would.@EXPORT and @EXPORT_OK, and which should you use?cpanfile and a Makefile.PL?BAIL_OUT do and when is it appropriate?perldoc get its text from?JSON lines, CSV, XML and an always-succeeds fallback, all behind the same parser contract, with a registry that sniffs an unknown file and picks one.
A dispatch table, a duck-typed parser contract, confidence scoring rather than boolean detection, and the discovery that two of these formats are not line-oriented at all.
The contract every parser satisfies, which is enforced by nothing except tests and consistency:
| Method | Purpose |
|---|---|
new(%args), name | construction and identification |
mode | "line" or "stream" |
detect(@lines) | class method returning confidence 0..1 |
parse($line, %ctx) | line-mode parsers: always returns a record |
parse_handle($fh, $name, $cb) | stream-mode parsers: calls back per record |
stats | parsed, failed, problems by kind |
The line-by-line model from Milestone 4 is correct for logs and wrong for CSV and XML:
"Multi\nline name" is one record spanning two lines. Splitting on newlines before parsing produces two broken records and there is no way to recover.The wrong fix is to slurp those formats, which abandons the memory guarantee. The fix taken here is a second mode: a parser may declare mode "stream", receive the filehandle, and call back once per record. Fifteen lines in the pipeline, and no other code changes:
sub run_handle ($self, $fh, $name) {
my $parser = $self->{parser};
if ($parser->can("parse_handle")) {
$parser->parse_handle($fh, $name, sub ($parsed) {
$self->{stats}{lines}++;
$self->_emit($parsed);
});
return $self;
}
while (defined(my $line = <$fh>)) {
$self->{stats}{lines}++;
$self->_emit($parser->parse($line, file => $name, lineno => $.));
}
return $self;
}$parser->can("parse_handle") is duck typing done properly: ask the object what it can do rather than what it is. Extracting _emit so both paths share the record-wrapping, stage-running and statistics is what keeps the two modes from drifting apart.
my $data = eval { $JSON->decode($line) };
if (!defined $data) {
my $why = $@ // "unknown";
$why =~ s/\s+at\s+\S+\s+line\s+\d+.*//s; # trim Perl's location noise
$rec{decode_error} = $why;
return $self->_problem(\%rec, "bad_json");
}
if (ref $data ne "HASH") {
return $self->_problem(\%rec, "not_an_object");
}
# Flatten one level so nested objects become dotted keys, which keeps
# every record a flat set of fields whatever the service emitted.
_flatten($data, "", \%rec);
{"ctx":{"region":"eu"}} to ctx.region keeps every record from every format the same shape: a flat bag of named fields. Correlation across formats is only possible if the records are comparable.JSON::PP::Boolean needs explicit handling or your "false" is a blessed object that stringifies to "" in some contexts and 0 in others. The test asserts it becomes plain 0.not_an_object rather than crashing or inventing a field is the third outcome from Milestone 3 in a new setting." at /path/to/JSON/PP.pm line 62" off the error message matters: that location is inside a library the user did not write, and leaving it in makes every error report look like a bug in your tool. while (1) {
my $row = $csv->getline($fh);
unless ($row) {
last if $csv->eof;
# A broken row: report it with the text that caused it and
# resynchronise, rather than abandoning the rest of the file.
my ($code, $str, $pos) = $csv->error_diag;
my $bad = $csv->error_input // "";
$cb->({ ... problem => "csv_error", csv_error => "$code: $str at char $pos" });
$csv->SetDiag(0); # clear and keep going
last if eof($fh);
next;
}
...
}
Splitting on commas is wrong and everyone knows it; what people forget is that error recovery is the other reason to use a library. Text::CSV tells you the error code, the character position, and the exact input that failed, and SetDiag(0) clears the error so parsing can continue. A hand-rolled splitter gives you none of that, and the first malformed quote ends your run.
Separator sniffing scores consistency rather than counting commas:
sub sniff_separator ($class, @lines) {
my %score;
for my $sep (",", ";", "\t", "|") {
my %counts;
for my $line (@lines) {
my $n = () = $line =~ /\Q$sep\E/g;
$counts{$n}++;
}
my ($mode) = sort { $counts{$b} <=> $counts{$a} } keys %counts;
$score{$sep} = $mode ? $counts{$mode} * $mode : 0; # consistent AND present
}
...
}
The insight is that a real separator appears the same number of times on almost every line. Semicolons scattered through prose score badly because their count varies; the actual delimiter scores highly because it does not. \Q...\E quotes the separator so a | is a literal rather than alternation.
naive split (typical) Perl, with Text::CSV
────────────────────── ─────────────────────
my @fields = split /,/, $line; my $row = $csv->getline($fh);
# "Multi\nline name" already broke # multi-line quoted fields,
# this before split ever ran # embedded commas and quotes,
# a stray quote just becomes # all handled; a broken row
# one more ordinary character # reports where and why
split /,/ cannot be fixed into correctness with a cleverer regex, because the problem is not the separator — it is that one CSV "line" can span several physical lines when a field is quoted, and split only ever sees one physical line at a time. Text::CSV is a real state machine that reads as many physical lines as one logical record needs, and when a row does break, it hands back the error code, the character position, and the exact text that failed. Reimplementing that by hand, one edge case at a time, is how a production log parser accumulates a decade of regex patches and still gets embedded quotes wrong.
# XML is not line-oriented either, and a DOM parser would load the whole
# document into memory. XML::LibXML::Reader is a pull parser: it walks the
# document element by element with constant memory, which is the only
# defensible way to read an XML file of unknown size.
while (eval { $reader->read }) {
next unless $reader->nodeType == XML_READER_TYPE_ELEMENT;
if (!defined $record) {
next if $reader->depth == 0; # skip the root itself
$record = $reader->name; # the first child names the record
}
next unless $reader->name eq $record;
...
}
Three decisions worth noting. The record element is inferred from the first child of the root, so <alerts><alert/></alerts> needs no configuration. Attributes are prefixed with @ (@id, @severity) so they cannot collide with child elements of the same name, which is a real XML idiom. And recover => 2 asks libxml2 to continue after errors rather than aborting, which is the same policy as everywhere else in this project.
sub detect ($class, @) { return 0.01 } # always applicable, never preferred
A nonzero score means it is always a candidate; a tiny score means anything else beats it. That one line replaces a special case in the registry, and it is a nice demonstration of why scoring beats a boolean: "can you parse this?" has no good answer for a fallback, but "how confident are you?" does.
share/fixtures/access.log => apache apache=1.15 csv=0.00 json_lines=0.00 xml=0.00
share/fixtures/alerts.xml => xml apache=0.00 csv=0.00 json_lines=0.00 xml=1.10
share/fixtures/events.jsonl => json_lines apache=0.00 csv=0.33 json_lines=0.65 xml=0.00
share/fixtures/users.csv => csv apache=0.00 csv=0.53 json_lines=0.00 xml=0.00
share/fixtures/mixed.txt => unstructured apache=0.00 csv=0.00 json_lines=0.00 xml=0.00
Correct on all five. Then I ran the finished tool over the awkward fixtures, and two entries were wrong:
share/fixtures/hard/access.log.gz csv 0 records 19.7KB [gzip]
share/fixtures/hard/latin1.log apache 2 records 60B
Bug: sniffing looked at the compressed bytesThe gzipped Apache log was detected as CSV with zero records. The registry opened the file itself and scored the raw bytes, which for a gzip file are compressed noise that happens to contain a consistent number of commas. Meanwhile the Source was decompressing correctly, so the parser was reading real log lines and finding no CSV in them.
The fix is a rule worth generalising: sniff the stream the parser will actually read, not the file on disk.
# peek_lines opens the file, reads a few decoded lines, and closes it
# again. Sniffing has to see the data the parser will see: peeking at the
# raw bytes of a gzip file tells you nothing except that it is a gzip file.
sub peek_lines ($class, $path, $n = 20) {
my $src = eval { $class->open($path) } or return ();
my @lines;
while (@lines < $n && defined(my $line = $src->next_line)) { push @lines, $line }
$src->close;
return @lines;
}It decompresses the first few lines twice, once for sniffing and once for real, which is a fair price for correctness. The alternative, a pushback buffer in front of the handle, is more efficient and considerably more code.
latin1.log is two lines of prose with a bad byte in each. Every parser scored it zero except the fallback at 0.01, and then the .log extension added 0.15 to Apache, which won. The tool confidently parsed a text file as Apache and produced two failures.
# A filename extension is advice about a file whose content we have
# already scored. It may break a tie; it must never promote a format the
# content gave no support for, or every unreadable ".log" becomes Apache.
sub _apply_hint ($class, $score, $path) {
return unless $path =~ /(\.[A-Za-z0-9]+)$/;
my $hinted = $HINT{ lc $1 } or return;
return unless ($score->{$hinted} // 0) > 0.05;
$score->{$hinted} += 0.15;
return;
}The general principle: metadata may adjust a judgement, never create one. The same mistake appears whenever a system trusts a Content-Type header, a file extension or a user-declared schema over the bytes actually present.
After both fixes:
$ ./bin/strata ingest share/fixtures/*.* share/fixtures/hard/*
share/fixtures/access.log apache 2,004 records 206.1KB
share/fixtures/alerts.xml xml 2 records 347B
share/fixtures/events.jsonl json_lines 6 records 416B
share/fixtures/mixed.txt unstructured 3 records 177B
share/fixtures/users.csv csv 5 records 266B
share/fixtures/hard/access.log.gz apache 2,004 records 19.7KB [gzip]
share/fixtures/hard/binary.log unstructured 2 records 30B [contains NUL bytes]
share/fixtures/hard/bom.csv csv 2 records 25B [BOM (UTF-8)]
share/fixtures/hard/giant.log unstructured 3 records 2.9MB
share/fixtures/hard/latin1.log unstructured 2 records 60B
share/fixtures/hard/truncated.log apache 2 records 85B
4,035 records from 11 files
problems
bad_json 2
blank 1
missing_bytes 1
no_match 8
not_an_object 1
short_row 1
Sep 12 13:44:09 api-3 kernel: [88231.4] Out of memory, including the RFC 5424 variant with a priority prefix (<34>1 2026-09-12T13:44:09Z ...). Its detect must not claim ordinary prose.Strata::Parser::Multi that holds several parsers and, per line, uses the one whose detect on that single line scores highest, recording which. Measure what it costs.detect against every fixture and asserts the winner. Then add a fixture that is genuinely ambiguous (a CSV of JSON strings) and decide what the right answer is.1. The detection is the interesting half. Syslog's classic format is just a date followed by a hostname and a tag, which prose can accidentally resemble, so require the structure:
sub detect ($class, @lines) {
my $hits = grep {
/^<\d{1,3}>\d?\s/ # RFC 5424 priority
|| /^\w{3}\s+\d{1,2}\s\d{2}:\d{2}:\d{2}\s+\S+\s+\S+/ # classic: date host tag
} @lines;
return @lines ? $hits / @lines : 0;
}Note that both variants belong in one parser rather than two: they are the same log, and which one you get depends on the daemon's configuration, not on the file.
2. Multi is straightforward and the cost is not:
sub parse ($self, $line, %ctx) {
my ($best, $score) = ("unstructured", 0);
for my $name (keys %{ $self->{parsers} }) {
my $s = ref($self->{parsers}{$name})->detect($line);
($best, $score) = ($name, $s) if $s > $score;
}
$self->{used}{$best}++;
return $self->{parsers}{$best}->parse($line, %ctx);
}Running every parser's detect on every line multiplies the per-line regex work by the number of formats, and in my measurements that roughly halved throughput. The standard mitigation is stickiness: remember the format that matched the previous line and try it first, falling back to the full scan only when it fails. Real logs are strongly clustered, so this recovers most of the cost.
3. A CSV whose fields contain JSON is genuinely ambiguous, and the right answer is CSV: the outer structure wins, because parsing it as JSON lines fails on every line while parsing it as CSV succeeds and leaves the JSON as a field value that a later stage can decode. That suggests a general rule for the confusion matrix: when two formats both score well, prefer the one that is the outer container, and provide a stage that parses field values rather than making the file parser guess.
detect return a score rather than true or false?One module that turns a path into a stream of text and absorbs everything the real world does to files: gzip, byte-order marks, the wrong encoding, embedded NULs, truncation, and a single line three megabytes long.
Magic-byte detection, PerlIO layers, Encode fallbacks, and bounding the damage a hostile file can do.
Every one of these problems has been solved ad hoc, in the middle of a loop, in a thousand scripts. Putting them in one place with a name means the rest of the program can assume it is reading text. The fixtures were written first, deliberately:
| Fixture | What breaks without the Source |
|---|---|
access.log.gz | binary noise, or a special case at every call site |
bom.csv | the first column is named \x{feff}id and never matches |
latin1.log | a decode error, or silently corrupted text |
truncated.log | a half line silently dropped, or an infinite loop |
binary.log | NULs and control characters through your pipeline and terminal |
giant.log | 3 MB in one scalar; at scale, memory exhaustion |
open my $raw, "<:raw", $path or die "Strata::Source: $path: $!\n";
my $head = "";
read $raw, $head, PEEK_BYTES;
seek $raw, 0, 0;
$self->{compressed} = substr($head, 0, 2) eq "\x1f\x8b";
$self->{binary} = !$self->{compressed} && index($head, "\0") >= 0;
($self->{encoding}, $self->{bom_bytes}) = $class->_sniff_encoding($head);
Detect by content, not by name. \x1f\x8b is gzip's magic number whatever the file is called, and plenty of gzipped logs are named .log. A NUL byte in the first 4 KB means this is not text; that heuristic is what grep and git use and it is right far more often than it is wrong.
sub _sniff_encoding ($class, $head) {
return ("UTF-8", 3) if substr($head, 0, 3) eq "\xef\xbb\xbf";
return ("UTF-16LE", 2) if substr($head, 0, 2) eq "\xff\xfe";
return ("UTF-16BE", 2) if substr($head, 0, 2) eq "\xfe\xff";
return ("UTF-8", 0);
}
And then the byte count is used to skip it:
# Skip the byte-order mark so it does not appear in the first field
# of the first record, which is a bug you can stare at for an hour.
seek $fh, $self->{bom_bytes}, 0 if $self->{bom_bytes};
The BOM bug is worth dwelling on because it is so common and so invisible: your CSV's first header becomes \x{feff}id instead of id, every lookup of id returns undef, and the file looks perfect in every editor you open it in. Excel writes these by default.
The finding I did not expect: Perl's default replacement is not U+FFFDMy test asserted that an invalid byte becomes the replacement character. It failed, and the actual decoded characters were:
U+005C U+0078 U+0045 U+0039 which is the four-character text \xE9PerlIO's :encoding layer, by default, substitutes a literal escape sequence for bytes it cannot decode. Your data now contains a backslash, an x, and two hex digits, which will flow into your regexes, your database and your reports, and which no one will recognise as a decoding failure.
The fix is one line, and it is not well known:
unless ($self->{binary}) {
local $PerlIO::encoding::fallback = Encode::FB_DEFAULT;
binmode $fh, ":encoding($self->{encoding})";
}Verified both ways:
fallback=default -> U+005C U+0078 U+0045 U+0039 ("\xE9" as text)
fallback=0x0 -> U+FFFD U+0020 U+006E U+006F (the replacement character)With the fallback set, bad bytes become U+FFFD, which is greppable, countable, and universally understood to mean "something was lost here". Counting them is then trivial and more reliable than counting warnings, which the fallback suppresses:
if (index($line, "\x{fffd}") >= 0) {
my $n = () = $line =~ /\x{fffd}/g;
$self->{decode_warnings} += $n;
push @{ $self->{notes} }, "invalid bytes replaced" if $self->{decode_warnings} == $n;
}typical (wrapper object) Perl (PerlIO layers)
───────────────────────── ─────────────────────
fh = gzip.open(path, "rb") open my $fh, "<:raw", $path;
text = io.TextIOWrapper(fh, # detect gzip, then:
encoding="utf-8", binmode $fh, ":encoding($enc)";
errors="replace")
for line in text: ... while (<$fh>) { ... }
Both end up with a handle that yields decoded text regardless of what is compressing or encoding it underneath, but they get there differently. The typical approach composes by wrapping one object in another, and every wrapper adds a method-call layer of indirection to every read. Perl's version composes by stacking string labels onto one handle — :raw, then separately :encoding(UTF-8) — pushed and popped like a real stack, while <$fh> stays the same one operator throughout. The honest cost is that the stack is global mutable state attached to the handle itself: two pieces of code that both call binmode on the same handle can step on each other in a way two independently-scoped wrapper objects cannot, which is exactly why Strata::Source keeps every layering decision in one place instead of letting callers add their own.
use constant {
PEEK_BYTES => 4096,
MAX_LINE_BYTES => 1_048_576, # a "line" longer than 1 MB is not a line
};
if (length($line) > $self->{max_line_bytes}) {
$self->{long_lines}++;
push @{ $self->{notes} }, "over-long line truncated" if $self->{long_lines} == 1;
$line = substr($line, 0, $self->{max_line_bytes}) . "\n";
}
A file with no newlines at all is a single "line" the size of the file, and <$fh> will happily read all of it into one scalar. That is how a streaming tool runs out of memory on a file it never loaded. Any reader of untrusted input needs a line-length bound, exactly as a network protocol needs a maximum frame size, and the test proves that reading continues correctly after a truncation rather than losing the rest of the file.
subtest "invalid bytes are replaced, counted, and never fatal" => sub {
my $src = Strata::Source->open("$H/latin1.log");
my @lines;
push @lines, $_ while defined($_ = $src->next_line);
is scalar @lines, 2, "both lines were read";
like $lines[0], qr/caf\x{fffd} not found/, "the bad byte became U+FFFD";
is $src->stats->{decode_warnings}, 2, "and was counted";
like join(",", $src->notes), qr/invalid bytes replaced/, "the file is flagged";
};
subtest "one enormous line cannot exhaust memory" => sub {
my $src = Strata::Source->open("$H/giant.log", max_line_bytes => 64 * 1024);
my @lines;
push @lines, $_ while defined($_ = $src->next_line);
is scalar @lines, 3, "three lines";
is length($lines[1]), 64 * 1024 + 1, "the giant was truncated to the limit";
is $src->stats->{long_lines}, 1, "and counted";
is $lines[2], "after the giant\n", "reading continued correctly afterwards";
};
$ prove -l t/50-source.t
t/50-source.t .. ok
All tests successful.
Six subtests, each one a category of file that has ruined somebody's afternoon. Writing the hostile fixtures before the module is the technique: it converts "be robust" from an aspiration into a checklist, and every new disaster you meet in production becomes one more fixture and one more passing test.
open my $fh, "-|", "zstd", "-dc", $path) when no Perl module is available. Handle the case where the external tool is missing, and note what changes about error reporting when your reader is a child process.notes, and make it overridable.tell/seek support so an interrupted ingest can restart at the last committed byte offset. Explain why this is easy for a plain file and hard for a gzip stream, and what real tools do about it.1. The magic bytes are BZh, \xfd7zXZ and \x28\xb5\x2f\xfd. The interesting part is the external-process fallback:
open my $fh, "-|", $tool, "-dc", "--", $path
or die "Strata::Source: cannot run $tool: $!\n";
Three things change. The child's error output goes to your stderr unless you redirect it. The exit status arrives at close, not at open, so a tool that dies halfway through looks like a clean end of file unless you check close and $?. And the list form of open (arguments as separate strings, with --) is essential: the string form goes through the shell, so a file called ; rm -rf ~ would be executed. Never build a shell command from a filename.
2. Strict UTF-8 first is the right order, because valid UTF-8 is unlikely to occur by accident:
my $ok = eval { Encode::decode("UTF-8", $head, Encode::FB_CROAK); 1 };
if (!$ok) {
my $suspicious = () = $head =~ /[\x80-\x9f]/g; # cp1252 punctuation range
$encoding = $suspicious ? "cp1252" : "latin1";
push @notes, "encoding guessed: $encoding";
}The \x80-\x9f range is unassigned in Latin-1 and holds curly quotes and dashes in Windows-1252, so its presence is strong evidence. Say "guessed" in the notes: a guess recorded as a guess is useful, and a guess recorded as a fact is a future bug report.
3. For a plain file, remember tell after each committed record and seek back on restart. For gzip you cannot: the decompressor's state depends on everything before the current point, so a byte offset into the compressed file is meaningless without replaying it. Real systems solve this three ways: record the offset in the decompressed stream and re-read (correct, and costs a full decompression); use a format with sync points (bgzip, or multi-member gzip, which is why MultiStream => 1 is in our constructor); or checkpoint the decompressor state itself, which is what zran does for random access into gzip. Compression and random access are in tension, and the resolution is always a format decision rather than a code one.
:encoding layer insert for undecodable bytes by default, and how do you change it?Pull the things worth correlating out of records (addresses, emails, hosts, paths, identifiers), canonicalise them, and convert every timestamp dialect into one comparable number.
Match, validate, normalise as three distinct steps; context as the only defence against false positives; calendar arithmetic; and the syslog year problem.
A regex finds candidates. It does not find entities. Every extractor here is three parts:
ipv4 => {
pattern => $PATTERN{ipv4},
validate => sub ($v) { 4 == grep { $_ <= 255 && !/^0\d/ } split /\./, $v },
normalise => sub ($v) { join ".", map { 0 + $_ } split /\./, $v },
# A dotted quad preceded by "version" or "v" is a version number.
# No regex can tell 1.2.3.4 from an address on its own; only the
# surrounding text can, and even then only sometimes.
reject_after => qr/(?:version|ver|v)\s*$/i,
},
The pattern matches 999.1.1.1; the validator rejects it. The pattern matches 010.000.000.001; the validator rejects it, because leading zeros mean octal in some resolvers and are a classic filter-evasion trick. Normalisation is what makes correlation possible at all: the same address written differently in two files must become the same string.
Version numbers are indistinguishable from addresses. My test asserted that scanning a line containing version 1.2.3.4 would yield one address, and it yielded two. There is no regex that separates them, because they are the same string. Only context helps, hence reject_after, and even that fails on upgraded to 1.2.3.4 from 10.0.0.1. The real defence is the next point.
HTTP/1.1 is not a path. Scanning a raw Apache line for paths produced /api/export and /1.1, the second from the protocol version. The fix is a context rule with a rationale:
# A slash glued to the end of a word is part of that word:
# "HTTP/1.1" is a protocol, not a path to a file called 1.1.
reject_after => qr/\w$/,Both bugs point the same way, which is why from_record is built as it is:
# from_record looks in the fields a parser produced first, and only falls
# back to scanning the raw line for types the fields did not supply. A
# parsed field is evidence; a regex over the whole line is a guess.Prefer structure to scanning, always. The Apache parser already knows the client address and the request path; asking it is exact, and scanning the same line is an inference that will sometimes be wrong. Scanning is for the text nobody parsed, which is exactly where you need it and exactly where it is least reliable.
Records from five sources are only comparable if their timestamps are. That means every dialect becomes one integer, and the integer has to be right.
# 12/Sep/2026:13:44:10 +0000
if ($raw =~ m{^(\d{2})/(\w{3})/(\d{4}):(\d{2}):(\d{2}):(\d{2})(?:\s([+-])(\d{2})(\d{2}))?}) {
my $mon = $MONTH{$2} or return $class->_strptime_fallback($raw, %opt);
my $offset = defined $7 ? ($8 * 3600 + $9 * 60) * ($7 eq "-" ? -1 : 1) : 0;
return _epoch_utc($3, $mon, $1, $4, $5, $6) - $offset;
}
The test that matters asserts that different spellings of the same instant produce the same number:
my %cases = (
"12/Sep/2026:13:44:10 +0000" => 1789220650,
"12/Sep/2026:14:44:10 +0100" => 1789220650, # same instant, other zone
"2026-09-12T13:44:10Z" => 1789220650,
"2026-09-12 13:44:10" => 1789220650,
"2026-09-12T15:44:10+02:00" => 1789220650,
"2026-09-12T13:44:10.123Z" => 1789220650,
);
All six pass. That table is the specification of the module, and it is the kind of test that pays for itself the first time someone adds a format.
# Sep 12 13:44:10 -- syslog, with no year at all.
if ($raw =~ m{^(\w{3})\s+(\d{1,2})\s+(\d{2}):(\d{2}):(\d{2})$}) {
my $year = $opt{year} // (localtime)[5] + 1900;
my $epoch = _epoch_utc($year, $mon, $2, $3, $4, $5);
# If assuming this year puts the event more than a day in the
# future, the log almost certainly rolled over from December.
if (defined $opt{reference} && $epoch > $opt{reference} + 86_400) {
$epoch = _epoch_utc($year - 1, $mon, $2, $3, $4, $5);
}
return $epoch;
}
Syslog's classic format omits the year, so a December line read in January lands twelve months in the future, and your incident timeline puts the cause after the effect. The heuristic (a timestamp more than a day ahead of a known reference belongs to last year) is what every log tool does, it is right almost always, and it is wrong for clock-skewed hosts. Write the heuristic down where it is applied, and give the caller a way to override it.
My first version of the fast path pattern-matched the timestamp and then called Time::Piece->strptime to do the arithmetic. The comment in the code claimed it was "about ten times faster". Measured:
hand-written fast path: 1.30s (77,208/sec)
Time::Piece strptime: 0.98s (102,042/sec)
ratio: 0.8xIt was slower than the library it was meant to beat, because it did the same work plus a regex. The comment was an assumption I had written down as a fact, which is the most expensive kind of comment.
The fix was to remove the object entirely and compute the epoch arithmetically, using the standard branch-free calendar algorithm:
# days_from_civil: the standard branch-free calendar algorithm (Howard
# Hinnant's). Converting a date to a day number with arithmetic avoids
# constructing an object per line, which is what actually costs.
sub _days_from_civil ($y, $m, $d) {
$y -= $m <= 2;
my $era = int(($y >= 0 ? $y : $y - 399) / 400);
my $yoe = $y - $era * 400; # [0, 399]
my $doy = int((153 * ($m + ($m > 2 ? -3 : 9)) + 2) / 5) + $d - 1;
my $doe = $yoe * 365 + int($yoe / 4) - int($yoe / 100) + $doy;
return $era * 146_097 + $doe - 719_468;
}
arithmetic fast path: 0.47s (425,794/sec)
Time::Piece strptime: 1.90s (105,445/sec)
speedup: 4.0x
Four times faster, and every correctness test still passes, which is the only reason the rewrite was safe to attempt. Three lessons, in order of importance: a comment claiming a speedup is a claim, and claims get measured; the cost was object construction rather than parsing, which the benchmark told me and intuition did not; and a table of correctness tests written before the optimisation is what turns a risky rewrite into a routine one.
$ ./bin/strata entities share/fixtures/access.log share/fixtures/mixed.txt
ipv4 (4 distinct)
10.14.22.9 1,022
10.14.22.31 344
192.168.4.7 324
172.16.0.99 312
path (6 distinct)
/api/export 343
/api/search 336
/static/app.js 323
/papers 320
/login 318
activity by hour
2026-09-12T13:00:00Z 1,418
2026-09-12T14:00:00Z 585
Entities ranked across two files of different formats, and a histogram over a time axis that did not exist until this milestone. That last block is the foundation of Milestone 9: once every record has a comparable instant, "what else happened within ninety seconds of this" becomes a query rather than a research project.
Exercise 8raw with stable pseudonyms (the same input always giving the same token) so that records can be shared. Use a keyed hash, and explain why an unkeyed hash of an email address is not anonymisation.tz option, use it when no offset is present, and write a test with a log recorded across a daylight-saving transition. Decide what to do about the hour that occurs twice.1. The worry list is the exercise. For URLs it is trailing punctuation: see http://example.com/api. ends with a full stop that is not part of the URL, so strip trailing .,;:)]} and document it. For IPv6 it is that :: can be expanded in only one place and the canonical form (RFC 5952) requires lowercase hex, no leading zeros, and the longest run of zero groups compressed; getting this wrong means the same address appears as two entities. For card numbers it is that a Luhn-valid sixteen-digit string is also a valid order number, so a match should be reported as a possible card, never as a certainty.
2. An unkeyed hash of an email address is not anonymisation because the input space is small enough to enumerate: an attacker with your redacted file hashes their own address list and matches. Use a keyed hash with a secret the recipient does not have:
use Digest::SHA qw(hmac_sha256_hex);
sub pseudonym ($value, $key) { "tok_" . substr(hmac_sha256_hex($value, $key), 0, 16) }Stable (the same input gives the same token, so correlation still works), unlinkable without the key, and per-dataset if you rotate the key. Note what it still leaks: frequency. If one token appears 90% of the time, its identity may be inferable from context regardless of the hash.
3. The ambiguous hour is genuinely unresolvable from the data: 01:30 occurs twice on the night the clocks go back, and a local-time log with no offset cannot say which. The defensible options are to pick the first occurrence and flag the record, or to mark the timestamp as ambiguous and let correlation treat it as a range. What you must not do is pick one silently, because an hour of duplicated timestamps in an incident timeline is exactly the kind of thing that sends an investigation down a wrong path. This is also the strongest possible argument for logging in UTC with an explicit offset, which is worth saying in your tool's documentation.
reject_after defend against, and why can it never be a complete defence on its own?from_record prefer a parser's own fields over scanning the raw line?Time::Piece. What was it actually spending its time on, and what fixed it?:encoding gives you U+FFFD for bad bytes.Milestone 7 is Perl at its best in a way that is hard to see unless you have written the equivalent elsewhere. PerlIO layers mean compression, encoding and buffering compose as a stack on a handle rather than as wrapper objects, so binmode $fh, ":encoding(UTF-8)" after opening a gzip stream is the whole of "decompress then decode". The fixtures, the magic-byte checks and the line bound are about a hundred lines total.
Milestone 6 is more mixed. The dispatch table and duck-typed contract are pleasant, and $parser->can("parse_handle") is exactly the right amount of ceremony. But nothing checks that a parser actually implements the contract, so a plugin missing stats fails at run time, in production, on the one file that used it. Go's interfaces or Ruby's contract tests both handle this better; the Perl answer is a test module that every parser's test file uses, which is Milestone 11.
Milestone 8 is the honest low point for the language, and the high point for the discipline. Perl gives you nothing for entity extraction that another language would not; the value is entirely in the three-step match-validate-normalise structure, the context rules, and the tests that caught two false positives. The 4× timestamp speedup came from writing arithmetic instead of using a library, which is a technique available everywhere.
strata/
├── Makefile.PL, cpanfile declared dependencies, installable
├── bin/strata ingest | entities
├── lib/Strata.pm POD, $VERSION
├── lib/Strata/
│ ├── Util.pm commify, top, human_bytes, truncate_str
│ ├── Pattern.pm named composable qr// building blocks
│ ├── Record.pm fields, provenance, problems, entities
│ ├── Pipeline.pm line and stream modes, stages, stats
│ ├── Source.pm gzip, BOMs, encodings, NULs, giant lines
│ ├── Extract.pm match -> validate -> normalise, with context
│ ├── Normalize.pm every timestamp dialect -> one integer
│ └── Parser/
│ ├── Registry.pm dispatch table and confidence scoring
│ ├── Apache.pm JsonLines.pm Csv.pm Xml.pm Unstructured.pm
└── t/ 7 files, 46 tests
└── share/fixtures/hard/ six files designed to break the reader
$ prove -l t/
All tests successful. Files=7, Tests=46
$ git commit -am "milestones 5-8: distribution, formats, hardened reading, entities"
Continue