Instalment 11 · Course 3 (Perl) · Parts 0–2
Go was about concurrency and Ruby about expressiveness. Perl is about the hour you spend with ten gigabytes of logs that nobody documented, half of which are truncated, and a question you need answered by lunchtime.
Every example below was run on Perl 5.38.2 and the outputs are copied from those runs. The modules used in the project (Text::CSV, DBI, DBD::SQLite, XML::LibXML, Try::Tiny) are installed in my sandbox, so the later milestones will be verified too.
A command-line tool called strata that you point at a directory of unexplained files and interrogate.
$ strata ingest ./incident-2026-09-12/ --recursive
apache/access.log 412,884 lines apache_combined 3.2s
apache/error.log 18,221 lines apache_error 0.4s
app/service.log 904,110 lines json_lines 8.1s
exports/users.csv 9,412 rows csv 0.3s
config/nginx.conf 420 lines nginx_config 0.0s
unknown/dump.txt 33,900 lines unstructured 1.1s
corrupt/partial.log 2,004 lines apache_combined 0.1s (91 malformed, kept)
1,381,051 records, 214,882 entities, 46,203 events in 13.2s (104k lines/sec)
$ strata entities --type ip --top 5
10.14.22.9 88,214 occurrences 6 files first 13:02:11 last 14:47:52
10.14.22.31 41,002 occurrences 4 files ...
$ strata timeline --entity ip:10.14.22.9 --around '13:44:10' --window 90s
13:43:58 apache/access.log:88214 GET /api/export 200 1.2MB
13:44:02 app/service.log:551203 export.start user=4412 rows=900000
13:44:09 apache/error.log:9902 upstream timed out
13:44:10 app/service.log:551288 ERROR OOM killed worker pid=8823
13:44:11 apache/access.log:88240 GET /api/export 502
$ strata graph --entity user:4412 --depth 2 --format dot | dot -Tsvg > incident.svg
Underneath: a streaming ingestion pipeline with pluggable format detectors, an entity extractor, a normaliser that turns eleven timestamp formats into one, a correlation engine that groups records into events, and a SQLite-backed store you can query, all of it able to survive files that are truncated, mis-encoded, or simply lying about their format.
Most data-processing tutorials use clean data, which is the one thing you will never be given. Real forensic work looks like this: eight formats, three of them undocumented, timestamps in four time zones, a log rotated mid-write so one line is half of two lines, an "XML" file that is actually XML fragments concatenated without a root element, and a CSV whose quoting breaks on row 40,000 because someone's surname contains a comma and a quotation mark.
The interesting engineering is not parsing any one format. It is building something that keeps going: that treats malformed input as expected rather than exceptional, that tells you exactly what it could not parse and why, and that never loads more than one line into memory so the same tool works on a 2 KB file and a 200 GB one.
if ($line =~ /^(\S+) (\S+)/) versus Python's m = re.match(r'^(\S+) (\S+)', line); if m: m.group(1). One of those disappears into the code and one of them is visible machinery. Over ten thousand lines of parsing that difference compounds. while (<>) { ... } reads standard input or every file named on the command line, one line at a time, with no imports and no ceremony. Your program is a filter before you have decided to write one.my @all = $text =~ /(\w+)@([\w.]+)/g; collects every capture from every match into a list, in one expression.perl -lane 'print $F[6] if $F[8] == 500' access.log is a complete program you type into a terminal while exploring, and the language you explore in is the language you write the tool in.
For the middle of this project (streaming, regex-heavy, line-oriented transformation with Unix plumbing), Perl is still the best tool in existence, and the reason is not nostalgia: no other language has put regular expressions, context, and the input loop into the syntax itself.
Where it is not the answer, and we will say so at the time:
use v5.36, named captures, /x, real data structures, and tests.The honest summary: Perl is the best language I know for the first two hours with unknown data, and you should be prepared to hand the result to something else afterwards. Knowing when to hand over is part of what this course teaches.
files (any shape, any size, some broken)
│
▼
┌───────────────┐ one line at a time, constant memory
│ Source │ handles gzip, encodings, truncation
└───────┬───────┘
▼
┌───────────────┐ sniffs the format, or is told
│ Detector │ apache | nginx | json_lines | csv | xml | conf | code | unknown
└───────┬───────┘
▼
┌───────────────┐ a plugin per format, each returning a Record
│ Parsers │ malformed lines become Records with a `problem` field,
└───────┬───────┘ never exceptions and never silent drops
▼
┌───────────────┐ ips, emails, paths, ids, versions, timestamps...
│ Extractors │ each is a named regex plus a normaliser
└───────┬───────┘
▼
┌───────────────┐ one canonical timestamp, one canonical host,
│ Normaliser │ one canonical path, whatever the source said
└───────┬───────┘
▼
┌───────────────┐ records near in time and sharing an entity
│ Correlator │ become events; events referencing each other
└───────┬───────┘ become a graph
▼
┌───────────────┐
│ Store │ SQLite: records, entities, occurrences, events, edges
└───────┬───────┘
▼
query · timeline · graph · report (and stdout, for piping)
| # | Milestone | What it teaches |
|---|---|---|
| 1 | A filter that counts what it reads | scalars, while (<>), $_, strict and warnings, exit codes |
| 2 | Summaries: top talkers, error rates | hashes, sorting, context, output formatting |
| 3 | An Apache/nginx parser | regexes, named captures, /x, qr//, failure counting |
| 4 | A Record model and a pipeline | references, nested data, subs, Data::Dumper |
| 5 | Splitting into modules | packages, lib/, cpanm, cpanfile, Test::More, prove |
| 6 | Pluggable formats: CSV, JSON, XML | dispatch tables, CPAN modules, format sniffing |
| 7 | Streaming 10 GB safely | encodings, gzip, truncation, memory discipline, malformed-input policy |
| 8 | Entity extraction and normalisation | IPs, timestamps, paths, ids; canonical forms; time zones |
| 9 | Correlation and sessionisation | time windows, joins, SQLite via DBI, indexes |
| 10 | A proper command-line tool | Getopt::Long, exit codes, signals, pipes, --help |
| 11 | Testing, fuzzing and profiling | Test::More, corrupt-input tests, Devel::NYTProf, optimisation |
| 12 | The knowledge graph | graph storage and queries, fork for parallelism, packaging |
How to build a streaming pipeline that survives input designed to break it; how to write regexes that a colleague can read a year later; why Perl's context rule exists and how to stop it surprising you; how to correlate events across sources with nothing but a time window and a shared identifier; and when to stop and hand the cleaned data to something else.
Perl is a dynamically typed, garbage-collected interpreted language from 1987, designed explicitly for report processing and system glue. It is not "Perl 6": that language was renamed Raku in 2019 and is a different thing. Perl 5 is alive, gets an annual release, and version 5.36 (2022) and 5.38 (2023) modernised it substantially, with subroutine signatures and a real class syntax.
Two facts shape everything about the language. First, values have no fixed type but variables have a fixed shape, marked by a sigil: $ for a single value, @ for a list, % for a key-value table. Second, every expression is evaluated in a context, either "give me one value" or "give me a list", and many operations do different things depending on which. Nothing else in mainstream programming works this way, and both are covered properly in Part 2.
Unlike Ruby, using the system Perl is usually fine: it is well-maintained, present everywhere, and most distributions keep it current enough. Perl's compatibility record is unusually good, so a script from 2010 generally still runs. You want 5.36 or newer to get signatures and the modern feature bundle; I verified everything on 5.38.2.
perl -v # almost certainly already there
sudo apt install perl cpanminus # Debian/Ubuntu: adds the cpanm installer
sudo dnf install perl perl-App-cpanminus # Fedora
# many CPAN modules exist as distribution packages, which is the easiest path:
sudo apt install libtext-csv-perl libdbd-sqlite3-perl libxml-libxml-perl libtry-tiny-perl
perl -v # Apple ships one, often a version or two behind
brew install perl # newer, installed under /opt/homebrew
:: Strawberry Perl bundles a C compiler, so modules with XS code build
winget install StrawberryPerl.StrawberryPerl
Strawberry Perl is genuinely good. That said, this course leans on Unix habits (pipes, STDIN, signals, file permissions) more than any other in the curriculum, and WSL2 will save you real friction.
If you need a version the system does not have, or you are installing modules on a machine where you have no root, use a version manager:
curl -L https://install.perlbrew.pl | bash # perlbrew
perlbrew install perl-5.38.2
perlbrew switch perl-5.38.2
# or plenv, if you liked rbenv
Perl's tooling is not as unified as a single-vendor language's, but two language-server projects cover most editors well enough to set up before Part 2, plus a static analyser and a formatter worth knowing about even before you wire either into anything.
# either one, from the Extensions panel or the command line:
code --install-extension bscan.perlnavigator # PerlNavigator: lighter, installs in seconds
code --install-extension richterger.perl # Perl::LanguageServer: deeper, needs a C toolchain
PerlNavigator is a TypeScript language server that shells out to perl -c and perlcritic for you, giving syntax checking, outlining and completion with almost no setup. Perl::LanguageServer goes further — real go-to-definition, variable inspection, a working debugger — because its backend is itself a Perl module, an XS one, which means it needs cpan Perl::LanguageServer to succeed on your machine before the extension works at all, and an XS build occasionally means fighting a missing compiler or a stale header before it does. Start with PerlNavigator; reach for Perl::LanguageServer only once you miss real navigation.
Neovim's built-in LSP client (via nvim-lspconfig) can drive either server above directly, the same way it would for any other language server. Plain Vim's long-standing option is perl-support.vim, for syntax, skeletons and running the current file, without full language-server intelligence. Emacs ships cperl-mode, the conventional choice over the older bundled perl-mode: it indents heredocs and POD correctly, which the older mode does not, and eglot or lsp-mode can point it at either server above for full completion.
perlcritic and perltidy: the closest thing Perl has to a linter and a formattercpanm Perl::Critic Perl::Tidy
perlcritic bin/strata # static analysis, default severity
perlcritic --brutal bin/strata # same file, strictest severity level
perltidy -b lib/Strata/Record.pm # reformat in place; -b keeps a .bak
perlcritic checks code against configurable rules drawn largely from Damian Conway's Perl Best Practices, at five severity levels from --gentle (only the most serious issues) to --brutal (everything); a project settles on a level and a .perlcriticrc once, rather than treating every possible warning as mandatory from day one. perltidy reformats to a configurable style the way gofmt does for Go, except it is not bundled with the interpreter and not run on every commit by default the way gofmt effectively is — both are a five-minute cpanm install away, and neither is the default a newer language's tooling would make it.
CPAN is Perl's module archive, and at around 220,000 modules it is one of the oldest and deepest package ecosystems in existence. The modern client is cpanm:
cpanm Text::CSV # install one module
cpanm --installdeps . # install everything a cpanfile lists
cpanm --local-lib=~/perl5 DBI # install into your home directory, no root
# then tell perl where to find them (add to ~/.bashrc):
eval "$(perl -I ~/perl5/lib/perl5 -Mlocal::lib)"
Never use sudo cpanm. It mixes modules you installed with modules your operating system manages, and the resulting breakage is tedious. Use local::lib, perlbrew, or your distribution's packages.
For a project, declare dependencies in a cpanfile:
# cpanfile
requires 'perl', '5.036';
requires 'Text::CSV', '2.00';
requires 'DBI';
requires 'DBD::SQLite';
requires 'Try::Tiny';
on 'test' => sub {
requires 'Test::More', '1.302';
};
Then cpanm --installdeps . installs them. If you want a lockfile and a bundled dependency tree (Bundler's role), that is Carton: carton install writes cpanfile.snapshot and carton exec runs with exactly those versions.
CPAN is the strongest argument for doing forensic text work in Perl at all: at roughly 220,000 modules covering three decades of "someone already hit this exact format and fixed it", the odds that today's undocumented log shape already has a parser somewhere are good. That is not spin — Text::CSV and XML::LibXML, which this project leans on from Milestone 6 onward, are mature, heavily-used modules, not toy packages someone abandoned in 2009.
Two honest costs sit right next to that strength. First, a real fraction of CPAN's best modules — XML::LibXML among them — are XS: Perl wrapped around C, which means installing one sometimes means having a working compiler and system headers rather than just downloading a file, and it is the single most common reason a cpanm install fails on a machine that "should" work. Second, cpanm --installdeps . alone gives you no lockfile: two checkouts on two machines can silently resolve a cpanfile's version ranges to two different sets of module versions unless you additionally reach for Carton, which is opt-in rather than the default the way Gemfile.lock or package-lock.json are. And perlbrew, while it works, compiles an entire Perl interpreter from source for every version you install — it is closer in spirit to pyenv or rbenv than to a version manager that downloads a prebuilt binary, so "just try 5.38" costs several minutes the first time, not seconds. None of this is disqualifying; all of it is worth knowing before the first cpanm command mysteriously fails on someone else's laptop.
perldoc -f split # one built-in function, with examples
perldoc perlre # the regex reference
perldoc perlretut # the regex tutorial
perldoc perlrequick # the regex quick start
perldoc perlvar # what $_, @ARGV, $! and the rest mean
perldoc perldsc # data structures cookbook: the one to read twice
perldoc perlop # operators and precedence
perldoc -q "how do I sort" # search the FAQ
perldoc Text::CSV # any installed module's own documentation
perldoc -l Text::CSV # where that module actually lives on disk
Perl's documentation ships with the interpreter, is written by people who use the language, and is complete. perldoc perlretut alone is a better regex tutorial than most books. Get into the habit of reading it in a terminal rather than searching the web, because the web has thirty years of outdated Perl advice on it and perldoc has none.
strata/
├── cpanfile dependencies
├── bin/
│ └── strata the command-line program (no .pl extension)
├── lib/
│ └── Strata/
│ ├── Record.pm Strata::Record
│ ├── Source.pm Strata::Source
│ └── Parser/
│ ├── Apache.pm Strata::Parser::Apache
│ └── CSV.pm
├── t/
│ ├── 00-load.t does everything compile?
│ ├── 10-record.t
│ └── 20-parser-apache.t
└── share/
└── fixtures/ sample logs, including deliberately broken ones
Three conventions, all enforced by tooling rather than taste:
Strata::Parser::Apache must live at lib/Strata/Parser/Apache.pm. use lib 'lib' or perl -Ilib puts lib/ on the search path (@INC).t/ and end in .t, and are run by prove. Numbering them controls order.1;. A module must return a true value or use fails, and a bare 1; is the conventional way. Forgetting it is every Perl programmer's first confusing error.#!/usr/bin/env perl
use v5.36;
my $name = shift @ARGV // "world";
say "hello, $name";
say "perl $^V, script $0, pid $$";
$ perl hello.pl
hello, world
perl v5.38.2, script hello.pl, pid 18244
$ perl hello.pl strata
hello, strata
#!/usr/bin/env perl — the shebang, so ./hello.pl works once the file is executable. /usr/bin/env perl finds whichever Perl is first on your PATH, which is what you want with perlbrew.
use v5.36; — the single most important line in modern Perl. It enables a whole feature bundle at once:
| What it turns on | Why it matters |
|---|---|
use strict | Undeclared variables are a compile error, not a new global |
use warnings | Undefined values, dubious conversions and typos in comparisons get reported |
say | print with a newline |
| subroutine signatures | sub f ($x, $y = 1) { ... } instead of unpacking @_ by hand |
isa, postfix deref, and more | modern conveniences without per-feature imports |
| disables indirect object syntax | removes a genuinely confusing old parsing rule |
Old Perl tutorials start with use strict; use warnings; as two separate lines. use v5.36; includes both and more, and it also declares your minimum Perl version, so running on something older fails immediately with a clear message rather than mysteriously.
my $name = shift @ARGV // "world"; — my declares a lexical variable, scoped to the enclosing block. @ARGV holds the command-line arguments (unlike C, it does not include the program name, which is in $0). shift removes and returns the first element. // is the defined-or operator: use the left side unless it is undef. It differs from || in that 0 and "" are kept, which matters constantly when parsing data where zero is a real value.
say "hello, $name" — double-quoted strings interpolate variables, including array and hash elements. Single quotes do not interpolate anything.
$^V, $0, $$ — Perl's punctuation variables: the version, the program name, the process id. There are dozens; perldoc perlvar lists them all, and Part 2 covers the six you actually need.
Before writing a script, Perl programmers interrogate data from the shell. This is a real part of the language's culture and the fastest way to build intuition:
# print lines matching a pattern (grep, but with Perl regexes)
perl -ne 'print if /ERROR/' access.log
# -l adds newline handling, -a splits each line into @F on whitespace
perl -lane 'print $F[0] if $F[8] == 500' access.log
# count by field: the Perl idiom you will use a thousand times
perl -lane '$c{$F[0]}++; END { print "$c{$_}\t$_" for sort { $c{$b} <=> $c{$a} } keys %c }' access.log
# in-place edit with a backup
perl -i.bak -pe 's/\bDEBUG\b/TRACE/g' service.log
# -F sets the split pattern: CSV-ish, badly, but instantly
perl -F, -lane 'print $F[2] if $F[4] > 100' export.csv
| Flag | Meaning |
|---|---|
-e | the program is on the command line |
-n | wrap it in while (<>) { ... } |
-p | same, and print $_ at the end of each iteration |
-l | strip the newline on input, add one on output |
-a | autosplit each line into @F |
-F | the pattern to autosplit on |
-i | edit files in place |
That third one-liner is a complete top-talkers report, and it is the seed of Milestone 2. Perl's design makes the exploratory version and the production version the same language, which is exactly what you want when the exploration turns out to be the tool.
use v5.36;
use lib "lib";
use Test::More tests => 4;
use Strata::Record;
my $r = Strata::Record->new(source => "apache", fields => { ip => "10.0.0.1" });
isa_ok $r, "Strata::Record";
is $r->source, "apache", "source is kept";
is $r->field("ip"), "10.0.0.1", "fields are readable";
like $r->to_line, qr/ip=10\.0\.0\.1/, "to_line renders fields";
$ perl -Ilib t/10-record.t
1..4
ok 1 - An object of class 'Strata::Record' isa 'Strata::Record'
ok 2 - source is kept
ok 3 - fields are readable
ok 4 - to_line renders fields
$ prove -l t/
t/10-record.t .. ok
All tests successful.
Files=1, Tests=4, 0 wallclock secs
Result: PASS
That output format is TAP, the Test Anything Protocol, which Perl invented in 1987 and which now has implementations in most languages. A test file is an ordinary program that prints ok and not ok lines; prove runs many of them and summarises. prove -l adds lib/ to the path, prove -lv shows every assertion, and prove -lj4 runs four files in parallel.
Write bin/logstat, a program that reads log lines from files named on the command line (or from standard input when none are named) and prints: the total line count, how many lines contain ERROR, and the percentage. It must exit 0 when there are no errors and 1 when there are, so it can be used in a shell conditional.
Then prove it works in a pipeline: cat *.log | ./bin/logstat and ./bin/logstat a.log b.log should both work without changing the code.
Hints: while (<>) handles both cases for free; $. is the current line number; exit sets the status; printf formats the percentage.
#!/usr/bin/env perl
use v5.36;
my ($lines, $errors) = (0, 0);
while (my $line = <>) {
$lines++;
$errors++ if $line =~ /\bERROR\b/;
}
if ($lines == 0) {
warn "logstat: no input\n";
exit 2;
}
printf "%d lines, %d errors (%.2f%%)\n", $lines, $errors, 100 * $errors / $lines;
exit($errors > 0 ? 1 : 0);
$ printf 'ERROR disk full\nINFO started\nERROR timeout\nWARN slow\n' | ./bin/logstat
4 lines, 2 errors (50.00%)
$ echo $?
1
Four things worth extracting.
while (my $line = <>) is the entire input story. No arguments means standard input; arguments mean read each of those files in turn. Your program is a Unix filter with no extra work, which is why cat x | prog and prog x both work.\b word boundaries matter. Without them, a line containing NOERROR or ERRORS_TOTAL=0 counts as an error. Ninety per cent of wrong log analysis is a missing word boundary.grep convention, and following it means your tool composes with shell scripts people already know.warn writes to standard error, print and say to standard output. Keeping diagnostics off stdout is what lets someone pipe your output into another program.One subtlety: while (<>) without assigning to a variable puts the line in $_, which is idiomatic and slightly risky, because anything you call inside the loop might also use $_. Assigning to a named variable, as here, is the safer habit in anything longer than a one-liner.
Common first-day errorsCan't locate Strata/Record.pm in @INC — lib/ is not on the search path. Use perl -Ilib, use lib 'lib';, or prove -l.Strata/Record.pm did not return a true value — you forgot the 1; at the end of the module.Global symbol "$count" requires explicit package name — use strict is doing its job: you forgot my.Use of uninitialized value in ... — a warning, not an error, and almost always a real bug: you used a value that was never set, usually a capture group from a match that failed.Can't call method "new" on an undefined value — you forgot to use the module that defines the class.syntax error at ... near "}" — a missing semicolon on the previous line. Perl's error points at where parsing failed, not where you went wrong.== to compare strings. == is numeric; eq is for strings. "abc" == "def" is true, because both convert to 0.use v5.36; switch on, and why is it better than use strict; use warnings;?1;?// and ||, and when does it matter?while (<>) read from?sudo cpanm?perldoc pages would you open first for a regex question?Only what the project needs, which is most of Perl's text machinery and none of its formats or tie magic. Every output below is from a real run. Keep a terminal open and type them.
use v5.36;
my $count = 42;
my @steps = ("fetch", "filter", "save");
my %options = (topic => "AI", limit => 10);
say "scalar: $count";
say "array has ", scalar(@steps), " elements, last index $#steps";
say "element: $steps[0], slice: @steps[0,1]";
say "hash value: $options{topic}, keys: ", join(",", sort keys %options);
my $n = @steps; # array in scalar context: its length
my ($first) = @steps; # list context: its first element
say "scalar context gives $n, list context gives $first";
scalar: 42
array has 3 elements, last index 2
element: fetch, slice: fetch filter
hash value: AI, keys: limit,topic
scalar context gives 3, list context gives fetch
Two rules explain that output, and together they are the thing that makes Perl feel alien for a week and obvious afterwards.
The sigil describes what you are asking for, not what the variable is. @steps is the whole array; $steps[0] is one scalar from it, so it takes $; @steps[0,1] is several, so it takes @. Likewise %options is the hash and $options{topic} is one value from it. The rule is consistent, and it is the opposite of what people assume ("@ means array"), which is why $steps[0] looks wrong to newcomers.
Every expression is evaluated in scalar or list context, and many behave differently in each. my $n = @steps asks for one value from an array, so you get its length; my ($first) = @steps has a list on the left, so the array is unpacked and the first element assigned. Those two lines differ only in parentheses and mean completely different things.
my @words = split /,/, "a,b,c";
my $words = split /,/, "a,b,c"; # scalar context: count
say "list: @words / scalar: $words";
say "reverse in list: ", join("", reverse @words);
say "reverse in scalar: ", scalar reverse "hello";
list: a b c / scalar: 3
reverse in list: cba
reverse in scalar: olleh
reverse reverses a list in list context and a string in scalar context. This is not a special case; it is the design. When a Perl function surprises you, the first question is always "what context is this in?", and perldoc -f reverse will tell you what it does in each.
In Python, len(xs) is a function call and there is no way for an expression to know whether its caller wants one value or many. Perl passes that information down, which makes code shorter (my ($x) = f() versus x = f()[0]) and makes a category of bug possible that exists nowhere else (my $x = f() quietly giving you a count instead of a value). The mitigation is the same as the diagnosis: when a value is wrong in a confusing way, print scalar(@thing) and check which context you are in.
my @nums = (5, 3, 9, 1);
push @nums, 7;
my $popped = pop @nums;
say "after push/pop: @nums (popped $popped)";
say "sorted numerically: ", join(",", sort { $a <=> $b } @nums);
say "sorted as strings: ", join(",", sort @nums);
say "grep > 3: ", join(",", grep { $_ > 3 } @nums);
say "map doubled: ", join(",", map { $_ * 2 } @nums);
my @spliced = splice(@nums, 1, 2);
say "spliced out @spliced leaving @nums";
after push/pop: 5 3 9 1 (popped 7)
sorted numerically: 1,3,5,9
sorted as strings: 1,3,5,9
grep > 3: 5,9
map doubled: 10,6,18,2
spliced out 3 9 leaving 5 1
sort compares as strings by default. The two sorts agree here by luck; try (5, 30, 9) and the default gives 30, 5, 9. Always write the comparator: { $a <=> $b } numeric, { $a cmp $b } string. $a and $b are package globals the sort block sees, which is why they need no my.grep and map set $_ to each element. grep keeps the elements whose block is true; map returns whatever the block returns, and a block returning two values makes the result longer than the input, which is occasionally exactly what you want.splice removes and optionally replaces a range in place. It is the general-purpose array surgery tool."@nums") joins with spaces. A hash does not interpolate at all.my %count;
$count{$_}++ for qw(fetch filter fetch save fetch);
for my $k (sort { $count{$b} <=> $count{$a} || $a cmp $b } keys %count) {
say " $k: $count{$k}";
}
say "exists: ", (exists $count{fetch} ? "yes" : "no");
delete $count{fetch};
say "after delete: ", join(",", sort keys %count);
my @wanted = @count{qw(filter save)}; # hash slice
say "slice: @wanted";
fetch: 3
filter: 1
save: 1
exists: yes
after delete: filter,save
slice: 1 1
$count{$_}++ for @things is the most-used line in Perl, and it is worth unpacking completely. %count starts empty. $count{$_} on a missing key is undef, and ++ on undef treats it as 0 and makes it 1, with no warning because incrementing undef is explicitly allowed. The for at the end is a statement modifier: a postfix loop, readable precisely because it is short.
The sort block chains two comparisons with ||: compare counts descending ($count{$b} <=> $count{$a}), and when they are equal (<=> returns 0, which is false) fall through to comparing keys alphabetically. That is the standard multi-key sort and it appears in every report you will ever write.
exists versus truth versus defined are three different questions: is the key there, is the value true, is the value not undef. A key with value 0 exists, is defined, and is false.sort keys %h when printing.@count{qw(filter save)} is a hash slice: several values at once, so the sigil is @.Arrays and hashes can only hold scalars, so nesting requires references: scalars that point at something.
my @list = (1, 2, 3);
my %opts = (a => 1);
my $aref = \@list; # reference to an existing array
my $href = \%opts;
my $anon = [ { name => "fetch", options => { from => "arxiv" } } ]; # anonymous
say "deref whole: @$aref / @{$aref}";
say "element: $aref->[0] and $$aref[0]";
say "nested: $anon->[0]{name} -> $anon->[0]{options}{from}";
say "ref types: ", join(",", map { ref } ($aref, $href, $anon, sub {}, \"x"));
push @$aref, 4;
say "the original array sees it: @list";
my %index;
push @{ $index{ai} }, "paper1"; # autovivification
push @{ $index{ai} }, "paper2";
say "autovivified: ", join(",", @{ $index{ai} });
deref whole: 1 2 3 / 1 2 3
element: 1 and 1
nested: fetch -> arxiv
ref types: ARRAY,HASH,ARRAY,CODE,SCALAR
the original array sees it: 1 2 3 4
autovivified: paper1,paper2
Four rules cover almost everything:
\ takes a reference; [ ... ] and { ... } create anonymous arrays and hashes directly. Use the anonymous forms in data structures.-> both dereferences and indexes. $aref->[0], $href->{key}. The older $$aref[0] means the same thing and is harder to read.$anon->[0]{name}{x} is the same as $anon->[0]->{name}->{x}. Write the first arrow, omit the rest; everyone does.@$aref, %$href, @{ $index{ai} }. The braces are for disambiguation and are never wrong.Autovivification is the fifth line's real subject: push @{ $index{ai} }, "paper1" works even though $index{ai} did not exist. Perl saw it being used as an array reference and created one. This is enormously convenient for building nested indexes (push @{ $by_ip{$ip}{$day} }, $record; just works) and it is a trap when you read a structure you thought was there: merely checking if ($index{missing}{deep}) creates $index{missing} as an empty hash. Use exists for tests, and remember it when a data structure grows keys nobody added.
C++ (raw pointer) Perl (reference)
────────────────── ─────────────────
int arr[] = {1,2,3}; my @list = (1,2,3);
int *p = arr; my $aref = \@list;
p++; // legal, $aref++; // does NOT walk the
// now aliases # array -- it is not a memory
// arr[1] # address you can walk
delete p; // your job, and forgetting # freed automatically once the
// it or doing it twice # last reference to @list is
// corrupts the heap # gone -- no delete, ever
A Perl reference is not a memory address you can walk with arithmetic; $aref++ just increments a number that has lost its "this is a reference" tag, which is a bug, not a way to advance to the next element. In that sense a Perl reference behaves less like a raw C++ pointer and more like a std::shared_ptr: it is reference-counted, the thing it points to is freed automatically the moment the last reference to it disappears, and there is no delete to forget and no double-free to cause. The cost of that safety is the same cost shared_ptr pays: a cycle — two structures each holding a reference to the other — is never collected by reference counting alone, and Perl leaks it silently unless you reach for Scalar::Util::weaken on one side, exactly the way C++ code reaches for std::weak_ptr for the same reason.
Data::Dumper is how you see what you have built:
use Data::Dumper;
$Data::Dumper::Indent = 1; $Data::Dumper::Sortkeys = 1;
print Dumper($anon);
$VAR1 = [
{
'name' => 'fetch',
'options' => {
'from' => 'arxiv'
}
}
];
Setting Sortkeys makes output deterministic, which matters if you ever compare dumps. perldoc perldsc is the data-structures cookbook and is worth an hour early on.
sub summarize ($text, $max_words = 5, %opts) {
my @words = split ' ', $text;
my $out = join " ", @words[0 .. ($max_words - 1 < $#words ? $max_words - 1 : $#words)];
return $opts{upper} ? uc $out : $out;
}
say summarize("the quick brown fox jumps over the lazy dog");
say summarize("the quick brown fox jumps", 3, upper => 1);
sub minmax (@values) {
my @sorted = sort { $a <=> $b } @values;
return wantarray ? ($sorted[0], $sorted[-1]) : $sorted[-1];
}
my ($min, $max) = minmax(4, 9, 1);
my $just_max = minmax(4, 9, 1);
say "list context: $min..$max / scalar context: $just_max";
the quick brown fox jumps
THE QUICK BROWN
list context: 1..9 / scalar context: 9
use v5.36. Older code unpacks the argument array by hand: my ($text, $max) = @_;. You will read a lot of that, so know what @_ is: all arguments, flattened into one list.f(@a, @b) passes one combined list and the callee cannot tell where one ended. To pass two arrays separately, pass references: f(\@a, \@b). This is the single most common source of "my function got the wrong arguments".wantarray asks what context the caller used, so one sub can return a pair or a single value. Use it sparingly; it is clever, and clever is expensive to read.split ' ' with a literal single-space string is a special case meaning "split on runs of whitespace, ignoring leading whitespace", which is almost always what you want for text.our $depth = 0;
sub show { say " depth is $depth" }
sub descend {
local $depth = $depth + 1; # dynamic scope: visible to callees
show();
}
descend();
show();
depth is 1
depth is 0
| Keyword | Creates | Visible to |
|---|---|---|
my | a lexical variable | the enclosing block and any closure made inside it. Use this by default. |
our | an alias to a package global | everything, by full name too. For package-level constants and configuration. |
local | a temporary value for an existing global | the rest of this block and everything it calls, restored on exit. |
local is dynamic scoping, which almost no modern language has, and it is not a way to make local variables (that is my). Its real use is temporarily changing Perl's special variables safely:
{
local $/ = undef; # slurp mode: read the whole file at once
my $whole = <$fh>;
} # $/ restored automatically, even on die
{
local @ARGV = ("sample.log"); # make <> read this file
while (<>) { ... }
}
That pattern appears throughout this project, and it is the right tool because it cannot leak: the old value comes back when the block exits by any route, including an exception.
This is why you are here.
my $line = '192.168.1.42 - alice [10/Oct/2026:13:55:36 +0000] "GET /papers?id=7 HTTP/1.1" 200 2326';
if ($line =~ /^(\S+) \S+ (\S+) \[([^\]]+)\] "(\w+) (\S+)[^"]*" (\d{3}) (\d+)$/) {
say "ip=$1 user=$2 when=$3 method=$4 path=$5 status=$6 bytes=$7";
}
ip=192.168.1.42 user=alice when=10/Oct/2026:13:55:36 +0000 method=GET path=/papers?id=7 status=200 bytes=2326
It works, and you should never ship it. Seven numbered captures means every future edit renumbers everything after it, and nobody reading this in a year can tell what $5 was. The maintainable version uses named captures and /x:
my $apache = qr{
^(?<ip>\S+) \s+ \S+ \s+ (?<user>\S+) \s+ # client, identd, user
\[(?<ts>[^\]]+)\] \s+ # [timestamp]
"(?<method>[A-Z]+) \s (?<path>\S+) [^"]*" \s+ # "GET /path HTTP/1.1"
(?<status>\d{3}) \s+ (?<bytes>\d+) # status and size
}x;
if ($line =~ $apache) {
say "named: $+{ip} asked for $+{path} and got $+{status}";
say "capture names: ", join(",", sort keys %+);
}
named: 192.168.1.42 asked for /papers?id=7 and got 200
capture names: bytes,ip,method,path,status,ts,user
Four features doing the work:
/x makes whitespace and # comments inside the pattern insignificant, so a regex can be laid out and annotated like code. With /x you must write real spaces as \s or [ ], which is the small price. Any regex longer than about forty characters should use /x.(?<name>...) names a capture, available afterwards in %+. Renumbering stops being a problem, and the parsing code reads like the data.qr{...} compiles a pattern once into a value you can store, pass around, interpolate into a bigger pattern, and reuse in a loop. In a parser that runs a million times, compiling once matters; in readability terms it matters more, because it lets you build a library of named patterns.{...} as the delimiter instead of /.../ avoids escaping every slash in a path pattern. Perl lets you use almost any delimiter for m, s, qr and tr, and choosing one that does not appear in the pattern is basic hygiene.my $text = "contact alice\@example.com or bob\@test.org today";
my @emails = $text =~ /([\w.]+@[\w.]+)/g; # list context + /g: every match
say "all emails: @emails";
while ($text =~ /(\w+)@([\w.]+)/g) { # scalar context + /g: iterate
say " user=$1 host=$2 at offset $-[0]";
}
(my $masked = $text) =~ s/([\w.]+)@([\w.]+)/[redacted]\@$2/g;
say "masked: $masked";
my $prices = "cost: 10, 20, 30";
(my $doubled = $prices) =~ s/(\d+)/$1 * 2/ge; # /e evaluates the replacement
say "doubled: $doubled";
all emails: alice@example.com bob@test.org
user=alice host=example.com at offset 8
user=bob host=test.org at offset 29
masked: contact [redacted]@example.com or [redacted]@test.org today
doubled: cost: 20, 40, 60
/g in list context returns every capture from every match. One line to harvest every email in a document. If the pattern has no captures, you get the whole matches instead. /g in scalar context is an iterator, resuming from where it stopped (the position is stored with the string, readable as pos $text). That is the while loop above, and it is how you walk a large string without copying it.(my $copy = $original) =~ s/.../.../ is the copy-then-modify idiom. Substitution modifies in place, so without the copy you destroy your input. Perl 5.14 added the /r flag for the same thing more clearly: my $copy = $original =~ s/a/b/gr;. /e treats the replacement as Perl code to evaluate. s/(\d+)/$1 * 2/ge doubles every number in a string. This is a small superpower for data cleanup: normalising units, decoding escapes, reformatting dates, all inline.$-[0] and @+ hold the start and end offsets of the match, which you need when reporting exactly where in a file something went wrong.my $html = '<b>bold</b> and <i>italic</i>';
my ($greedy) = $html =~ /<(.+)>/;
my ($lazy) = $html =~ /<(.+?)>/;
say "greedy: $greedy";
say "lazy: $lazy";
say "count of tags: ", scalar(() = $html =~ /<[^>]+>/g);
greedy: b>bold</b> and <i>italic</i
lazy: b
count of tags: 4
.+ is greedy: it takes as much as it can and gives back only as needed, so it ran to the last > in the string. .+? is lazy and stops at the first. The third and best option is usually neither: [^>]+ says what you mean (characters that are not the terminator), is faster because it cannot backtrack, and does not depend on remembering which flavour of . you wanted.
scalar(() = $html =~ /.../g) is the countof idiom: assign the match list to an empty list in scalar context, which yields the number of elements. Ugly, universal, worth recognising.
Regex mistakes that cost the most time\b. /ERROR/ matches NOERROR and ERRORS=0./(\d{3})/ against a log line finds the first three digits anywhere, which may be part of the date. Anchor with ^, $, or surrounding context.. where a negated class belongs. "([^"]*)" beats "(.*?)" for a quoted field: clearer and it cannot backtrack catastrophically.(\s*\w+)*$ on a long non-matching line can take exponential time and hang your program. If a parser mysteriously stalls on one file, suspect this first.XML::LibXML; we will in Milestone 6.$1. On failure the capture variables keep their previous values, so a failed match silently reuses the last line's data. Always if ($line =~ ...) { ... }.open my $fh, "<", "sample.log" or die "cannot open sample.log: $!";
my $errors = 0;
while (my $line = <$fh>) {
chomp $line;
$errors++ if $line =~ /^ERROR\b/;
}
close $fh;
say "errors: $errors";
open my $out, ">", "summary.txt" or die "cannot write: $!";
say {$out} "errors=$errors";
close $out;
errors: 2
wrote: errors=2
open with a lexical filehandle is the only correct form. The mode is separate from the filename, so a file called >evil cannot become a redirection. Two-argument open with the mode inside the string is a genuine security hole, and you will see it in old code.or die "...: $!" — $! is the system error message ("No such file or directory"). An error message without $! tells you something failed but not why. while (my $line = <$fh>) reads one line at a time, so memory is constant whether the file is 2 KB or 200 GB. This is the whole reason the project can claim to stream, and it is the default rather than something you opt into.chomp removes the trailing newline (strictly, the current value of $/). Forgetting it means your "ip" ends with a newline and every comparison fails mysteriously.say {$out} "..." — the braces disambiguate a filehandle expression. With a simple handle, say $out "..." also works, with no comma, which looks wrong forever.Encoding deserves a line now and a milestone later. A file is bytes; treating it as text requires knowing the encoding:
open my $fh, "<:encoding(UTF-8)", $path or die "$path: $!";
# and for data that claims to be UTF-8 and is not, in milestone 7:
open my $fh, "<:raw", $path or die "$path: $!"; # bytes, decode manually
my $result = eval {
die "something broke\n";
1;
};
say "eval returned ", (defined $result ? $result : "undef"), " and \$\@ is: $@";
eval returned undef and $@ is: something broke
Perl's exception mechanism is die to throw and eval { } to catch. The block returns undef on failure and the error lands in $@. Two conventions: end your message with \n or Perl appends " at script.pl line 12", which is useful for bugs and noise for expected failures; and put 1; as the last statement of the eval block so success is unambiguous.
$@ is a global, and it is easy to clobber between the eval and the check (a destructor running, a cleanup call). The community solution is Try::Tiny:
use Try::Tiny;
try {
die { code => 503, message => "service unavailable" };
} catch {
my $err = $_;
say "Try::Tiny caught a ", ref($err), " with code $err->{code}";
};
Try::Tiny caught a HASH with code 503
Note that die can throw any reference, not just a string, which is how Perl does structured exceptions: a hash reference with a code and a message, or an object from a class like Throwable. For this project, structured errors matter because "line 40,112 of users.csv had unbalanced quotes" needs to be data, not prose.
package Strata::Record;
use v5.36;
sub new ($class, %args) {
my $self = {
source => $args{source} // "unknown",
fields => $args{fields} // {},
};
return bless $self, $class;
}
sub source ($self) { return $self->{source} }
sub field ($self, $name) { return $self->{fields}{$name} }
sub to_line ($self) {
my $f = $self->{fields};
return join " ", map { "$_=$f->{$_}" } sort keys %$f;
}
1; # a module must return a true value
apache: ip=10.0.0.1 status=200
ref: Strata::Record isa: yes
Perl's object system is three rules. A class is a package. An object is a reference that has been blessed into that package. A method call $obj->method(@args) calls the package's sub with the object as the first argument. That is all bless does: it writes the package name onto the reference so method lookup knows where to go.
It is minimal to the point of being spartan: no attribute declarations, no encapsulation (anyone can reach into $self->{fields}), no type checking. In production Perl most people use Moo or Moose, which add attributes, types, roles and defaults on top. And since 5.38 there is a real class syntax, still marked experimental:
use v5.38;
use experimental 'class';
class Strata::Entity {
field $type :param;
field $value :param;
field $count = 1;
method type { $type }
method seen { $count++; $self }
method to_string { "$type($value) x$count" }
}
my $e = Strata::Entity->new(type => "ip", value => "10.0.0.1");
$e->seen->seen;
say $e->to_string;
ip(10.0.0.1) x3
Genuinely pleasant, genuinely encapsulated (those fields are not reachable from outside), and genuinely experimental: the syntax may still change and it will warn unless you ask for it. This project uses plain bless, because it is what you will meet in existing code and because understanding it explains how Moo, Moose and the new class all work underneath.
Write a sub parse_apache_line($line) that returns a hash reference of named fields for a combined-format Apache line, or undef if the line does not match. Requirements: use qr// with /x and named captures; handle the - that appears for a missing user or a zero byte count by normalising it to undef and 0 respectively; and split the request field into method, path and protocol.
Then write a second sub summarise(@lines) returning a hash reference with total lines, parsed lines, failed lines, and a count by status code. Test it with three good lines and two deliberately broken ones.
use v5.36;
my $APACHE = qr{
^ (?<ip>\S+) \s+ (?<identd>\S+) \s+ (?<user>\S+) \s+
\[ (?<ts>[^\]]+) \] \s+
" (?<request>[^"]*) " \s+
(?<status>\d{3}) \s+ (?<bytes>\d+|-)
(?: \s+ " (?<referer>[^"]*) " \s+ " (?<agent>[^"]*) " )? # combined format
\s* $
}x;
sub parse_apache_line ($line) {
return undef unless $line =~ $APACHE;
my %f = %+; # copy: %+ is reset by the next match
# "-" is Apache's way of saying "nothing here".
$f{user} = undef if $f{user} eq "-";
$f{bytes} = 0 if $f{bytes} eq "-";
if ($f{request} =~ m{^(?<method>[A-Z]+) \s+ (?<path>\S+) (?: \s+ (?<proto>\S+))?$}x) {
@f{qw(method path proto)} = @+{qw(method path proto)};
} else {
$f{problem} = "unparsable request line";
}
return \%f;
}
sub summarise (@lines) {
my %out = (total => 0, parsed => 0, failed => 0, by_status => {});
for my $line (@lines) {
$out{total}++;
my $rec = parse_apache_line($line);
if ($rec) {
$out{parsed}++;
$out{by_status}{ $rec->{status} }++;
} else {
$out{failed}++;
}
}
return \%out;
}
Five things worth taking from this.
my %f = %+; copies the capture hash immediately. %+ is global and is reset by the next successful match anywhere, including the one two lines later that splits the request. Copying first is not optional; forgetting it produces a bug that appears only when you add a second regex.(?: ... )? parses both, and the fields are simply absent for the shorter one.- at the boundary means the rest of the program never thinks about Apache's conventions. Every parser in this project will do this: the Record that comes out should not betray which format it came from.problem field, not a discarded line. This is the project's central policy in miniature: keep the record, mark what is wrong with it, and let the caller decide. A line you throw away is a line you cannot investigate.@f{qw(method path proto)} = @+{qw(method path proto)} is a hash slice on both sides: three assignments in one statement. This is where Perl's sigil rules start paying you back.my %bytes_by_ip = ("10.0.0.1" => 4_112_883, "10.0.0.2" => 55_201, "10.0.0.9" => 913_004);
for my $ip (sort { $bytes_by_ip{$b} <=> $bytes_by_ip{$a} } keys %bytes_by_ip) {
printf " %-15s %10s\n", $ip, commify($bytes_by_ip{$ip});
}
sub commify ($n) {
1 while $n =~ s/^(\d+)(\d{3})/$1,$2/;
return $n;
}
10.0.0.1 4,112,883
10.0.0.9 913,004
10.0.0.2 55,201
printf with %-15s (left-aligned, 15 wide) and %10s (right-aligned) is how every Perl report is formatted. commify is a classic: 1 while s/.../.../ repeats a substitution until it stops matching, which is a loop written as an expression. Numeric literals can contain underscores for readability.
When sorting by an expensive computed key, use the Schwartzian transform, which computes each key once:
my @sorted = map { $_->[1] }
sort { $a->[0] <=> $b->[0] }
map { [ expensive_key($_), $_ ] } @records;
Read it bottom-up: decorate each record with its key, sort by the key, undecorate. It is the standard idiom precisely because a naive sort { expensive($a) <=> expensive($b) } calls the expensive function O(n log n) times instead of n.
Write a single program, bin/toptalkers, that reads Apache logs from files or standard input and prints a report. Requirements:
Test it on a file with a few thousand lines, including some you have deliberately truncated mid-line.
#!/usr/bin/env perl
use v5.36;
my $APACHE = qr{
^ (?<ip>\S+) \s+ \S+ \s+ (?<user>\S+) \s+
\[ (?<ts>[^\]]+) \] \s+ " (?<request>[^"]*) " \s+
(?<status>\d{3}) \s+ (?<bytes>\d+|-)
}x;
my (%hits, %bytes, %errors, %by_minute, %status_class);
my ($total, $failed, @first_failures) = (0, 0);
while (my $line = <>) {
$total++;
unless ($line =~ $APACHE) {
$failed++;
push @first_failures, "$ARGV:$." if @first_failures < 3;
next;
}
my %f = %+;
my $bytes = $f{bytes} eq "-" ? 0 : $f{bytes};
$hits{ $f{ip} }++;
$bytes{ $f{ip} } += $bytes;
$errors{ $f{ip} }++ if $f{status} >= 400;
$status_class{ substr($f{status}, 0, 1) . "xx" }++;
# 10/Oct/2026:13:55:36 +0000 -> 10/Oct/2026:13:55
$by_minute{$1}++ if $f{ts} =~ /^(\d+\/\w+\/\d+:\d+:\d+)/;
}
say "top talkers";
my @top = (sort { $hits{$b} <=> $hits{$a} || $a cmp $b } keys %hits)[0 .. 4];
for my $ip (grep { defined } @top) {
printf " %-15s %8d hits %12s bytes %5.1f%% errors\n",
$ip, $hits{$ip}, commify($bytes{$ip}),
100 * ($errors{$ip} // 0) / $hits{$ip};
}
say "\nstatus classes";
printf " %s %8d\n", $_, $status_class{$_} for sort keys %status_class;
say "\nbusiest minutes";
my @busy = (sort { $by_minute{$b} <=> $by_minute{$a} } keys %by_minute)[0 .. 4];
printf " %-22s %8d\n", $_, $by_minute{$_} for grep { defined } @busy;
my $rate = $total ? 100 * $failed / $total : 0;
printf "\n%d lines, %d unparsed (%.2f%%)%s\n", $total, $failed, $rate,
@first_failures ? " first at: " . join(", ", @first_failures) : "";
exit($rate > 1 ? 1 : 0);
sub commify ($n) { 1 while $n =~ s/^(\d+)(\d{3})/$1,$2/; return $n }
Six details that are the actual lesson.
$ARGV and $. give you free provenance. Inside <>, $ARGV is the file currently being read and $. is the line number, so "where did this come from" costs nothing. (One caveat worth knowing: $. does not reset between files unless you close ARGV at eof.)(sort ...)[0 .. 4] takes a slice of a list without an intermediate array, and grep { defined } handles the case where fewer than five exist. Slicing past the end gives undef, not an error, which is convenient and requires the guard.($errors{$ip} // 0) because an IP with no errors has no key, and arithmetic on undef warns. Defined-or is the right operator: a genuine zero must survive.If you built something close to this, Milestones 1 through 3 will feel like tidying rather than learning, which is the intention.
== on strings or eq on numbers. "10" == "10.0" is true; "10" eq "10.0" is false. Both are sometimes what you want.chomp, then wondering why $fields[-1] never matches anything.$1 without checking the match succeeded. It holds the previous match's value.%+ before the next match. Same class of bug, more surprising.exists.open, ever.my @lines = <$fh>) out of habit. It works until the file is 40 GB.$steps[0] and not @steps[0]?$count{$_}++ for @items do, step by step?local rather than my?/x, qr// and (?<name>...) each buy you?/g in list context and in scalar context?%+ immediately?open the only acceptable form?bless actually do?Look back at the capstone solution. The regex is a first-class value with named fields laid out over five commented lines; the input loop handles files and pipes with no imports; five aggregations happen in one pass with hashes that spring into existence as needed; provenance comes free in $ARGV and $.; and the whole thing is sixty lines and streams a file of any size. The equivalent Python is perhaps twice as long and has re.compile, fileinput, defaultdict and argparse visible in it. That difference is Perl's argument, and for this kind of work it is a strong one.
The costs are equally visible. Context means my $x = f() and my ($x) = f() are different programs. Capture variables are global and get clobbered. local is a scoping rule most programmers have never met. Sigils change with what you are asking for rather than what the variable is. None of these is hard once learned, and all of them are sharp edges that a language designed in 2010 would not have.
The discipline that makes Perl maintainable is not subtle, and it is all in this part: use v5.36, named captures, /x on anything long, real data structures instead of clever parallel arrays, and tests from day one.
Milestone 1 turns the one-liner instinct into a program: a filter with proper option handling, exit codes and tests. Milestone 2 adds the aggregation you just wrote by hand. Milestone 3 is where the parser becomes serious, with a pattern library, failure accounting, and the first fixtures of deliberately broken input.
Before then: run the one-liners from Part 1 against a log file on your own machine, and read perldoc perlretut. It is the best forty minutes available to you at this point.