Milestones 1–4

Instalment 11 · Course 3 (Perl) · Parts 0–2

The language that assumes your data is a mess

Go was about concurrency and Ruby about expressiveness. Perl is about the hour you spend with ten gigabytes of logs that nobody documented, half of which are truncated, and a question you need answered by lunchtime.

Verification note

Every example below was run on Perl 5.38.2 and the outputs are copied from those runs. The modules used in the project (Text::CSV, DBI, DBD::SQLite, XML::LibXML, Try::Tiny) are installed in my sandbox, so the later milestones will be verified too.

Course 3 · Part 0What are we building?

The final result

The Mewlang cat, looking up curiouslyA command-line tool called strata that you point at a directory of unexplained files and interrogate.

$ strata ingest ./incident-2026-09-12/ --recursive
  apache/access.log        412,884 lines   apache_combined     3.2s
  apache/error.log          18,221 lines   apache_error        0.4s
  app/service.log          904,110 lines   json_lines          8.1s
  exports/users.csv          9,412 rows    csv                 0.3s
  config/nginx.conf            420 lines   nginx_config        0.0s
  unknown/dump.txt          33,900 lines   unstructured        1.1s
  corrupt/partial.log        2,004 lines   apache_combined     0.1s  (91 malformed, kept)

  1,381,051 records, 214,882 entities, 46,203 events in 13.2s (104k lines/sec)

$ strata entities --type ip --top 5
  10.14.22.9        88,214 occurrences   6 files   first 13:02:11  last 14:47:52
  10.14.22.31       41,002 occurrences   4 files   ...

$ strata timeline --entity ip:10.14.22.9 --around '13:44:10' --window 90s
  13:43:58  apache/access.log:88214   GET /api/export  200  1.2MB
  13:44:02  app/service.log:551203    export.start  user=4412 rows=900000
  13:44:09  apache/error.log:9902     upstream timed out
  13:44:10  app/service.log:551288    ERROR OOM killed worker pid=8823
  13:44:11  apache/access.log:88240   GET /api/export  502

$ strata graph --entity user:4412 --depth 2 --format dot | dot -Tsvg > incident.svg

Underneath: a streaming ingestion pipeline with pluggable format detectors, an entity extractor, a normaliser that turns eleven timestamp formats into one, a correlation engine that groups records into events, and a SQLite-backed store you can query, all of it able to survive files that are truncated, mis-encoded, or simply lying about their format.

Why this project is interesting

Most data-processing tutorials use clean data, which is the one thing you will never be given. Real forensic work looks like this: eight formats, three of them undocumented, timestamps in four time zones, a log rotated mid-write so one line is half of two lines, an "XML" file that is actually XML fragments concatenated without a root element, and a CSV whose quoting breaks on row 40,000 because someone's surname contains a comma and a quotation mark.

The interesting engineering is not parsing any one format. It is building something that keeps going: that treats malformed input as expected rather than exceptional, that tells you exactly what it could not parse and why, and that never loads more than one line into memory so the same tool works on a 2 KB file and a 200 GB one.

Why Perl in particular

Why are we using this language here?

The Mewlang cat, wearing glasses and looking confidentFor the middle of this project (streaming, regex-heavy, line-oriented transformation with Unix plumbing), Perl is still the best tool in existence, and the reason is not nostalgia: no other language has put regular expressions, context, and the input loop into the syntax itself.

Where it is not the answer, and we will say so at the time:

  • Structured analysis. Once the data is clean and rectangular, Python with pandas or Polars is better. Grouping, joining and statistics are libraries there and hand-written loops here.
  • Raw throughput. Go or Rust would be several times faster on the same regexes and would parallelise more easily. We will measure this rather than assert it.
  • Long-lived services. Perl is a fine language for programs that start, work and exit. For something that runs for six months, the ecosystem's tooling for observability and deployment is thinner than Go's.
  • Team maintenance. Perl earns its reputation for write-only code when written badly, and badly is easy. A large part of this course is the discipline that prevents it: use v5.36, named captures, /x, real data structures, and tests.

The honest summary: Perl is the best language I know for the first two hours with unknown data, and you should be prepared to hand the result to something else afterwards. Knowing when to hand over is part of what this course teaches.

Architecture we are building toward

   files (any shape, any size, some broken)
        │
        ▼
  ┌───────────────┐   one line at a time, constant memory
  │  Source       │   handles gzip, encodings, truncation
  └───────┬───────┘
          ▼
  ┌───────────────┐   sniffs the format, or is told
  │  Detector     │   apache | nginx | json_lines | csv | xml | conf | code | unknown
  └───────┬───────┘
          ▼
  ┌───────────────┐   a plugin per format, each returning a Record
  │  Parsers      │   malformed lines become Records with a `problem` field,
  └───────┬───────┘   never exceptions and never silent drops
          ▼
  ┌───────────────┐   ips, emails, paths, ids, versions, timestamps...
  │  Extractors   │   each is a named regex plus a normaliser
  └───────┬───────┘
          ▼
  ┌───────────────┐   one canonical timestamp, one canonical host,
  │  Normaliser   │   one canonical path, whatever the source said
  └───────┬───────┘
          ▼
  ┌───────────────┐   records near in time and sharing an entity
  │  Correlator   │   become events; events referencing each other
  └───────┬───────┘   become a graph
          ▼
  ┌───────────────┐
  │  Store        │   SQLite: records, entities, occurrences, events, edges
  └───────┬───────┘
          ▼
   query · timeline · graph · report        (and stdout, for piping)

The twelve milestones

#MilestoneWhat it teaches
1A filter that counts what it readsscalars, while (<>), $_, strict and warnings, exit codes
2Summaries: top talkers, error rateshashes, sorting, context, output formatting
3An Apache/nginx parserregexes, named captures, /x, qr//, failure counting
4A Record model and a pipelinereferences, nested data, subs, Data::Dumper
5Splitting into modulespackages, lib/, cpanm, cpanfile, Test::More, prove
6Pluggable formats: CSV, JSON, XMLdispatch tables, CPAN modules, format sniffing
7Streaming 10 GB safelyencodings, gzip, truncation, memory discipline, malformed-input policy
8Entity extraction and normalisationIPs, timestamps, paths, ids; canonical forms; time zones
9Correlation and sessionisationtime windows, joins, SQLite via DBI, indexes
10A proper command-line toolGetopt::Long, exit codes, signals, pipes, --help
11Testing, fuzzing and profilingTest::More, corrupt-input tests, Devel::NYTProf, optimisation
12The knowledge graphgraph storage and queries, fork for parallelism, packaging

What you will know afterwards

How to build a streaming pipeline that survives input designed to break it; how to write regexes that a colleague can read a year later; why Perl's context rule exists and how to stop it surprising you; how to correlate events across sources with nothing but a time window and a shared identifier; and when to stop and hand the cleaned data to something else.


Course 3 · Part 1Install and first program

What Perl is

Perl is a dynamically typed, garbage-collected interpreted language from 1987, designed explicitly for report processing and system glue. It is not "Perl 6": that language was renamed Raku in 2019 and is a different thing. Perl 5 is alive, gets an annual release, and version 5.36 (2022) and 5.38 (2023) modernised it substantially, with subroutine signatures and a real class syntax.

Two facts shape everything about the language. First, values have no fixed type but variables have a fixed shape, marked by a sigil: $ for a single value, @ for a list, % for a key-value table. Second, every expression is evaluated in a context, either "give me one value" or "give me a list", and many operations do different things depending on which. Nothing else in mainstream programming works this way, and both are covered properly in Part 2.

Installing

Unlike Ruby, using the system Perl is usually fine: it is well-maintained, present everywhere, and most distributions keep it current enough. Perl's compatibility record is unusually good, so a script from 2010 generally still runs. You want 5.36 or newer to get signatures and the modern feature bundle; I verified everything on 5.38.2.

Linux
perl -v                                 # almost certainly already there
sudo apt install perl cpanminus         # Debian/Ubuntu: adds the cpanm installer
sudo dnf install perl perl-App-cpanminus # Fedora

# many CPAN modules exist as distribution packages, which is the easiest path:
sudo apt install libtext-csv-perl libdbd-sqlite3-perl libxml-libxml-perl libtry-tiny-perl
macOS
perl -v                    # Apple ships one, often a version or two behind
brew install perl          # newer, installed under /opt/homebrew
Windows
:: Strawberry Perl bundles a C compiler, so modules with XS code build
winget install StrawberryPerl.StrawberryPerl

Strawberry Perl is genuinely good. That said, this course leans on Unix habits (pipes, STDIN, signals, file permissions) more than any other in the curriculum, and WSL2 will save you real friction.

When you need your own Perl

If you need a version the system does not have, or you are installing modules on a machine where you have no root, use a version manager:

curl -L https://install.perlbrew.pl | bash     # perlbrew
perlbrew install perl-5.38.2
perlbrew switch perl-5.38.2

# or plenv, if you liked rbenv

Editor and tooling

Perl's tooling is not as unified as a single-vendor language's, but two language-server projects cover most editors well enough to set up before Part 2, plus a static analyser and a formatter worth knowing about even before you wire either into anything.

VS Code
# either one, from the Extensions panel or the command line:
code --install-extension bscan.perlnavigator     # PerlNavigator: lighter, installs in seconds
code --install-extension richterger.perl         # Perl::LanguageServer: deeper, needs a C toolchain

PerlNavigator is a TypeScript language server that shells out to perl -c and perlcritic for you, giving syntax checking, outlining and completion with almost no setup. Perl::LanguageServer goes further — real go-to-definition, variable inspection, a working debugger — because its backend is itself a Perl module, an XS one, which means it needs cpan Perl::LanguageServer to succeed on your machine before the extension works at all, and an XS build occasionally means fighting a missing compiler or a stale header before it does. Start with PerlNavigator; reach for Perl::LanguageServer only once you miss real navigation.

Vim, Neovim, Emacs

Neovim's built-in LSP client (via nvim-lspconfig) can drive either server above directly, the same way it would for any other language server. Plain Vim's long-standing option is perl-support.vim, for syntax, skeletons and running the current file, without full language-server intelligence. Emacs ships cperl-mode, the conventional choice over the older bundled perl-mode: it indents heredocs and POD correctly, which the older mode does not, and eglot or lsp-mode can point it at either server above for full completion.

perlcritic and perltidy: the closest thing Perl has to a linter and a formatter
cpanm Perl::Critic Perl::Tidy
perlcritic bin/strata                # static analysis, default severity
perlcritic --brutal bin/strata       # same file, strictest severity level
perltidy -b lib/Strata/Record.pm     # reformat in place; -b keeps a .bak

perlcritic checks code against configurable rules drawn largely from Damian Conway's Perl Best Practices, at five severity levels from --gentle (only the most serious issues) to --brutal (everything); a project settles on a level and a .perlcriticrc once, rather than treating every possible warning as mandatory from day one. perltidy reformats to a configurable style the way gofmt does for Go, except it is not bundled with the interpreter and not run on every commit by default the way gofmt effectively is — both are a five-minute cpanm install away, and neither is the default a newer language's tooling would make it.

CPAN, cpanm, and not using sudo

CPAN is Perl's module archive, and at around 220,000 modules it is one of the oldest and deepest package ecosystems in existence. The modern client is cpanm:

cpanm Text::CSV                  # install one module
cpanm --installdeps .            # install everything a cpanfile lists
cpanm --local-lib=~/perl5 DBI    # install into your home directory, no root

# then tell perl where to find them (add to ~/.bashrc):
eval "$(perl -I ~/perl5/lib/perl5 -Mlocal::lib)"

Never use sudo cpanm. It mixes modules you installed with modules your operating system manages, and the resulting breakage is tedious. Use local::lib, perlbrew, or your distribution's packages.

For a project, declare dependencies in a cpanfile:

# cpanfile
requires 'perl', '5.036';
requires 'Text::CSV', '2.00';
requires 'DBI';
requires 'DBD::SQLite';
requires 'Try::Tiny';

on 'test' => sub {
    requires 'Test::More', '1.302';
};

Then cpanm --installdeps . installs them. If you want a lockfile and a bundled dependency tree (Bundler's role), that is Carton: carton install writes cpanfile.snapshot and carton exec runs with exactly those versions.

Why are we using this language here?

CPAN is the strongest argument for doing forensic text work in Perl at all: at roughly 220,000 modules covering three decades of "someone already hit this exact format and fixed it", the odds that today's undocumented log shape already has a parser somewhere are good. That is not spin — Text::CSV and XML::LibXML, which this project leans on from Milestone 6 onward, are mature, heavily-used modules, not toy packages someone abandoned in 2009.

Two honest costs sit right next to that strength. First, a real fraction of CPAN's best modules — XML::LibXML among them — are XS: Perl wrapped around C, which means installing one sometimes means having a working compiler and system headers rather than just downloading a file, and it is the single most common reason a cpanm install fails on a machine that "should" work. Second, cpanm --installdeps . alone gives you no lockfile: two checkouts on two machines can silently resolve a cpanfile's version ranges to two different sets of module versions unless you additionally reach for Carton, which is opt-in rather than the default the way Gemfile.lock or package-lock.json are. And perlbrew, while it works, compiles an entire Perl interpreter from source for every version you install — it is closer in spirit to pyenv or rbenv than to a version manager that downloads a prebuilt binary, so "just try 5.38" costs several minutes the first time, not seconds. None of this is disqualifying; all of it is worth knowing before the first cpanm command mysteriously fails on someone else's laptop.

Documentation: the best thing about Perl

perldoc -f split          # one built-in function, with examples
perldoc perlre            # the regex reference
perldoc perlretut         # the regex tutorial
perldoc perlrequick       # the regex quick start
perldoc perlvar           # what $_, @ARGV, $! and the rest mean
perldoc perldsc           # data structures cookbook: the one to read twice
perldoc perlop            # operators and precedence
perldoc -q "how do I sort"  # search the FAQ
perldoc Text::CSV         # any installed module's own documentation
perldoc -l Text::CSV      # where that module actually lives on disk

Perl's documentation ships with the interpreter, is written by people who use the language, and is complete. perldoc perlretut alone is a better regex tutorial than most books. Get into the habit of reading it in a terminal rather than searching the web, because the web has thirty years of outdated Perl advice on it and perldoc has none.

Project layout

strata/
├── cpanfile              dependencies
├── bin/
│   └── strata            the command-line program (no .pl extension)
├── lib/
│   └── Strata/
│       ├── Record.pm     Strata::Record
│       ├── Source.pm     Strata::Source
│       └── Parser/
│           ├── Apache.pm Strata::Parser::Apache
│           └── CSV.pm
├── t/
│   ├── 00-load.t         does everything compile?
│   ├── 10-record.t
│   └── 20-parser-apache.t
└── share/
    └── fixtures/         sample logs, including deliberately broken ones

Three conventions, all enforced by tooling rather than taste:

Hello, archaeology

#!/usr/bin/env perl
use v5.36;

my $name = shift @ARGV // "world";
say "hello, $name";
say "perl $^V, script $0, pid $$";
$ perl hello.pl
hello, world
perl v5.38.2, script hello.pl, pid 18244

$ perl hello.pl strata
hello, strata

Every line, explained

#!/usr/bin/env perl — the shebang, so ./hello.pl works once the file is executable. /usr/bin/env perl finds whichever Perl is first on your PATH, which is what you want with perlbrew.

use v5.36; — the single most important line in modern Perl. It enables a whole feature bundle at once:

What it turns onWhy it matters
use strictUndeclared variables are a compile error, not a new global
use warningsUndefined values, dubious conversions and typos in comparisons get reported
sayprint with a newline
subroutine signaturessub f ($x, $y = 1) { ... } instead of unpacking @_ by hand
isa, postfix deref, and moremodern conveniences without per-feature imports
disables indirect object syntaxremoves a genuinely confusing old parsing rule

Old Perl tutorials start with use strict; use warnings; as two separate lines. use v5.36; includes both and more, and it also declares your minimum Perl version, so running on something older fails immediately with a clear message rather than mysteriously.

my $name = shift @ARGV // "world"; — my declares a lexical variable, scoped to the enclosing block. @ARGV holds the command-line arguments (unlike C, it does not include the program name, which is in $0). shift removes and returns the first element. // is the defined-or operator: use the left side unless it is undef. It differs from || in that 0 and "" are kept, which matters constantly when parsing data where zero is a real value.

say "hello, $name" — double-quoted strings interpolate variables, including array and hash elements. Single quotes do not interpolate anything.

$^V, $0, $$ — Perl's punctuation variables: the version, the program name, the process id. There are dozens; perldoc perlvar lists them all, and Part 2 covers the six you actually need.

One-liners, which are not a gimmick

Before writing a script, Perl programmers interrogate data from the shell. This is a real part of the language's culture and the fastest way to build intuition:

# print lines matching a pattern (grep, but with Perl regexes)
perl -ne 'print if /ERROR/' access.log

# -l adds newline handling, -a splits each line into @F on whitespace
perl -lane 'print $F[0] if $F[8] == 500' access.log

# count by field: the Perl idiom you will use a thousand times
perl -lane '$c{$F[0]}++; END { print "$c{$_}\t$_" for sort { $c{$b} <=> $c{$a} } keys %c }' access.log

# in-place edit with a backup
perl -i.bak -pe 's/\bDEBUG\b/TRACE/g' service.log

# -F sets the split pattern: CSV-ish, badly, but instantly
perl -F, -lane 'print $F[2] if $F[4] > 100' export.csv
FlagMeaning
-ethe program is on the command line
-nwrap it in while (<>) { ... }
-psame, and print $_ at the end of each iteration
-lstrip the newline on input, add one on output
-aautosplit each line into @F
-Fthe pattern to autosplit on
-iedit files in place

That third one-liner is a complete top-talkers report, and it is the seed of Milestone 2. Perl's design makes the exploratory version and the production version the same language, which is exactly what you want when the exploration turns out to be the tool.

Testing, from the first day

use v5.36;
use lib "lib";
use Test::More tests => 4;
use Strata::Record;

my $r = Strata::Record->new(source => "apache", fields => { ip => "10.0.0.1" });
isa_ok $r, "Strata::Record";
is $r->source, "apache", "source is kept";
is $r->field("ip"), "10.0.0.1", "fields are readable";
like $r->to_line, qr/ip=10\.0\.0\.1/, "to_line renders fields";
$ perl -Ilib t/10-record.t
1..4
ok 1 - An object of class 'Strata::Record' isa 'Strata::Record'
ok 2 - source is kept
ok 3 - fields are readable
ok 4 - to_line renders fields

$ prove -l t/
t/10-record.t .. ok
All tests successful.
Files=1, Tests=4,  0 wallclock secs
Result: PASS

That output format is TAP, the Test Anything Protocol, which Perl invented in 1987 and which now has implementations in most languages. A test file is an ordinary program that prints ok and not ok lines; prove runs many of them and summarises. prove -l adds lib/ to the path, prove -lv shows every assertion, and prove -lj4 runs four files in parallel.

Exercise 1

Write bin/logstat, a program that reads log lines from files named on the command line (or from standard input when none are named) and prints: the total line count, how many lines contain ERROR, and the percentage. It must exit 0 when there are no errors and 1 when there are, so it can be used in a shell conditional.

Then prove it works in a pipeline: cat *.log | ./bin/logstat and ./bin/logstat a.log b.log should both work without changing the code.

Hints: while (<>) handles both cases for free; $. is the current line number; exit sets the status; printf formats the percentage.

Solution 1 — open after trying
#!/usr/bin/env perl
use v5.36;

my ($lines, $errors) = (0, 0);

while (my $line = <>) {
    $lines++;
    $errors++ if $line =~ /\bERROR\b/;
}

if ($lines == 0) {
    warn "logstat: no input\n";
    exit 2;
}

printf "%d lines, %d errors (%.2f%%)\n", $lines, $errors, 100 * $errors / $lines;
exit($errors > 0 ? 1 : 0);
$ printf 'ERROR disk full\nINFO started\nERROR timeout\nWARN slow\n' | ./bin/logstat
4 lines, 2 errors (50.00%)
$ echo $?
1

Four things worth extracting.

  • while (my $line = <>) is the entire input story. No arguments means standard input; arguments mean read each of those files in turn. Your program is a Unix filter with no extra work, which is why cat x | prog and prog x both work.
  • \b word boundaries matter. Without them, a line containing NOERROR or ERRORS_TOTAL=0 counts as an error. Ninety per cent of wrong log analysis is a missing word boundary.
  • Three exit codes, deliberately: 0 nothing found, 1 something found, 2 could not run. That is the grep convention, and following it means your tool composes with shell scripts people already know.
  • warn writes to standard error, print and say to standard output. Keeping diagnostics off stdout is what lets someone pipe your output into another program.

One subtlety: while (<>) without assigning to a variable puts the line in $_, which is idiomatic and slightly risky, because anything you call inside the loop might also use $_. Assigning to a named variable, as here, is the safer habit in anything longer than a one-liner.

The Mewlang cat, giving an unimpressed side-eyeCommon first-day errors
  • Can't locate Strata/Record.pm in @INC — lib/ is not on the search path. Use perl -Ilib, use lib 'lib';, or prove -l.
  • Strata/Record.pm did not return a true value — you forgot the 1; at the end of the module.
  • Global symbol "$count" requires explicit package name — use strict is doing its job: you forgot my.
  • Use of uninitialized value in ... — a warning, not an error, and almost always a real bug: you used a value that was never set, usually a capture group from a match that failed.
  • Can't call method "new" on an undefined value — you forgot to use the module that defines the class.
  • syntax error at ... near "}" — a missing semicolon on the previous line. Perl's error points at where parsing failed, not where you went wrong.
  • Using == to compare strings. == is numeric; eq is for strings. "abc" == "def" is true, because both convert to 0.

Checkpoint

  1. What does use v5.36; switch on, and why is it better than use strict; use warnings;?
  2. Why must a module end with 1;?
  3. What is the difference between // and ||, and when does it matter?
  4. What does while (<>) read from?
  5. Why should you never run sudo cpanm?
  6. Which three perldoc pages would you open first for a regex question?

Course 3 · Part 2Language crash course

Only what the project needs, which is most of Perl's text machinery and none of its formats or tie magic. Every output below is from a real run. Keep a terminal open and type them.

2.1 Sigils and context, the rule with no equivalent elsewhere

use v5.36;

my $count   = 42;
my @steps   = ("fetch", "filter", "save");
my %options = (topic => "AI", limit => 10);

say "scalar: $count";
say "array has ", scalar(@steps), " elements, last index $#steps";
say "element: $steps[0], slice: @steps[0,1]";
say "hash value: $options{topic}, keys: ", join(",", sort keys %options);

my $n = @steps;              # array in scalar context: its length
my ($first) = @steps;        # list context: its first element
say "scalar context gives $n, list context gives $first";
scalar: 42
array has 3 elements, last index 2
element: fetch, slice: fetch filter
hash value: AI, keys: limit,topic
scalar context gives 3, list context gives fetch

Two rules explain that output, and together they are the thing that makes Perl feel alien for a week and obvious afterwards.

The sigil describes what you are asking for, not what the variable is. @steps is the whole array; $steps[0] is one scalar from it, so it takes $; @steps[0,1] is several, so it takes @. Likewise %options is the hash and $options{topic} is one value from it. The rule is consistent, and it is the opposite of what people assume ("@ means array"), which is why $steps[0] looks wrong to newcomers.

Every expression is evaluated in scalar or list context, and many behave differently in each. my $n = @steps asks for one value from an array, so you get its length; my ($first) = @steps has a list on the left, so the array is unpacked and the first element assigned. Those two lines differ only in parentheses and mean completely different things.

my @words = split /,/, "a,b,c";
my $words = split /,/, "a,b,c";     # scalar context: count
say "list: @words / scalar: $words";
say "reverse in list: ", join("", reverse @words);
say "reverse in scalar: ", scalar reverse "hello";
list: a b c / scalar: 3
reverse in list: cba
reverse in scalar: olleh

reverse reverses a list in list context and a string in scalar context. This is not a special case; it is the design. When a Perl function surprises you, the first question is always "what context is this in?", and perldoc -f reverse will tell you what it does in each.

Typical language vs Perl

In Python, len(xs) is a function call and there is no way for an expression to know whether its caller wants one value or many. Perl passes that information down, which makes code shorter (my ($x) = f() versus x = f()[0]) and makes a category of bug possible that exists nowhere else (my $x = f() quietly giving you a count instead of a value). The mitigation is the same as the diagnosis: when a value is wrong in a confusing way, print scalar(@thing) and check which context you are in.

2.2 Arrays

my @nums = (5, 3, 9, 1);
push @nums, 7;
my $popped = pop @nums;
say "after push/pop: @nums (popped $popped)";
say "sorted numerically: ", join(",", sort { $a <=> $b } @nums);
say "sorted as strings:  ", join(",", sort @nums);
say "grep > 3: ", join(",", grep { $_ > 3 } @nums);
say "map doubled: ", join(",", map { $_ * 2 } @nums);
my @spliced = splice(@nums, 1, 2);
say "spliced out @spliced leaving @nums";
after push/pop: 5 3 9 1 (popped 7)
sorted numerically: 1,3,5,9
sorted as strings:  1,3,5,9
grep > 3: 5,9
map doubled: 10,6,18,2
spliced out 3 9 leaving 5 1

2.3 Hashes, and the counting idiom

my %count;
$count{$_}++ for qw(fetch filter fetch save fetch);

for my $k (sort { $count{$b} <=> $count{$a} || $a cmp $b } keys %count) {
    say "  $k: $count{$k}";
}

say "exists: ", (exists $count{fetch} ? "yes" : "no");
delete $count{fetch};
say "after delete: ", join(",", sort keys %count);
my @wanted = @count{qw(filter save)};      # hash slice
say "slice: @wanted";
  fetch: 3
  filter: 1
  save: 1
exists: yes
after delete: filter,save
slice: 1 1

$count{$_}++ for @things is the most-used line in Perl, and it is worth unpacking completely. %count starts empty. $count{$_} on a missing key is undef, and ++ on undef treats it as 0 and makes it 1, with no warning because incrementing undef is explicitly allowed. The for at the end is a statement modifier: a postfix loop, readable precisely because it is short.

The sort block chains two comparisons with ||: compare counts descending ($count{$b} <=> $count{$a}), and when they are equal (<=> returns 0, which is false) fall through to comparing keys alphabetically. That is the standard multi-key sort and it appears in every report you will ever write.

2.4 References and real data structures

Arrays and hashes can only hold scalars, so nesting requires references: scalars that point at something.

my @list = (1, 2, 3);
my %opts = (a => 1);
my $aref = \@list;              # reference to an existing array
my $href = \%opts;
my $anon = [ { name => "fetch", options => { from => "arxiv" } } ];   # anonymous

say "deref whole: @$aref / @{$aref}";
say "element: $aref->[0] and $$aref[0]";
say "nested: $anon->[0]{name} -> $anon->[0]{options}{from}";
say "ref types: ", join(",", map { ref } ($aref, $href, $anon, sub {}, \"x"));
push @$aref, 4;
say "the original array sees it: @list";

my %index;
push @{ $index{ai} }, "paper1";      # autovivification
push @{ $index{ai} }, "paper2";
say "autovivified: ", join(",", @{ $index{ai} });
deref whole: 1 2 3 / 1 2 3
element: 1 and 1
nested: fetch -> arxiv
ref types: ARRAY,HASH,ARRAY,CODE,SCALAR
the original array sees it: 1 2 3 4
autovivified: paper1,paper2

Four rules cover almost everything:

  1. \ takes a reference; [ ... ] and { ... } create anonymous arrays and hashes directly. Use the anonymous forms in data structures.
  2. -> both dereferences and indexes. $aref->[0], $href->{key}. The older $$aref[0] means the same thing and is harder to read.
  3. The arrow is optional between subscripts. $anon->[0]{name}{x} is the same as $anon->[0]->{name}->{x}. Write the first arrow, omit the rest; everyone does.
  4. To use a reference as a whole aggregate, put the right sigil in front: @$aref, %$href, @{ $index{ai} }. The braces are for disambiguation and are never wrong.

Autovivification is the fifth line's real subject: push @{ $index{ai} }, "paper1" works even though $index{ai} did not exist. Perl saw it being used as an array reference and created one. This is enormously convenient for building nested indexes (push @{ $by_ip{$ip}{$day} }, $record; just works) and it is a trap when you read a structure you thought was there: merely checking if ($index{missing}{deep}) creates $index{missing} as an empty hash. Use exists for tests, and remember it when a data structure grows keys nobody added.

References vs C++ pointers
C++ (raw pointer)                         Perl (reference)
──────────────────                         ─────────────────
int arr[] = {1,2,3};                      my @list = (1,2,3);
int *p = arr;                             my $aref = \@list;
p++;                     // legal,        $aref++;   // does NOT walk the
                         // now aliases   # array -- it is not a memory
                         // arr[1]        # address you can walk
delete p;   // your job, and forgetting   # freed automatically once the
            // it or doing it twice       # last reference to @list is
            // corrupts the heap          # gone -- no delete, ever

A Perl reference is not a memory address you can walk with arithmetic; $aref++ just increments a number that has lost its "this is a reference" tag, which is a bug, not a way to advance to the next element. In that sense a Perl reference behaves less like a raw C++ pointer and more like a std::shared_ptr: it is reference-counted, the thing it points to is freed automatically the moment the last reference to it disappears, and there is no delete to forget and no double-free to cause. The cost of that safety is the same cost shared_ptr pays: a cycle — two structures each holding a reference to the other — is never collected by reference counting alone, and Perl leaks it silently unless you reach for Scalar::Util::weaken on one side, exactly the way C++ code reaches for std::weak_ptr for the same reason.

Data::Dumper is how you see what you have built:

use Data::Dumper;
$Data::Dumper::Indent = 1; $Data::Dumper::Sortkeys = 1;
print Dumper($anon);
$VAR1 = [
  {
    'name' => 'fetch',
    'options' => {
      'from' => 'arxiv'
    }
  }
];

Setting Sortkeys makes output deterministic, which matters if you ever compare dumps. perldoc perldsc is the data-structures cookbook and is worth an hour early on.

2.5 Subroutines

sub summarize ($text, $max_words = 5, %opts) {
    my @words = split ' ', $text;
    my $out = join " ", @words[0 .. ($max_words - 1 < $#words ? $max_words - 1 : $#words)];
    return $opts{upper} ? uc $out : $out;
}

say summarize("the quick brown fox jumps over the lazy dog");
say summarize("the quick brown fox jumps", 3, upper => 1);

sub minmax (@values) {
    my @sorted = sort { $a <=> $b } @values;
    return wantarray ? ($sorted[0], $sorted[-1]) : $sorted[-1];
}

my ($min, $max) = minmax(4, 9, 1);
my $just_max    = minmax(4, 9, 1);
say "list context: $min..$max / scalar context: $just_max";
the quick brown fox jumps
THE QUICK BROWN
list context: 1..9 / scalar context: 9

2.6 my, our, and local

our $depth = 0;
sub show { say "  depth is $depth" }

sub descend {
    local $depth = $depth + 1;     # dynamic scope: visible to callees
    show();
}

descend();
show();
  depth is 1
  depth is 0
KeywordCreatesVisible to
mya lexical variablethe enclosing block and any closure made inside it. Use this by default.
ouran alias to a package globaleverything, by full name too. For package-level constants and configuration.
locala temporary value for an existing globalthe rest of this block and everything it calls, restored on exit.

local is dynamic scoping, which almost no modern language has, and it is not a way to make local variables (that is my). Its real use is temporarily changing Perl's special variables safely:

{
    local $/ = undef;           # slurp mode: read the whole file at once
    my $whole = <$fh>;
}                               # $/ restored automatically, even on die

{
    local @ARGV = ("sample.log");   # make <> read this file
    while (<>) { ... }
}

That pattern appears throughout this project, and it is the right tool because it cannot leak: the old value comes back when the block exits by any route, including an exception.

2.7 Regular expressions, part one: matching and capturing

This is why you are here.

my $line = '192.168.1.42 - alice [10/Oct/2026:13:55:36 +0000] "GET /papers?id=7 HTTP/1.1" 200 2326';

if ($line =~ /^(\S+) \S+ (\S+) \[([^\]]+)\] "(\w+) (\S+)[^"]*" (\d{3}) (\d+)$/) {
    say "ip=$1 user=$2 when=$3 method=$4 path=$5 status=$6 bytes=$7";
}
ip=192.168.1.42 user=alice when=10/Oct/2026:13:55:36 +0000 method=GET path=/papers?id=7 status=200 bytes=2326

It works, and you should never ship it. Seven numbered captures means every future edit renumbers everything after it, and nobody reading this in a year can tell what $5 was. The maintainable version uses named captures and /x:

my $apache = qr{
    ^(?<ip>\S+) \s+ \S+ \s+ (?<user>\S+) \s+      # client, identd, user
    \[(?<ts>[^\]]+)\] \s+                          # [timestamp]
    "(?<method>[A-Z]+) \s (?<path>\S+) [^"]*" \s+  # "GET /path HTTP/1.1"
    (?<status>\d{3}) \s+ (?<bytes>\d+)             # status and size
}x;

if ($line =~ $apache) {
    say "named: $+{ip} asked for $+{path} and got $+{status}";
    say "capture names: ", join(",", sort keys %+);
}
named: 192.168.1.42 asked for /papers?id=7 and got 200
capture names: bytes,ip,method,path,status,ts,user

Four features doing the work:

2.8 Regular expressions, part two: global matching and substitution

my $text = "contact alice\@example.com or bob\@test.org today";

my @emails = $text =~ /([\w.]+@[\w.]+)/g;      # list context + /g: every match
say "all emails: @emails";

while ($text =~ /(\w+)@([\w.]+)/g) {           # scalar context + /g: iterate
    say "  user=$1 host=$2 at offset $-[0]";
}

(my $masked = $text) =~ s/([\w.]+)@([\w.]+)/[redacted]\@$2/g;
say "masked: $masked";

my $prices = "cost: 10, 20, 30";
(my $doubled = $prices) =~ s/(\d+)/$1 * 2/ge;   # /e evaluates the replacement
say "doubled: $doubled";
all emails: alice@example.com bob@test.org
  user=alice host=example.com at offset 8
  user=bob host=test.org at offset 29
masked: contact [redacted]@example.com or [redacted]@test.org today
doubled: cost: 20, 40, 60

2.9 Greedy, lazy, and the mistakes everyone makes

my $html = '<b>bold</b> and <i>italic</i>';
my ($greedy) = $html =~ /<(.+)>/;
my ($lazy)   = $html =~ /<(.+?)>/;
say "greedy: $greedy";
say "lazy:   $lazy";
say "count of tags: ", scalar(() = $html =~ /<[^>]+>/g);
greedy: b>bold</b> and <i>italic</i
lazy:   b
count of tags: 4

.+ is greedy: it takes as much as it can and gives back only as needed, so it ran to the last > in the string. .+? is lazy and stops at the first. The third and best option is usually neither: [^>]+ says what you mean (characters that are not the terminator), is faster because it cannot backtrack, and does not depend on remembering which flavour of . you wanted.

scalar(() = $html =~ /.../g) is the countof idiom: assign the match list to an empty list in scalar context, which yields the number of elements. Ugly, universal, worth recognising.

The Mewlang cat, glancing sideways with annoyanceRegex mistakes that cost the most time
  • Forgetting \b. /ERROR/ matches NOERROR and ERRORS=0.
  • Unanchored patterns on structured data. /(\d{3})/ against a log line finds the first three digits anywhere, which may be part of the date. Anchor with ^, $, or surrounding context.
  • Using . where a negated class belongs. "([^"]*)" beats "(.*?)" for a quoted field: clearer and it cannot backtrack catastrophically.
  • Catastrophic backtracking. Nested quantifiers like (\s*\w+)*$ on a long non-matching line can take exponential time and hang your program. If a parser mysteriously stalls on one file, suspect this first.
  • Parsing nested structures with regexes. HTML and XML are not regular. Use XML::LibXML; we will in Milestone 6.
  • Not checking whether the match succeeded before using $1. On failure the capture variables keep their previous values, so a failed match silently reuses the last line's data. Always if ($line =~ ...) { ... }.

2.10 Files and streams

open my $fh, "<", "sample.log" or die "cannot open sample.log: $!";
my $errors = 0;
while (my $line = <$fh>) {
    chomp $line;
    $errors++ if $line =~ /^ERROR\b/;
}
close $fh;
say "errors: $errors";

open my $out, ">", "summary.txt" or die "cannot write: $!";
say {$out} "errors=$errors";
close $out;
errors: 2
wrote: errors=2

Encoding deserves a line now and a milestone later. A file is bytes; treating it as text requires knowing the encoding:

open my $fh, "<:encoding(UTF-8)", $path or die "$path: $!";
# and for data that claims to be UTF-8 and is not, in milestone 7:
open my $fh, "<:raw", $path or die "$path: $!";   # bytes, decode manually

2.11 Errors

my $result = eval {
    die "something broke\n";
    1;
};
say "eval returned ", (defined $result ? $result : "undef"), " and \$\@ is: $@";
eval returned undef and $@ is: something broke

Perl's exception mechanism is die to throw and eval { } to catch. The block returns undef on failure and the error lands in $@. Two conventions: end your message with \n or Perl appends " at script.pl line 12", which is useful for bugs and noise for expected failures; and put 1; as the last statement of the eval block so success is unambiguous.

$@ is a global, and it is easy to clobber between the eval and the check (a destructor running, a cleanup call). The community solution is Try::Tiny:

use Try::Tiny;

try {
    die { code => 503, message => "service unavailable" };
} catch {
    my $err = $_;
    say "Try::Tiny caught a ", ref($err), " with code $err->{code}";
};
Try::Tiny caught a HASH with code 503

Note that die can throw any reference, not just a string, which is how Perl does structured exceptions: a hash reference with a code and a message, or an object from a class like Throwable. For this project, structured errors matter because "line 40,112 of users.csv had unbalanced quotes" needs to be data, not prose.

2.12 Packages and modules

package Strata::Record;
use v5.36;

sub new ($class, %args) {
    my $self = {
        source => $args{source} // "unknown",
        fields => $args{fields} // {},
    };
    return bless $self, $class;
}

sub source ($self) { return $self->{source} }
sub field  ($self, $name) { return $self->{fields}{$name} }

sub to_line ($self) {
    my $f = $self->{fields};
    return join " ", map { "$_=$f->{$_}" } sort keys %$f;
}

1;   # a module must return a true value
apache: ip=10.0.0.1 status=200
ref: Strata::Record isa: yes

Perl's object system is three rules. A class is a package. An object is a reference that has been blessed into that package. A method call $obj->method(@args) calls the package's sub with the object as the first argument. That is all bless does: it writes the package name onto the reference so method lookup knows where to go.

It is minimal to the point of being spartan: no attribute declarations, no encapsulation (anyone can reach into $self->{fields}), no type checking. In production Perl most people use Moo or Moose, which add attributes, types, roles and defaults on top. And since 5.38 there is a real class syntax, still marked experimental:

use v5.38;
use experimental 'class';

class Strata::Entity {
    field $type :param;
    field $value :param;
    field $count = 1;

    method type  { $type }
    method seen  { $count++; $self }
    method to_string { "$type($value) x$count" }
}

my $e = Strata::Entity->new(type => "ip", value => "10.0.0.1");
$e->seen->seen;
say $e->to_string;
ip(10.0.0.1) x3

Genuinely pleasant, genuinely encapsulated (those fields are not reachable from outside), and genuinely experimental: the syntax may still change and it will warn unless you ask for it. This project uses plain bless, because it is what you will meet in existing code and because understanding it explains how Moo, Moose and the new class all work underneath.

Exercise 2.A

Write a sub parse_apache_line($line) that returns a hash reference of named fields for a combined-format Apache line, or undef if the line does not match. Requirements: use qr// with /x and named captures; handle the - that appears for a missing user or a zero byte count by normalising it to undef and 0 respectively; and split the request field into method, path and protocol.

Then write a second sub summarise(@lines) returning a hash reference with total lines, parsed lines, failed lines, and a count by status code. Test it with three good lines and two deliberately broken ones.

Solution 2.A — open after trying
use v5.36;

my $APACHE = qr{
    ^ (?<ip>\S+) \s+ (?<identd>\S+) \s+ (?<user>\S+) \s+
    \[ (?<ts>[^\]]+) \] \s+
    " (?<request>[^"]*) " \s+
    (?<status>\d{3}) \s+ (?<bytes>\d+|-)
    (?: \s+ " (?<referer>[^"]*) " \s+ " (?<agent>[^"]*) " )?   # combined format
    \s* $
}x;

sub parse_apache_line ($line) {
    return undef unless $line =~ $APACHE;

    my %f = %+;                       # copy: %+ is reset by the next match

    # "-" is Apache's way of saying "nothing here".
    $f{user}  = undef if $f{user} eq "-";
    $f{bytes} = 0     if $f{bytes} eq "-";

    if ($f{request} =~ m{^(?<method>[A-Z]+) \s+ (?<path>\S+) (?: \s+ (?<proto>\S+))?$}x) {
        @f{qw(method path proto)} = @+{qw(method path proto)};
    } else {
        $f{problem} = "unparsable request line";
    }

    return \%f;
}

sub summarise (@lines) {
    my %out = (total => 0, parsed => 0, failed => 0, by_status => {});

    for my $line (@lines) {
        $out{total}++;
        my $rec = parse_apache_line($line);
        if ($rec) {
            $out{parsed}++;
            $out{by_status}{ $rec->{status} }++;
        } else {
            $out{failed}++;
        }
    }
    return \%out;
}

Five things worth taking from this.

  • my %f = %+; copies the capture hash immediately. %+ is global and is reset by the next successful match anywhere, including the one two lines later that splits the request. Copying first is not optional; forgetting it produces a bug that appears only when you add a second regex.
  • The optional group makes one pattern handle two formats. Common and combined log formats differ only by the trailing referer and user-agent, so (?: ... )? parses both, and the fields are simply absent for the shorter one.
  • Normalising - at the boundary means the rest of the program never thinks about Apache's conventions. Every parser in this project will do this: the Record that comes out should not betray which format it came from.
  • A failed sub-parse becomes a problem field, not a discarded line. This is the project's central policy in miniature: keep the record, mark what is wrong with it, and let the caller decide. A line you throw away is a line you cannot investigate.
  • @f{qw(method path proto)} = @+{qw(method path proto)} is a hash slice on both sides: three assignments in one statement. This is where Perl's sigil rules start paying you back.

2.13 Sorting, formatting, and the report idioms

my %bytes_by_ip = ("10.0.0.1" => 4_112_883, "10.0.0.2" => 55_201, "10.0.0.9" => 913_004);

for my $ip (sort { $bytes_by_ip{$b} <=> $bytes_by_ip{$a} } keys %bytes_by_ip) {
    printf "  %-15s %10s\n", $ip, commify($bytes_by_ip{$ip});
}

sub commify ($n) {
    1 while $n =~ s/^(\d+)(\d{3})/$1,$2/;
    return $n;
}
  10.0.0.1          4,112,883
  10.0.0.9            913,004
  10.0.0.2             55,201

printf with %-15s (left-aligned, 15 wide) and %10s (right-aligned) is how every Perl report is formatted. commify is a classic: 1 while s/.../.../ repeats a substitution until it stops matching, which is a loop written as an expression. Numeric literals can contain underscores for readability.

When sorting by an expensive computed key, use the Schwartzian transform, which computes each key once:

my @sorted = map  { $_->[1] }
             sort { $a->[0] <=> $b->[0] }
             map  { [ expensive_key($_), $_ ] } @records;

Read it bottom-up: decorate each record with its key, sort by the key, undecorate. It is the standard idiom precisely because a naive sort { expensive($a) <=> expensive($b) } calls the expensive function O(n log n) times instead of n.

Exercise 2.B — the capstone of Part 2

The Mewlang cat, thinking with a paw to its chinWrite a single program, bin/toptalkers, that reads Apache logs from files or standard input and prints a report. Requirements:

  • Top 5 client IPs by request count, with their byte totals and error rates.
  • A count by status class (2xx, 3xx, 4xx, 5xx).
  • The five slowest-growing minutes by request volume (bucket timestamps to the minute).
  • A trailing line reporting how many lines could not be parsed, with the first three offending line numbers.
  • It must never load the whole file into memory, must work in a pipeline, and must exit 0 unless more than 1% of lines failed to parse, in which case exit 1.

Test it on a file with a few thousand lines, including some you have deliberately truncated mid-line.

Solution 2.B — open after trying
#!/usr/bin/env perl
use v5.36;

my $APACHE = qr{
    ^ (?<ip>\S+) \s+ \S+ \s+ (?<user>\S+) \s+
    \[ (?<ts>[^\]]+) \] \s+ " (?<request>[^"]*) " \s+
    (?<status>\d{3}) \s+ (?<bytes>\d+|-)
}x;

my (%hits, %bytes, %errors, %by_minute, %status_class);
my ($total, $failed, @first_failures) = (0, 0);

while (my $line = <>) {
    $total++;

    unless ($line =~ $APACHE) {
        $failed++;
        push @first_failures, "$ARGV:$." if @first_failures < 3;
        next;
    }

    my %f = %+;
    my $bytes = $f{bytes} eq "-" ? 0 : $f{bytes};

    $hits{ $f{ip} }++;
    $bytes{ $f{ip} } += $bytes;
    $errors{ $f{ip} }++ if $f{status} >= 400;
    $status_class{ substr($f{status}, 0, 1) . "xx" }++;

    # 10/Oct/2026:13:55:36 +0000  ->  10/Oct/2026:13:55
    $by_minute{$1}++ if $f{ts} =~ /^(\d+\/\w+\/\d+:\d+:\d+)/;
}

say "top talkers";
my @top = (sort { $hits{$b} <=> $hits{$a} || $a cmp $b } keys %hits)[0 .. 4];
for my $ip (grep { defined } @top) {
    printf "  %-15s %8d hits  %12s bytes  %5.1f%% errors\n",
        $ip, $hits{$ip}, commify($bytes{$ip}),
        100 * ($errors{$ip} // 0) / $hits{$ip};
}

say "\nstatus classes";
printf "  %s %8d\n", $_, $status_class{$_} for sort keys %status_class;

say "\nbusiest minutes";
my @busy = (sort { $by_minute{$b} <=> $by_minute{$a} } keys %by_minute)[0 .. 4];
printf "  %-22s %8d\n", $_, $by_minute{$_} for grep { defined } @busy;

my $rate = $total ? 100 * $failed / $total : 0;
printf "\n%d lines, %d unparsed (%.2f%%)%s\n", $total, $failed, $rate,
    @first_failures ? " first at: " . join(", ", @first_failures) : "";

exit($rate > 1 ? 1 : 0);

sub commify ($n) { 1 while $n =~ s/^(\d+)(\d{3})/$1,$2/; return $n }

Six details that are the actual lesson.

  • Five hashes, one pass. Everything is accumulated in a single traversal, so memory is proportional to the number of distinct IPs and minutes, not to the file. That is the shape of every streaming aggregation you will write.
  • $ARGV and $. give you free provenance. Inside <>, $ARGV is the file currently being read and $. is the line number, so "where did this come from" costs nothing. (One caveat worth knowing: $. does not reset between files unless you close ARGV at eof.)
  • Unparsed lines are counted and located, never silently dropped. Keeping the first three line numbers rather than all of them bounds the memory a pathological file can cost you.
  • (sort ...)[0 .. 4] takes a slice of a list without an intermediate array, and grep { defined } handles the case where fewer than five exist. Slicing past the end gives undef, not an error, which is convenient and requires the guard.
  • ($errors{$ip} // 0) because an IP with no errors has no key, and arithmetic on undef warns. Defined-or is the right operator: a genuine zero must survive.
  • The exit code encodes a judgement (more than 1% unparsed means something is wrong), which makes the tool usable in a cron job that alerts when a log format changes underneath you. That is a real failure mode and this is how you catch it.

If you built something close to this, Milestones 1 through 3 will feel like tidying rather than learning, which is the intention.

Common mistakes in Part 2
  • Using == on strings or eq on numbers. "10" == "10.0" is true; "10" eq "10.0" is false. Both are sometimes what you want.
  • Forgetting chomp, then wondering why $fields[-1] never matches anything.
  • Using $1 without checking the match succeeded. It holds the previous match's value.
  • Not copying %+ before the next match. Same class of bug, more surprising.
  • Sorting without a comparator and getting string order on numbers.
  • Passing two arrays to a sub and receiving one flattened list. Pass references.
  • Accidental autovivification when testing nested keys. Use exists.
  • Two-argument open, ever.
  • Slurping a file (my @lines = <$fh>) out of habit. It works until the file is 40 GB.

Part 2 checkpoint

  1. Why is it $steps[0] and not @steps[0]?
  2. Give two expressions that behave differently in scalar and list context, and say what each does.
  3. What does $count{$_}++ for @items do, step by step?
  4. Why does nesting data require references, and what is autovivification?
  5. When would you use local rather than my?
  6. What do /x, qr// and (?<name>...) each buy you?
  7. What is the difference between /g in list context and in scalar context?
  8. Why must you copy %+ immediately?
  9. Why is three-argument open the only acceptable form?
  10. What does bless actually do?
Why are we using this language here? (Part 2 summary)

Look back at the capstone solution. The regex is a first-class value with named fields laid out over five commented lines; the input loop handles files and pipes with no imports; five aggregations happen in one pass with hashes that spring into existence as needed; provenance comes free in $ARGV and $.; and the whole thing is sixty lines and streams a file of any size. The equivalent Python is perhaps twice as long and has re.compile, fileinput, defaultdict and argparse visible in it. That difference is Perl's argument, and for this kind of work it is a strong one.

The costs are equally visible. Context means my $x = f() and my ($x) = f() are different programs. Capture variables are global and get clobbered. local is a scoping rule most programmers have never met. Sigils change with what you are asking for rather than what the variable is. None of these is hard once learned, and all of them are sharp edges that a language designed in 2010 would not have.

The discipline that makes Perl maintainable is not subtle, and it is all in this part: use v5.36, named captures, /x on anything long, real data structures instead of clever parallel arrays, and tests from day one.

What is next

Milestone 1 turns the one-liner instinct into a program: a filter with proper option handling, exit codes and tests. Milestone 2 adds the aggregation you just wrote by hand. Milestone 3 is where the parser becomes serious, with a pattern library, failure accounting, and the first fixtures of deliberately broken input.

Before then: run the one-liners from Part 1 against a log file on your own machine, and read perldoc perlretut. It is the best forty minutes available to you at this point.

The Mewlang cat, walking away in a rear viewInstalment 11 of the five-course curriculum. Next: Perl Milestones 1–4, where the filter becomes a tool, hashes become reports, regexes become a parser, and records become a data model.

Continue