HOWL

Your machines howl home.

A small agent on every machine ships its logs to one receiver, alerts you in seconds, and updates itself. Nothing ever dials in.

The pack

One howl becomes a pack.

Add a machine and it simply starts howling too. No wolf talks to another wolf - every howl goes to the same moon. That is the whole topology.

The metaphor is the architecture

The moon is the server.

Howl is a tiny log-shipping fleet. Every machine runs a small agent that reads its own log files and pushes them out. Nothing ever connects to a machine - no open ports, no VPN, works from any network.

Wolf→agent

A ~56 MB process tailing log files with byte-exact checkpoints. Measured at 0.2% of one core.

Howl→HTTPS push

Batches ship within about a second. If the moon is unreachable, the log file itself is the buffer.

Moon→receiver

A Cloudflare Worker + D1 + R2. It lives outside your house on purpose - the monitor never shares fate with the monitored.

Pack→fleet

Wolves never talk to each other. Adding a machine is registering a token and letting it howl.

Watch one line travel

Read, ship, then move the cursor.

The agent keeps a byte cursor in every file it watches. It reads what is past the cursor, ships it, and waits for the receiver to say OK - and only then does the cursor move. That ordering is the whole delivery guarantee.

/var/log/syslogcursor 4,096
read up to here

writea process appends bytes to the file

Bytes in, lines out. The agent never asks for “the next line” - it reads a window of bytes from its cursor and splits what comes back on newlines. If the window ends mid-write, that trailing fragment is left alone and the cursor stops at the last newline, so a half-written line is never shipped as if it were whole.

Why the pause matters. The cursor moves only after the receiver acknowledges. Crash in between and the same batch simply goes again - duplicates are possible, gaps are not. If the receiver is unreachable, the cursor stops advancing and the log file itself becomes the buffer.
Under the fur

How the data flows.

Data pushes out; code pulls in. Agents checkpoint after every acknowledged batch, so crashes and restarts re-send instead of losing. The receiver classifies, stores hot, rolls cold, and watches for silence.

machine 1agent tails + checkpointsmachine 2agent tails + checkpointsmachine 3agent tails + checkpointsHTTPS push, ~1sthe receiverclassify error / warn / infohot: 7 days queryablecold: nightly gzip archivewatchdog: silence = alertserverless - no box to babysityour phoneerrors + silence, in secondsquery APIhot + cold, one viewgit mainagents self-update every 15 mincode pulls in

The life of a line

A line is queryable about a second after it is written, stays hot and filterable for a week, then rolls itself into cheap cold storage overnight - and it only leaves the hot store once the archive is confirmed written.

HOT · D1 · 7 DAYSCOLD · R2 · KEPTline writtent = 0shipped + acked~1 sclassified + indexedqueryable nowstill hotday 1 - 7gzip archiveone file per machine-dayfetched by datedownload + grepnightly rollwrite → verify → delete: rows leave D1 only after the archive is confirmed
hot · D1queryable, any filter
cold · R2gzip, by date

ingestlines land in today’s bucket — indexed and queryable immediately

Built like it matters

Small agent, serious guarantees.

Never loses a line

Checkpoints advance only after the server acknowledges. Rotation, deletion, truncation, even in-place rewrites - drained and fingerprinted, proven by a zero-loss test harness.

Alerts that respect you

One push per incident, not five hundred. Errors buzz in seconds; a machine going quiet buzzes once; a source that stops shipping while its machine looks healthy gets caught too.

Updates itself

Push to main and every wolf is running it within fifteen minutes. Failed builds keep the old agent alive. Restarts are lossless by construction.

0.2%

of one CPU core, measured on a live machine. 56 MB of memory. Under 2 MB of network a day when idle.

~1s

from a line hitting a log file to it being queryable at the receiver, classified and stored.

$0

of always-on servers. The receiver is a Cloudflare Worker; a home fleet fits comfortably in free tiers.

Not crying wolf

The wolf that howls at everything gets ignored.

Every monitoring tool can send you an email. The hard part is being worth reading. Point a naive rule at a real fleet and it finds the word “error” everywhere - in success messages, in service descriptions, in prose - and buries the one line that mattered under hundreds that did not. Mute it once and you have no monitoring at all, only the feeling of it.

So Howl alerts on novelty, not volume. A line is an error because of where its severity sits, not because a word appears in it. Errors are reduced to their shape - timestamps, pids and addresses stripped away - and a shape wakes you exactly once. Everything after that is a number in the next digest.

incoming lines
what reaches you
nothing yet - that is the point
0 never classed an error0 folded into a shape you know
tomorrow’s digest: certbot renew failure ×1 - still broken, still not news

Position, not presence

Finished unit.service - Download data for packages that failed is systemd announcing success; only the unit’s description says “failed”. Real logs are full of these. Severity is read where log formats put it.

A shape, not a line

One service failing twice a day for six days is one fact, not seventy-seven. It pages the first time, then reports as a count - so “still broken” stays visible without ever pretending to be news.

Silence is a signal too

A machine that stops talking is alerted on within minutes, and a source that was always chatty and went quiet is caught separately - a wedged agent still sending heartbeats is the failure nobody notices.

Volume has an opinion

Each machine is measured against its own median hour. Eight times its normal chatter raises an alert; past a hard ceiling, ingest answers 429 and the agent simply waits - the log file was always the buffer.

The measure of this is uncomfortable and worth stating plainly: on the fleet that Howl was built for, a day’s logs held 179 lines that looked like errors and about two that deserved a human. Both were real - a certificate renewal that had been failing silently for six days, and a VPN that had quietly logged itself out. Everything else was noise wearing the word “error”.

The trust model

A log agent sees a lot. It should be able to do very little.

You are installing a long-running process on every machine you own, so blast radius deserves a straight answer. Howl never runs as root: the agent is an unprivileged service that can read the log files you name and talk HTTPS outbound - nothing else.

Nothing can dial in

Agents open outbound HTTPS and nothing else. No listening socket, no port to forward, no VPN, nothing to scan. A machine behind hotel wifi works exactly like one in your rack.

One token per machine, hashed

Each machine gets its own bearer token, shown exactly once at registration; only a SHA-256 hash is stored. Revoke or rotate one machine without touching the rest, and a stolen token can ship logs but never read them.

The admin key never leaves your desk

Registering machines and reading logs use a separate admin key that no fleet machine ever holds. Compromising an agent gets you the ability to send lines - not the fleet.

Read-only deploy keys

Machines auto-execute whatever they pull, so each gets a read-only key of its own. A write-capable key on any box would let one break-in push code the whole fleet runs fifteen minutes later. That path simply does not exist.

No root, anywhere

The agent runs as its own system account with no capabilities, no shell, a read-only view of the filesystem, and write access to one state directory. It reads the logs you name because that account is in the log-reading group - not because it is powerful. Even self-updating needs no privilege: the updater signals the agent and the service manager restarts it.

Failures are loud

A wedged or killed agent stops heartbeating, and the watchdog pages you within ten minutes. Silence is treated as a fault, not as good news - the failure mode most monitoring gets wrong.

1.7
Scored by systemd's own auditor, not by us.

systemd-analyze security howl-agent rates the shipped unit 1.7 “OK”. Most of what remains is the job itself: half the residual score is simply having network access, which a log shipper cannot do without. Run it on your own machine and check.

Verify the footprint yourself. The resource claims above are not marketing numbers - the unit ships with kernel accounting on, so one command tells you the truth on your own machine: systemctl show howl-agent -p CPUUsageNSec -p MemoryCurrent -p IPEgressBytes
Join the pack

A new wolf in one command.

# 1. give the machine its own read-only deploy key, register the .pub
ssh-keygen -t ed25519 -N '' -f ~/.ssh/howl_ed25519

# 2. register the machine, copy the token it returns
curl -X POST https://your-moon.example/api/machines -d '{"name":"newbox"}'

# 3. install: service account, sandboxed units, config, start
sudo bash deploy/install.sh --key ~/.ssh/howl_ed25519 --token howl_...

The installer creates the unprivileged service account, places the key, builds, installs the sandboxed units, writes the config and starts the agent - then prints its first shipped lines and its security score. It is idempotent, so re-running it heals a half-finished machine.

Open source, soon. Howl runs a real fleet today. The repo goes public once the rough edges are sanded - the design writeup is already out: Working Through Log Data at Datacenter Scale.