Skip to content
Infrastructure · 5 min read

Put your logs somewhere the burning host cannot reach

The claim Logs that live only on the machine that produced them are useless in exactly the situation you most need them: when that machine has failed. The disk fills, the host is t...

A Written by Administrator
Put your logs somewhere the burning host cannot reach

The claim

Logs that live only on the machine that produced them are useless in exactly the situation you most need them: when that machine has failed. The disk fills, the host is terminated and replaced, the container is rebuilt, the instance is compromised and you can no longer trust anything on it — and in every one of those cases the record of what happened is gone with the thing that failed. Logs must be shipped off the host as they are written, to somewhere the host's failure cannot touch, or you are keeping a diary you will lose on the day it matters.

The failures that erase local logs

Four ordinary events destroy on-host logs, and none of them is exotic. A full disk stops the application and often corrupts the very log that would explain why. An auto-scaling group or orchestrator terminates an unhealthy instance and spins up a fresh one, taking the failed instance's logs with it — the logs describing the failure that triggered the termination. A container restart wipes anything not on a mounted volume. And a compromised host means you cannot trust its logs at all, because the first thing a competent intruder does is edit them. In each case, the moment you reach for the logs is the moment they are unavailable.

Ship as you write

The principle is that a log line should leave the host almost as soon as it is written, so that the copy you rely on already lives somewhere else before the host has any chance to fail. A lightweight forwarder reads the log stream and sends it to a central collector continuously:

# Vector, a small forwarder, tailing the journal and shipping onward
[sources.app]
type = "journald"

[sinks.central]
type = "loki"                      # or elasticsearch, s3, a hosted service
inputs = ["app"]
endpoint = "https://logs.internal.example.ca"

The forwarder buffers locally for the seconds or minutes a network blip might last, then delivers when the link returns, so a brief outage of the log destination does not lose data and does not block the application. What you are buying is that the authoritative copy is never the on-host copy — the host's copy becomes a convenience, and its loss becomes survivable.

Central logging is also how you search across hosts

The disaster-recovery argument is the headline, but centralised logs solve a second problem that shows up every day: correlating an event across more than one machine. When a request passes through a load balancer, two application instances, and a database, its story is spread across four hosts, and reconstructing it by SSHing into each one in turn is slow at the best of times and impossible during an incident when one of them is down. A central store lets you query all of them at once:

{service="orders"} | json | req_id="01J2K9X4"
# one query, every host that touched the request, in order

This is the same request-ID correlation that is hard to do on a single host and impossible to do across several without shipping the logs somewhere common first.

What to send, and what never to

Centralising logs multiplies the number of places they exist and the number of people who can read them, which sharpens the rule about what belongs in a log at all. Never ship secrets, full payment details, session tokens, authorization headers, or complete personal records — the accident of logging an entire request body on error becomes far more serious when that body lands in a searchable central store that contractors and support staff can query. Structure the logs as JSON with named fields so the central store can index them, keep identifiers and last-four digits rather than values, and decide deliberately, in a written retention policy, what personal data may appear and for how long.

Retention and cost, decided on purpose

Central logging has a cost that on-host logging hides, and the cost is real enough that it must be a decision rather than a surprise. Hosted log services typically charge by volume ingested, and a chatty application can generate more gigabytes per day than anyone expected. Split retention by how long each tier is actually useful:

TierRetentionWhere
Hot, searchable14–30 daysIndexed log store
Warm, retrievable3–12 monthsCompressed in object storage
Cold, compliance onlyAs requiredCheapest archive tier

Most searching happens against the last few days, so keep only that hot and indexed, and roll the rest to object storage where a gigabyte costs a few cents rather than the premium an indexed store charges. This keeps the bill proportional to the value each tier delivers.

The test

Ask one question of your current setup: if your most important host disappeared right now — terminated, wiped, gone — could you still read the last hour of its logs? If the answer is yes, your logs are shipped and you are covered. If the answer is that you would have to recover the host first to read the logs explaining why the host failed, then your logging is circular, and the day you need it most is the day it will not be there.

#logging #observability #disaster recovery #operations

Keep reading