Container crashes were the least debuggable failure we had, and the reason is structural rather than technical: the evidence dies with the process.
A container exits. The orchestrator does exactly what you asked it to and restarts it. By the time a human opens a terminal — minutes later, hours later, next morning — the logs are gone, the memory state is gone, the process is gone. What's left is a service that's running fine and a vague report that something was weird earlier.
Every investigation started from zero. Most of them ended there too.
The idea
Aircraft solved this in 1958. You don't reconstruct the failure afterwards; you record continuously and read the recorder.
So: a small utility installed as a system service on every host, watching the container runtime's event stream. The instant a container exits, it captures everything a person would want and posts it to the team channel — usually before anyone has noticed anything happened.
What it captures is just the set of things I found myself asking for every single time:
- Exit code, and whether the kernel killed it
- Memory at the moment of exit, against the limit
- How long it had been up
- How many times it had restarted in the last day
- The image and tag, so "which version was this" is never a question
- The final lines of log output, which is the part that actually dies
Why the last log lines matter most
Everything else in that list can be reconstructed later from somewhere. The last few log lines usually can't — they're the ones buffered at the moment the process died, the ones that never made it anywhere durable.
They're also, almost always, the answer. A worker logging allocating parse buffer and then cannot allocate memory immediately before an exit code 137 is not a mystery requiring investigation. It's a sentence.
The gap this closes isn't we had no data. It's we had the data for four seconds and nobody was looking.
What it found immediately
This is the part I didn't expect, and it's the argument for building this kind of thing early rather than when you think you need it.
Within days it surfaced problems that had been happening for months:
- Memory-hungry jobs being killed mid-work, silently, on a schedule
- Restart loops that had never reached a single person's attention
- Compute-heavy tasks that had outgrown the machines they'd been placed on years earlier
None of that was new. Every one of those had been happening the whole time. The only thing that changed is that we now knew.
That's worth being precise about, because it's easy to present this as we fixed a lot of bugs when what actually happened is we discovered a backlog that already existed. Monitoring doesn't reduce your failure rate on the day you install it. It converts unknown failures into known ones, and the first week's numbers look like a regression if you don't say so out loud.
The mode change
Before, a class of failure existed that we simply could not investigate. Not "hadn't got to" — couldn't. The information required didn't survive long enough to be looked at, so those incidents got closed as unexplained and quietly forgotten.
Now the first thing anyone sees is the answer rather than the question, and the work moves from firefighting to scheduling: a capacity trend gets a ticket instead of an outage getting an apology.
If you build one
Two things I'd do the same way again.
Push it where people already are. The diagnostics land in the team channel, not in a dashboard someone has to remember to open. Something that requires an act of remembering will not be remembered at 3am.
Capture more than you think you need. Storage is free and the failure is not reproducible — that's the entire premise. Every field I hesitated over has since been the one that mattered at least once.
And one thing I'd change: I should have written it two years earlier. The cost was a couple of days. The thing it was protecting had been broken, invisibly, for most of the life of the platform.