A year ago we found out production was down when a customer told us.
That sentence is worth sitting with, because it describes an ordering problem, not a monitoring problem. Every outage was a support incident before it was an engineering incident. The customer's experience of the failure preceded ours. Everything downstream of that — the scramble, the guessing, the apology written before anyone knew what broke — follows from those two events being in the wrong order.
There were no metrics, no alerts, no dashboards. Not inadequate ones. None.
Four questions
I found it useful to stop thinking about tools and start thinking about the questions a platform has to be able to answer, because the tools follow trivially once the question is clear and not at all if it isn't.
See
Do we know what production is doing? Metrics, probes and alerts on every machine — each product with a page of its own.
Catch
Do we find out before the customer does? Diagnostics that arrive on their own, with the evidence already attached.
Prove
Can we show it works, or do we just believe it? A suite running continuously, and a CI loop short enough to wait for.
Defend
Do we find our own weaknesses first? Scanning as a gate, supply chain closed, audits pointed inward.
This post is about the first one.
What got built
Prometheus and Grafana across the whole fleet — metrics from every machine, continuously. Blackbox probes on top, because CPU and memory tell you a server is alive and tell you nothing about whether the service is answering. Those are different questions and only one of them is the customer's.
Alert routing with thresholds chosen to arrive early rather than often: disk at 80% warns and at 90% pages, CPU sustained over ten minutes, certificates flagged days before expiry rather than the morning a browser starts refusing the page.
And a consolidated inventory of the entire estate — roughly fifty machines, what each one is, what it costs, what it's doing right now. That last one started as an accounting exercise and turned out to be the most useful artifact of the whole project, because "is there anything running that nobody remembers" has an answer now, and for a while the answer was yes.
Nothing configured by hand
Every dashboard and every alert rule lives in git and ships through the same reviewed pipeline as product code.
The rule is that the Grafana you're looking at and the Grafana in the repository are the same Grafana. Console-clicked monitoring drifts — someone silences an alert during an incident, or bumps a threshold to stop the noise, and eighteen months later nobody can tell you why the disk warning is at 95%. Making it a pull request costs about ninety seconds and buys you a permanent answer to "why is this like this."
It also makes the whole thing handover-shaped, which matters more to me than any of the graphs. Adding a machine, a dashboard or an alert is a small PR with an obvious shape to copy. Deliberately hard to get wrong.
Every project gets its own page
The instinct is to build one big dashboard with everything on it. Don't.
A dashboard everybody owns is a dashboard nobody opens. It's too dense to scan, most of it is irrelevant to whoever's looking, and there's no moment where a specific person feels responsible for a specific red square.
Each of our products has its own view: its own service availability, its own endpoints, its own machines, its own environments. A team recognises its own page, and recognition is the whole mechanism — data that exists but never gets looked at is a rounding error away from no data at all.
What changed
The mode of operating, more than any single number.
Before, someone reported a problem, an engineer dropped what they were doing, the evidence had usually already expired, the fix was a guess applied under time pressure, and the same class of problem came back unrecognised a month later.
Now the system says something is wrong and brings the evidence with it. Recurring patterns are visible, so they get fixed at the root instead of being rediscovered. Capacity problems show up as trends weeks before they're outages. Certificate and disk expiry — the classic everything broke on a Saturday category — became calendar items.
We've also caught cloud-provider incidents on our own dashboards and raised them with the provider ourselves, with our own evidence, rather than waiting to be told. That felt like the moment the thing was actually working.
The honest caveat
None of this is glamorous and none of it demos well. There's nothing to show at a sprint review, because the demo is that your Tuesday was uneventful.
That's the deal with this category of work. It's invisible when it succeeds, which means the only way it survives is if someone can explain, afterwards and in plain sentences, what would have happened otherwise.