← All writing

Production lessons

The green dashboard lied to us

Every metric was within threshold. The service was down for half the users. Availability and health are not the same thing.

The status page was green. CPU was fine. Memory was fine. Error rate was below the alert threshold. Support tickets were not fine.

Half our users could not complete checkout. The other half had no idea anything was wrong. We spent forty minutes proving the system was healthy before we started looking for who it was healthy for.

Aggregates hide the users who matter

A 0.3% error rate sounds acceptable until you learn the errors are concentrated on one payment provider, one region, or one API version. Averages are comfortable. Incidents are specific.

We had a dependency that started failing for requests carrying a particular header. Global error rate barely moved because most traffic did not send that header. The checkout flow did. The dashboard averaged checkout failures with health-check successes and called it a good hour.

Since then I watch error rates broken down by route, by downstream dependency, and by customer segment when we have enough volume. One graph that says “all good” is not evidence. It is a summary that may be lying.

Health checks measure the wrong thing

Our liveness probe hit /health and got 200. The endpoint checked that the process was running and the database connection pool could open a connection. It did not check that we could write an order, that the message broker accepted a publish, or that the feature flag service returned a config newer than last week.

A service can be alive and useless. Health checks should reflect what “working” means for that service, not what is cheap to ping.

We added a readiness check that runs a lightweight version of the critical path — read config, touch the database with a real query shape, verify the broker connection. It fails more often. That is the point. I would rather take a pod out of rotation during a partial outage than keep sending it traffic it cannot serve.

Alerts should fire on symptoms users feel

We used to alert on infrastructure: disk, CPU, pod restarts. Those are useful for capacity planning. They are late indicators for user-facing failure.

Now we alert on things closer to the user: checkout success rate, payment confirmation latency, message processing lag for the queue that sends receipts. Infrastructure alerts still exist, but they are secondary.

The shift was uncomfortable because user-symptom alerts are noisier at first. You have to tune them. You have to accept that “error budget” is a product conversation, not only an ops one. The alternative is learning about problems from Twitter before you learn from PagerDuty.

What we changed after that incident

Correlation ids on every request, propagated through messages. Not because it is best practice on a slide, but because support ticket 8841 had no thread through the logs.

Synthetic checks that run checkout every few minutes from outside the cluster. Cheaper than an on-call page, more honest than an internal health endpoint.

Runbooks that start with “which slice of traffic is affected?” instead of “restart the service”.

The dashboard still goes green. I just no longer treat green as the end of the question. Green means the aggregates we chose to watch look fine. It does not mean the system is fine for everyone using it right now.

That distinction cost us forty minutes once. Remembering it costs nothing.