Picture this: It’s 3:00 AM. Your pager is screaming. You open your observability platform and are greeted by a wall of 100 perfectly green dashboard widgets. The CPU is fine. Memory is fine. Network throughput is steady.
Meanwhile, on Twitter, customers are screaming that your checkout page is throwing 500 errors.
If this has happened to you, you don’t have an observability platform; you have a data hoarding problem.
The “Log Everything” Trap
In cloud-native, distributed architectures, the default instinct is to log everything and emit metrics for every conceivable variable. We hoard data like digital packrats, assuming that when an incident happens, the answer will be buried somewhere in the noise.
This leads to three massive problems:
- Alert Fatigue: Engineers get conditioned to ignore alerts because 90% of them are non-actionable noise.
- Massive Costs: Ingesting and storing terabytes of useless logs burns infrastructure budgets rapidly.
- High MTTR (Mean Time to Recovery): When something actually breaks, finding the root cause is like finding a needle in a haystack of your own making.
Shifting to User-Centric Metrics
We need to stop obsessing over infrastructure metrics. Your users do not care if a node’s CPU is at 90%. They only care if the application is fast and functional.
Instead of tracking infrastructure, track the Golden Signals:
- Latency: How long does it take to service a request?
- Traffic: How much demand is being placed on your system?
- Errors: What is the rate of failed requests?
- Saturation: How “full” is your system compared to its maximum capacity?
When you align your telemetry with the actual user experience, your dashboards instantly become meaningful.
Anatomy of a High-Trust Alert
The golden rule of alerting is simple: If an alert fires and a human does not need to take immediate action, delete the alert.
A high-trust alert provides absolute clarity at the moment of peak stress. Every page that wakes an engineer up should contain three things:
- What is broken? (e.g., “Checkout API error rate is > 5%”).
- Who is impacted? (e.g., “EU customers cannot complete purchases”).
- What do I do? (A direct link to a runbook and the relevant trace dashboard).
The Takeaway
Great telemetry isn’t about capturing every single micro-event in your cluster. It’s about surfacing the right context at the exact moment an engineer needs to make a critical decision. Cut the noise, focus on the user, and make every alert actionable.