How Alloy's Default max_shards Turned a Mimir Blip Into a Production Monitoring Blackout
When Grafana Mimir slowed down, Alloy's default retry configuration — 200 parallel shard workers — consumed the majority of CPU on a shared VM, starving the co-located business process and forcing an engineer to kill the observability agent during an active incident. Here's the exact three-layer fix and why shared-VM deployments require explicit resource budgets at every level.
Bridging OT and IT Observability: You Meet at the Historian, Not the PLC
Manufacturing observability tempts you to instrument the plant floor directly — point a collector at the PLCs, scrape everything. That path runs into a safety boundary, an air gap, and a cardinality wall: an unscoped Windows process collector was already emitting 3,900 series per plant mid-rollout. The integration point that actually works is the process historian, and the fragile link is a service account nobody tested.
An SLO Without an Error-Budget Policy Is Just a Chart
Most teams that 'do SLOs' have a number in a doc and a Grafana panel showing attainment. What they don't have is the thing that makes an SLO a business instrument: a written policy that changes what the team does when the budget runs out. Without the freeze-or-ship consequence, an SLO is a vanity metric with a burn-rate alert nobody has agreed to obey.
One Un-Instrumented Hop Breaks the Whole Trace
Distributed tracing gives you the illusion of end-to-end visibility right up until a request crosses a message queue, a legacy proxy, or a cron-triggered batch job that doesn't forward the trace context. The span chain snaps, Tempo stores two unrelated fragments, and the 3am question — where did the time go — has no answer. Context propagation is the whole game.
Every SRE Practice Is Downstream of Observability — Here's the Dependency Chain
SLOs, error budgets, burn-rate alerting, incident response, blameless postmortems, capacity planning, chaos engineering — each one consumes observability data as input. On a signal layer you can't trust, they degrade to opinion with a dashboard. The Reactive-to-Autonomous journey is really the maturity of that signal layer, and the 80%+ alert-noise cut came from making it trustworthy enough for automation to stand on.