SREObservabilityGrafanaAlloy
How Alloy's Default max_shards Turned a Mimir Blip Into a Production Monitoring Blackout
When Grafana Mimir slowed down, Alloy's default retry configuration — 200 parallel shard workers — consumed the majority of CPU on a shared VM, starving the co-located business process and forcing an engineer to kill the observability agent during an active incident. Here's the exact three-layer fix and why shared-VM deployments require explicit resource budgets at every level.