All posts

#Mimir

5 posts

OpenTelemetryGrafanaMimirLoki

How a 'Consistency Fix' in Program.cs Silently Broke Label Promotion Across Mimir and Loki

Renaming an OTel resource attribute from deployment.environment to deployment_environment in application code looked like the fix for inconsistent labels across Grafana Cloud's three signals. It's backwards: Mimir and Loki both promote and convert dotted attribute names on ingestion, so a pre-converted attribute silently fails to match and drops out of both label sets. The real fix belongs in the Alloy pipeline, scoped to traces only.

Read
SREObservabilityPrometheusGrafana

Unbounded Cardinality, Zero Alerts, and 14-Second Dashboard Loads: The Structural Limits of the Grafana Azure Monitor Plugin

The Grafana Azure Monitor datasource plugin connects directly to Azure Monitor at query time. On a 12-panel dashboard, that means 12 sequential ARM API calls on every open — 14 seconds before the first chart renders. It cannot feed Mimir alerting rules. It cannot attach environment labels without custom queries per panel. And promoting Azure tags as Prometheus labels explodes cardinality to thousands of series per resource group. This post documents the five structural gaps and the specific design decisions in a push-based exporter that close each one.

Read
ObservabilityPrometheusMimirFinOps

Why Cardinality Kills Observability Platforms (and How to Stop It)

Cardinality is the silent killer of Prometheus-based observability platforms. Here's how it happens, how to detect it early, and the label schema discipline that keeps ingestion costs sane at scale.

Read
ObservabilityOpenTelemetryPrometheusMimir

Push vs Pull Was Never the Point — Rethinking Metrics for the OTLP Era

The push-versus-pull debate is settled for application metrics: OTLP push through a collector. The migration that actually matters is what comes with it — resource-attribute promotion instead of relabel configs, exemplars linking metrics to traces, and a deliberate delta-vs-cumulative temporality choice. Lift-and-shift your Prometheus scrape config into OTLP and you keep the mechanics while losing the point.

Read
ObservabilityOpenTelemetryTempoLoki

The Three Pillars of Observability Are a Storage Detail, Not a Strategy

Metrics, logs, and traces name three storage engines with different index models and cost curves — not three things to instrument separately. Teams that organize around the pillars build three disconnected silos and still can't say why one request was slow. The unit that matters is the correlated event: one trace_id and an exemplar that stitches a histogram bucket to the trace to the logs.

Read