All posts

#Grafana

11 posts

OpenTelemetryGrafanaMimirLoki

How a 'Consistency Fix' in Program.cs Silently Broke Label Promotion Across Mimir and Loki

Renaming an OTel resource attribute from deployment.environment to deployment_environment in application code looked like the fix for inconsistent labels across Grafana Cloud's three signals. It's backwards: Mimir and Loki both promote and convert dotted attribute names on ingestion, so a pre-converted attribute silently fails to match and drops out of both label sets. The real fix belongs in the Alloy pipeline, scoped to traces only.

Read
SREObservabilityPrometheusGrafana

Unbounded Cardinality, Zero Alerts, and 14-Second Dashboard Loads: The Structural Limits of the Grafana Azure Monitor Plugin

The Grafana Azure Monitor datasource plugin connects directly to Azure Monitor at query time. On a 12-panel dashboard, that means 12 sequential ARM API calls on every open — 14 seconds before the first chart renders. It cannot feed Mimir alerting rules. It cannot attach environment labels without custom queries per panel. And promoting Azure tags as Prometheus labels explodes cardinality to thousands of series per resource group. This post documents the five structural gaps and the specific design decisions in a push-based exporter that close each one.

Read
SREObservabilityGrafanaAlloy

How Alloy's Default max_shards Turned a Mimir Blip Into a Production Monitoring Blackout

When Grafana Mimir slowed down, Alloy's default retry configuration — 200 parallel shard workers — consumed the majority of CPU on a shared VM, starving the co-located business process and forcing an engineer to kill the observability agent during an active incident. Here's the exact three-layer fix and why shared-VM deployments require explicit resource budgets at every level.

Read
ObservabilityOpenTelemetrySRESystem Design

You Don't Get Root on SAP RISE: Observability Inside a Managed Black Box

RISE with SAP hands you a mission-critical ERP landscape and takes away the thing every observability playbook assumes: OS access to put an agent on the host. The workable model is to monitor the contract, not the internals — synthetic business transactions, the interfaces SAP does expose, and the integration layers you still own. Plan to drop Alloy on the app servers and the project stalls at 'we can't get in.'

Read
ObservabilityPlatform EngineeringGrafanaSRE

Self-Service Observability Is a Paved Road, Not an Explore Button

Handing developers an Editor role and an empty Explore tab is not self-service — it's abdication. Real self-service is a paved road: a golden-signal dashboard generated from the service name, an SLO template, and label-based access control that scopes a team to its own data. Without the paving you get hundreds of one-off dashboards, cardinality bombs nobody owns, and every team able to read every other team's telemetry.

Read
GrafanaTerraformObservabilityPlatform Engineering

The Grafana Terraform Provider Silently Drops LBAC Rules — and the Three-Layer Fix

When we codified Grafana Cloud multi-tenancy as Terraform, label-based access control rules applied cleanly in the plan and silently did nothing at runtime. The provider ignores LBAC config on Terraform-provisioned datasources. This post covers what actually works: a three-layer architecture splitting access control across provider aliases, manually-created datasource UIDs, and Alloy write-time label enforcement.

Read
GrafanaTerraformPlatform EngineeringObservability

The Terraform Module That Has No Provider: How a Pure YAML Registry Drives a Multi-Tenant Grafana Platform

We onboard products to two Grafana Cloud stacks — dashboards, alert rules, RBAC, LBAC, contact points, OnCall — by editing one YAML file. No Terraform file changes. The key architectural decision was a provider-free Terraform module that reads the registry and derives everything else. This post covers the pattern, why it works, and where it breaks.

Read
SREObservabilityIncident ResponseSystem Design

An SLO Without an Error-Budget Policy Is Just a Chart

Most teams that 'do SLOs' have a number in a doc and a Grafana panel showing attainment. What they don't have is the thing that makes an SLO a business instrument: a written policy that changes what the team does when the budget runs out. Without the freeze-or-ship consequence, an SLO is a vanity metric with a burn-rate alert nobody has agreed to obey.

Read
ObservabilityAlloyGrafanaPlatform Engineering

Hybrid Observability Scales on the Label Schema, Not the Collector Count

The hard part of hybrid and multi-cloud observability isn't running collectors in every environment — it's making one query resolve identically whether the data came from AKS, an on-prem plant VM, or SAP RISE. That needs a single enforced label schema, set at the collector layer, with no segment optional. Without it, cross-environment queries become per-source OR branches and the cost dashboard becomes an unattributable blob.

Read
ObservabilityOpenTelemetryAlloyGrafana

Learn the Observability Pipeline, Not the Tools — a Map That Survives a Vendor Swap

Beginners drown memorizing Prometheus vs Loki vs Tempo vs Jaeger vs Alloy vs Fluent Bit. The durable model is a six-stage pipeline — instrument, collect, process, store, query, alert — where every tool is a swappable implementation of one stage. When we evaluated four whole observability stacks for the platform I lead, the pipeline shape was identical across all four; only the slots changed.

Read
ObservabilityOpenTelemetryPrometheusGrafana

Monitoring Answers the Questions You Wrote Down. Observability Answers the Ones You Didn't.

A PromQL query grouped by deployment_environment silently returned nothing — no error, no failed scrape, no alert, every dashboard green. Nothing was broken in the way monitoring understands broken; the label had just stopped being promoted. This is the practical line between monitoring (a fixed set of questions frozen at design time) and observability (asking new questions of data you already collected), and what it costs to confuse the two.

Read