All posts

#SRE

16 posts

SREPlatformEngineeringEngineeringPrioritization

Prioritization Frameworks for Engineering: Make Decisions That Actually Move the Needle

In on-call engineering cultures, urgency and importance feel identical—until you're weeks into work no one needed. This post combines the Complexity/Effort Matrix with the Eisenhower Decision Matrix into a five-tier execution system, with a ready-made 4×4 template for your Monday backlog triage.

Read
LLMPlatform EngineeringSystem DesignSRE

Agent Memory as Infrastructure: The Context Injection Pattern That Replaces Fine-Tuning

Built four production AI specialists off a single LLM using plain markdown files and a 20-line fetch call — no fine-tuning, no frameworks. The real insight is that LLM behavioral specialization is a configuration problem, not a training problem, and applying IaC discipline to memory files gives you diffability, rollback, and CI-gated behavioral regression testing that fine-tuning can never provide. At MAANG scale, this extends to prompt caching economics, governance workflows, distributed session state, and multi-model routing.

Read
SREAIOpsLLMKubernetes

Building an SRE Agent: From Playbook to Autonomous Incident Response

What happens when you put an LLM in the incident response loop? We built an SRE Agent on AKS that runs playbooks, queries Grafana, and generates post-mortems. Here's what worked, what didn't, and what we learned about trusting AI in production.

Read
SREObservabilityPrometheusGrafana

Unbounded Cardinality, Zero Alerts, and 14-Second Dashboard Loads: The Structural Limits of the Grafana Azure Monitor Plugin

The Grafana Azure Monitor datasource plugin connects directly to Azure Monitor at query time. On a 12-panel dashboard, that means 12 sequential ARM API calls on every open — 14 seconds before the first chart renders. It cannot feed Mimir alerting rules. It cannot attach environment labels without custom queries per panel. And promoting Azure tags as Prometheus labels explodes cardinality to thousands of series per resource group. This post documents the five structural gaps and the specific design decisions in a push-based exporter that close each one.

Read
SREObservabilityGrafanaAlloy

How Alloy's Default max_shards Turned a Mimir Blip Into a Production Monitoring Blackout

When Grafana Mimir slowed down, Alloy's default retry configuration — 200 parallel shard workers — consumed the majority of CPU on a shared VM, starving the co-located business process and forcing an engineer to kill the observability agent during an active incident. Here's the exact three-layer fix and why shared-VM deployments require explicit resource budgets at every level.

Read
ObservabilityOpenTelemetrySRESystem Design

In FinTech, Your Trace Attributes Are a Compliance Liability

The observability data that makes a payments system debuggable — full request bodies in logs, account numbers on spans, card data in error messages — is the same data that turns your telemetry backend into an unregulated copy of your system of record. Trust in fintech isn't a status page; it's provable reliability plus a guarantee that sensitive data never left the boundary it was supposed to stay inside.

Read
ObservabilityOpenTelemetrySRESystem Design

You Don't Get Root on SAP RISE: Observability Inside a Managed Black Box

RISE with SAP hands you a mission-critical ERP landscape and takes away the thing every observability playbook assumes: OS access to put an agent on the host. The workable model is to monitor the contract, not the internals — synthetic business transactions, the interfaces SAP does expose, and the integration layers you still own. Plan to drop Alloy on the app servers and the project stalls at 'we can't get in.'

Read
ObservabilityAlloyPlatform EngineeringIncident Response

Bridging OT and IT Observability: You Meet at the Historian, Not the PLC

Manufacturing observability tempts you to instrument the plant floor directly — point a collector at the PLCs, scrape everything. That path runs into a safety boundary, an air gap, and a cardinality wall: an unscoped Windows process collector was already emitting 3,900 series per plant mid-rollout. The integration point that actually works is the process historian, and the fragile link is a service account nobody tested.

Read
ObservabilityAlloySREPlatform Engineering

On-Prem Observability Breaks Every Assumption Your Cloud Collector Made

The Alloy config that works flawlessly as an AKS DaemonSet becomes a liability on a plant-floor VM. No elastic compute means a retry storm starves the workload it shares a host with. No managed identity means static token rotation. Egress restrictions mean the Grafana Cloud endpoint isn't reachable the way you assume. On-prem isn't cloud with worse latency — it's a different set of constraints.

Read
SREGrafanaAlloyPlatformEngineeringObservability

The Self-Silencing Anti-Pattern: Why Your Observability Stack Goes Blind When You Need It Most

The system effectively silences its own diagnostics during the moment of maximum operational need.

Read
ObservabilityOpenTelemetrySREPlatform Engineering

Retrofitting Observability Costs 10x — What 'Day One' Actually Means

We cut service onboarding from three days to thirty minutes, but only for services that adopt the template on day one. The three days is where retrofit lives: reverse-engineering what to instrument, adding correlation IDs to a printf codebase, backfilling resource attributes, and finding cardinality bombs in production. Day one is a concrete checklist, not a good intention.

Read
ObservabilityPlatform EngineeringGrafanaSRE

Self-Service Observability Is a Paved Road, Not an Explore Button

Handing developers an Editor role and an empty Explore tab is not self-service — it's abdication. Real self-service is a paved road: a golden-signal dashboard generated from the service name, an SLO template, and label-based access control that scopes a team to its own data. Without the paving you get hundreds of one-off dashboards, cardinality bombs nobody owns, and every team able to read every other team's telemetry.

Read
GrafanaTerraformObservabilityPlatform Engineering

The Grafana Terraform Provider Silently Drops LBAC Rules — and the Three-Layer Fix

When we codified Grafana Cloud multi-tenancy as Terraform, label-based access control rules applied cleanly in the plan and silently did nothing at runtime. The provider ignores LBAC config on Terraform-provisioned datasources. This post covers what actually works: a three-layer architecture splitting access control across provider aliases, manually-created datasource UIDs, and Alloy write-time label enforcement.

Read
SREObservabilityIncident ResponseSystem Design

An SLO Without an Error-Budget Policy Is Just a Chart

Most teams that 'do SLOs' have a number in a doc and a Grafana panel showing attainment. What they don't have is the thing that makes an SLO a business instrument: a written policy that changes what the team does when the budget runs out. Without the freeze-or-ship consequence, an SLO is a vanity metric with a burn-rate alert nobody has agreed to obey.

Read
SREObservabilityIncident ResponseAIOps

Every SRE Practice Is Downstream of Observability — Here's the Dependency Chain

SLOs, error budgets, burn-rate alerting, incident response, blameless postmortems, capacity planning, chaos engineering — each one consumes observability data as input. On a signal layer you can't trust, they degrade to opinion with a dashboard. The Reactive-to-Autonomous journey is really the maturity of that signal layer, and the 80%+ alert-noise cut came from making it trustworthy enough for automation to stand on.

Read
ObservabilityOpenTelemetryPrometheusGrafana

Monitoring Answers the Questions You Wrote Down. Observability Answers the Ones You Didn't.

A PromQL query grouped by deployment_environment silently returned nothing — no error, no failed scrape, no alert, every dashboard green. Nothing was broken in the way monitoring understands broken; the label had just stopped being promoted. This is the practical line between monitoring (a fixed set of questions frozen at design time) and observability (asking new questions of data you already collected), and what it costs to confuse the two.

Read