All posts

#System Design

5 posts

LLMPlatform EngineeringSystem DesignSRE

Agent Memory as Infrastructure: The Context Injection Pattern That Replaces Fine-Tuning

Built four production AI specialists off a single LLM using plain markdown files and a 20-line fetch call — no fine-tuning, no frameworks. The real insight is that LLM behavioral specialization is a configuration problem, not a training problem, and applying IaC discipline to memory files gives you diffability, rollback, and CI-gated behavioral regression testing that fine-tuning can never provide. At MAANG scale, this extends to prompt caching economics, governance workflows, distributed session state, and multi-model routing.

Read
ObservabilityOpenTelemetrySRESystem Design

In FinTech, Your Trace Attributes Are a Compliance Liability

The observability data that makes a payments system debuggable — full request bodies in logs, account numbers on spans, card data in error messages — is the same data that turns your telemetry backend into an unregulated copy of your system of record. Trust in fintech isn't a status page; it's provable reliability plus a guarantee that sensitive data never left the boundary it was supposed to stay inside.

Read
ObservabilityOpenTelemetrySRESystem Design

You Don't Get Root on SAP RISE: Observability Inside a Managed Black Box

RISE with SAP hands you a mission-critical ERP landscape and takes away the thing every observability playbook assumes: OS access to put an agent on the host. The workable model is to monitor the contract, not the internals — synthetic business transactions, the interfaces SAP does expose, and the integration layers you still own. Plan to drop Alloy on the app servers and the project stalls at 'we can't get in.'

Read
SREObservabilityIncident ResponseSystem Design

An SLO Without an Error-Budget Policy Is Just a Chart

Most teams that 'do SLOs' have a number in a doc and a Grafana panel showing attainment. What they don't have is the thing that makes an SLO a business instrument: a written policy that changes what the team does when the budget runs out. Without the freeze-or-ship consequence, an SLO is a vanity metric with a burn-rate alert nobody has agreed to obey.

Read
SREObservabilityIncident ResponseAIOps

Every SRE Practice Is Downstream of Observability — Here's the Dependency Chain

SLOs, error budgets, burn-rate alerting, incident response, blameless postmortems, capacity planning, chaos engineering — each one consumes observability data as input. On a signal layer you can't trust, they degrade to opinion with a dashboard. The Reactive-to-Autonomous journey is really the maturity of that signal layer, and the 80%+ alert-noise cut came from making it trustworthy enough for automation to stand on.

Read