LLM-backed SRE agent deployed on AKS. Receives alert webhooks from Grafana IRM, queries Grafana/Loki/Kubernetes, executes approved remediation playbooks, and generates draft post-mortems — autonomously, before a human wakes up.
// observability architect
Amit
Singh
14+ years building developer and observability platforms across global financial and enterprise organisations. M.Tech, BITS Pilani. Patent holder. 8+ industry awards.
About
Building platforms that
outlast the incidents
I'm an Observability Architect & Platform Engineer with 14+ years of experience building developer and observability platforms across global financial and enterprise organisations. M.Tech from BITS, Pilani. Currently leading McCain Foods' global observability transformation — a unified telemetry platform spanning Azure, On-Prem, and SAP RISE environments.
My work sits at the intersection of platform engineering, SRE, and applied AI. I think in signal-to-noise ratios, cardinality budgets, and MTTR. I hold a granted patent in the observability and performance engineering space and have been recognised with 8+ industry awards including two org-wide "Outstanding Contribution" honours.
Outside work, I'm preparing for a principal/staff role at MAANG — sharpening algorithmic depth, building this portfolio, and documenting hard-won knowledge through a personal LLM wiki pipeline.
System and Method for Monitoring Performance of Applications for an Entity
An automated global performance reporting system that reduced manual reporting effort from 31 hours → 2 minutes per cycle — saving 6,000+ engineering hours/year and generating $1.2M in annual savings across 20+ global teams. Standardised performance analytics using ASP.NET + SQL Server with automated ingestion pipelines.
Projects
Things I've built
Configurable OpenTelemetry signal generator for validating Alloy pipelines, testing cardinality budgets, and firing alert rules under controlled conditions. Used to validate observability configs before they reach production.
Full internal developer platform on a 5-cluster k3d substrate: Backstage (service catalog + scaffolder), ArgoCD (GitOps), Crossplane (infrastructure provisioning), Kyverno (policy), Linkerd (service mesh), and Grafana Cloud integration. Phase 0 → Phase 1 shipped.
Stack
Tools I trust in production
Observability Platform
Collection & Pipeline
Cloud & Platform
IaC & Automation
Languages
Principles & Practices
All production observability work uses OpenTelemetry-native instrumentation — vendor neutrality, open standards, long-term portability.
GitHub Actions Certified · Microsoft Azure Fundamentals (AZ-900) · M.Tech Software Engineering, BITS Pilani
Career
14 years of building things
SRE & Observability Architect
McCain Foods · Global SRE Team · Remote, India
- › Designed global observability platform (Grafana Cloud + OTel) — real-time visibility across 200+ workloads and multi-region
- › Observability-as-Code via Terraform + GitHub Actions: 100% reproducible, zero drift across all environments
- › Telemetry pipelines for Azure Container Apps, Function Apps, SQL MI, OT workloads — MTTA ↓ 50%, detection ↑ 40%
- › SLO/SLA governance with BigPanda AIOps and ServiceNow — alert noise ↓ 60%, actionable alerts ↑ 4×
- › Extended observability to factory systems and SAP RISE, bridging IT + OT telemetry
- › Mentored 8 SREs across India and Canada; co-led postmortem automation and observability maturity roadmap
Expert Site Reliability Engineer (SRE)
Finastra · Remote, India · [Global Financial Software]
- › Built IaC-based Grafana Cloud observability stack (Terraform, Helm, GitHub Actions) for 100+ microservices and 1,000+ containers — 30+ environments provisioned 95% faster
- › Developed OpenTelemetry tracing toggle utility for Java microservices — dynamic trace control, saving ~10 engineering hours per release
- › Implemented label-based RBAC and reusable monitoring templates across 12 product teams
- › Defined SLO framework and standard alerting taxonomy used org-wide
- › Collaborated with SRE Guild to automate incident taxonomy — postmortem efficiency ↑ 70%
Sr. Performance Engineer
Wolters Kluwer · Pune, India · [Tax & Accounting]
- › Designed ELK-based observability pipeline integrated with testing frameworks — analysis time ↓ 60%
- › Delivered capacity optimisation framework — $250K/year infrastructure cost savings
- › Implemented performance regression automation — Sev-1 incidents ↓ 35%
- › Established baseline SLIs and real-time performance telemetry with QA and DevOps teams
- › Winner — iLab Code Games 2018
Infrastructure Technology Specialist
Cognizant Technology Solutions · Pune, India
- › Built Splunk dashboards for distributed systems — investigation time ↓ 90%
- › Developed .NET-based transaction tracing tool enabling cross-service visibility
- › JVM and GC tuning — throughput ↑ 25%
- › Automated provisioning with Ansible and shell scripts — QA setup time ↓ 50%
- › "Best Newcomer" award — Cognizant Infrastructure Services PACE Practice
Sr. Project Engineer
Wipro Technologies · Pune, India
- › Invented and patented Global Performance Test Reporting Tool — reporting time 31 hours → 2 minutes, saving $1.2M+/year
- › Automated report ingestion and analytics using ASP.NET + SQL Server across 20+ global teams
- › Reduced manual effort from 8,000 → 2,000 hours/year through automation
- › Multiple innovation awards for automation-led delivery excellence
- › "Thanks, A Zillion" award — Citi US Client Appreciation
Speaking
Sharing the hard parts
From Reactive to Resilient — McCain's Observability in Motion
McCain Foods
Presented the global observability transformation narrative to CTO leadership — framing the journey from reactive monitoring to resilient, autonomous reliability across hybrid cloud workloads. Covered key metrics, architecture decisions, and the roadmap to autonomous SRE.
On the radar
- › KubeCon / CNCF — observability at scale
- › SREcon — autonomous incident response with LLMs
- › Grafana ObservabilityCON — Alloy pipeline patterns
Blog
Writing from the trenches
The Productivity Trap Nobody Talks About: When "Productive" Work Becomes Procrastination
Reinstalling Linux, rebuilding dotfiles, reorganizing repos — it all feels like progress, until you notice the thing you actually needed to learn never got touched. A framework for telling optimization work apart from avoidance.
Build Day: Agent Harness & Memory — Key Learnings on Agent Harnesses, Context Engineering, and Memory Systems
Reflections from Build Club & Mem0, Pune (June 2026)
Unbounded Cardinality, Zero Alerts, and 14-Second Dashboard Loads: The Structural Limits of the Grafana Azure Monitor Plugin
The Grafana Azure Monitor datasource plugin connects directly to Azure Monitor at query time. On a 12-panel dashboard, that means 12 sequential ARM API calls on every open — 14 seconds before the first chart renders. It cannot feed Mimir alerting rules. It cannot attach environment labels without custom queries per panel. And promoting Azure tags as Prometheus labels explodes cardinality to thousands of series per resource group. This post documents the five structural gaps and the specific design decisions in a push-based exporter that close each one.
Open Source
Code in the open
Monorepo learning lab — 14 pillars spanning observability, SRE, platform engineering, and autonomous reliability.
Experimental SRE agent for autonomous incident response — LLM-backed triage, playbook execution, and post-mortem generation on AKS.
OpenTelemetry validation lab — configurable signal generator for testing Alloy pipelines, cardinality budgets, and alert rules.
Contact
Let's connect
I'm actively exploring principal-level and staff engineering roles at MAANG and Tier-1 tech companies. If you're building something ambitious in the observability, reliability, or platform engineering space — let's talk.
Open to
Based in India (IST, UTC+5:30) · Open to remote globally