Configurable OpenTelemetry signal generator for validating Alloy pipelines, testing cardinality budgets, and firing alert rules under controlled conditions. Used to validate observability configs before they reach production.
// observability architect
Amit
Singh
14+ years building developer and observability platforms across global financial and enterprise organisations. M.Tech, BITS Pilani. Patent holder. 8+ industry awards.
About
Building platforms that
outlast the incidents
I'm an Observability Architect & Platform Engineer with 14+ years of experience building developer and observability platforms across global financial and enterprise organisations. M.Tech from BITS, Pilani. Currently leading McCain Foods' global observability transformation — a unified telemetry platform spanning Azure, On-Prem, and SAP RISE environments.
My work sits at the intersection of platform engineering, SRE, and applied AI. I think in signal-to-noise ratios, cardinality budgets, and MTTR. I hold a granted patent in the observability and performance engineering space and have been recognised with 8+ industry awards including two org-wide "Outstanding Contribution" honours.
Outside work, I'm preparing for a principal/staff role at MAANG — sharpening algorithmic depth, building this portfolio, and documenting hard-won knowledge through a personal LLM wiki pipeline.
System and Method for Monitoring Performance of Applications for an Entity
An automated global performance reporting system that reduced manual reporting effort from 31 hours → 2 minutes per cycle — saving 6,000+ engineering hours/year and generating $1.2M in annual savings across 20+ global teams. Standardised performance analytics using ASP.NET + SQL Server with automated ingestion pipelines.
Projects
Things I've built
CLI + GitHub Action that reconciles branch protection and repository rulesets against a versioned policy.yml — validate/audit/plan/apply, dual enforcement backends, fail-closed schema validation. Published to PyPI; governs its own repo.
LLM-backed SRE agent deployed on AKS. Receives alert webhooks from Grafana IRM, queries Grafana/Loki/Kubernetes, executes approved remediation playbooks, and generates draft post-mortems — autonomously, before a human wakes up.
Stack
Tools I trust in production
Observability Platform
Collection & Pipeline
Cloud & Platform
IaC & Automation
Languages
Principles & Practices
All production observability work uses OpenTelemetry-native instrumentation — vendor neutrality, open standards, long-term portability.
GitHub Actions Certified · Microsoft Azure Fundamentals (AZ-900) · M.Tech Software Engineering, BITS Pilani
Career
14 years of building things
SRE & Observability Architect
McCain Foods · Global SRE Team · Remote, India
- ›Designed global observability platform (Grafana Cloud + OTel) — real-time visibility across 200+ workloads and multi-region
- ›Observability-as-Code via Terraform + GitHub Actions: 100% reproducible, zero drift across all environments
- ›Telemetry pipelines for Azure Container Apps, Function Apps, SQL MI, OT workloads — MTTA ↓ 50%, detection ↑ 40%
- ›SLO/SLA governance with BigPanda AIOps and ServiceNow — alert noise ↓ 60%, actionable alerts ↑ 4×
- ›Extended observability to factory systems and SAP RISE, bridging IT + OT telemetry
- ›Mentored 8 SREs across India and Canada; co-led postmortem automation and observability maturity roadmap
Expert Site Reliability Engineer (SRE)
Finastra · Remote, India · [Global Financial Software]
- ›Built IaC-based Grafana Cloud observability stack (Terraform, Helm, GitHub Actions) for 100+ microservices and 1,000+ containers — 30+ environments provisioned 95% faster
- ›Developed OpenTelemetry tracing toggle utility for Java microservices — dynamic trace control, saving ~10 engineering hours per release
- ›Implemented label-based RBAC and reusable monitoring templates across 12 product teams
- ›Defined SLO framework and standard alerting taxonomy used org-wide
- ›Collaborated with SRE Guild to automate incident taxonomy — postmortem efficiency ↑ 70%
Sr. Performance Engineer
Wolters Kluwer · Pune, India · [Tax & Accounting]
- ›Designed ELK-based observability pipeline integrated with testing frameworks — analysis time ↓ 60%
- ›Delivered capacity optimisation framework — $250K/year infrastructure cost savings
- ›Implemented performance regression automation — Sev-1 incidents ↓ 35%
- ›Established baseline SLIs and real-time performance telemetry with QA and DevOps teams
- ›Winner — iLab Code Games 2018
Infrastructure Technology Specialist
Cognizant Technology Solutions · Pune, India
- ›Built Splunk dashboards for distributed systems — investigation time ↓ 90%
- ›Developed .NET-based transaction tracing tool enabling cross-service visibility
- ›JVM and GC tuning — throughput ↑ 25%
- ›Automated provisioning with Ansible and shell scripts — QA setup time ↓ 50%
- ›"Best Newcomer" award — Cognizant Infrastructure Services PACE Practice
Sr. Project Engineer
Wipro Technologies · Pune, India
- ›Invented and patented Global Performance Test Reporting Tool — reporting time 31 hours → 2 minutes, saving $1.2M+/year
- ›Automated report ingestion and analytics using ASP.NET + SQL Server across 20+ global teams
- ›Reduced manual effort from 8,000 → 2,000 hours/year through automation
- ›Multiple innovation awards for automation-led delivery excellence
- ›"Thanks, A Zillion" award — Citi US Client Appreciation
Speaking
Sharing the hard parts
From Reactive to Resilient — McCain's Observability in Motion
McCain Foods
Presented the global observability transformation narrative to CTO leadership — framing the journey from reactive monitoring to resilient, autonomous reliability across hybrid cloud workloads. Covered key metrics, architecture decisions, and the roadmap to autonomous SRE.
On the radar
- ›KubeCon / CNCF — observability at scale
- ›SREcon — autonomous incident response with LLMs
- ›Grafana ObservabilityCON — Alloy pipeline patterns
Blog
Writing from the trenches
The Productivity Trap Nobody Talks About: When "Productive" Work Becomes Procrastination
Reinstalling Linux, rebuilding dotfiles, reorganizing repos — it all feels like progress, until you notice the thing you actually needed to learn never got touched. A framework for telling optimization work apart from avoidance.
Prioritization Frameworks for Engineering: Make Decisions That Actually Move the Needle
In on-call engineering cultures, urgency and importance feel identical—until you're weeks into work no one needed. This post combines the Complexity/Effort Matrix with the Eisenhower Decision Matrix into a five-tier execution system, with a ready-made 4×4 template for your Monday backlog triage.
Build Day: Agent Harness & Memory — Key Learnings on Agent Harnesses, Context Engineering, and Memory Systems
Reflections from Build Club & Mem0, Pune (June 2026)
Open Source
Code in the open
Monorepo learning lab — 14 pillars spanning observability, SRE, platform engineering, and autonomous reliability.
Experimental SRE agent for autonomous incident response — LLM-backed triage, playbook execution, and post-mortem generation on AKS.
OpenTelemetry validation lab — configurable signal generator for testing Alloy pipelines, cardinality budgets, and alert rules.
Contact
Let's connect
I'm actively exploring principal-level and staff engineering roles at MAANG and Tier-1 tech companies. If you're building something ambitious in the observability, reliability, or platform engineering space — let's talk.
Open to
Based in India (IST, UTC+5:30) · Open to remote globally