// observability architect

Amit
Singh

14+ years building developer and observability platforms across global financial and enterprise organisations. M.Tech, BITS Pilani. Patent holder. 8+ industry awards.

14+ Years
90% MTTD reduction
70% MTTR reduction
$1.5M+/yr Cost savings
grafana-cloud ❯ mccain-global
$ query --scope=career --metric=impact
onboarding_time      3d → 30 min ↓96%
mttd_reduction       90%
mttr_reduction       70%
cost_savings         $1.5M+/year
engineers_trained    50+
patent.status        IN + US ✓
awards              8+ ✓
// Reactive → Resilient → Autonomous
scroll

Building platforms that
outlast the incidents

I'm an Observability Architect & Platform Engineer with 14+ years of experience building developer and observability platforms across global financial and enterprise organisations. M.Tech from BITS, Pilani. Currently leading McCain Foods' global observability transformation — a unified telemetry platform spanning Azure, On-Prem, and SAP RISE environments.

My work sits at the intersection of platform engineering, SRE, and applied AI. I think in signal-to-noise ratios, cardinality budgets, and MTTR. I hold a granted patent in the observability and performance engineering space and have been recognised with 8+ industry awards including two org-wide "Outstanding Contribution" honours.

Outside work, I'm preparing for a principal/staff role at MAANG — sharpening algorithmic depth, building this portfolio, and documenting hard-won knowledge through a personal LLM wiki pipeline.

M.Tech, Software Engineering — BITS, Pilani (2015) GitHub Actions Certified Azure AZ-900 Certified
14+
Years in platform & observability
1
Patent (India + US)
90%
MTTD reduction
70%
MTTR reduction
$1.5M+
Annualized savings
50+
Engineers mentored
Patent Holder Granted · India + US

System and Method for Monitoring Performance of Applications for an Entity

An automated global performance reporting system that reduced manual reporting effort from 31 hours → 2 minutes per cycle — saving 6,000+ engineering hours/year and generating $1.2M in annual savings across 20+ global teams. Standardised performance analytics using ASP.NET + SQL Server with automated ingestion pipelines.

India Patent No.
3327/CHE/2015
US Patent No.
KNS.BFSI.0353IN1
Annual savings
$1.2M+

Things I've built

LLM-backed SRE agent deployed on AKS. Receives alert webhooks from Grafana IRM, queries Grafana/Loki/Kubernetes, executes approved remediation playbooks, and generates draft post-mortems — autonomously, before a human wakes up.

PythonFastAPIClaude APIAKSKubernetes

Configurable OpenTelemetry signal generator for validating Alloy pipelines, testing cardinality budgets, and firing alert rules under controlled conditions. Used to validate observability configs before they reach production.

OpenTelemetryAlloyPrometheusKubernetesPython

Full internal developer platform on a 5-cluster k3d substrate: Backstage (service catalog + scaffolder), ArgoCD (GitOps), Crossplane (infrastructure provisioning), Kyverno (policy), Linkerd (service mesh), and Grafana Cloud integration. Phase 0 → Phase 1 shipped.

BackstageArgoCDCrossplaneKyvernoLinkerd

Tools I trust in production

Observability Platform

Grafana Cloud Prometheus Loki Tempo Mimir Grafana IRM AppDynamics Splunk

Collection & Pipeline

OpenTelemetry Grafana Alloy BigPanda AIOps ServiceNow ELK Stack

Cloud & Platform

Azure AKS Azure ARO Azure Container Apps Function Apps SQL MI Logic Apps SAP RISE

IaC & Automation

Terraform Ansible Helm FluxCD Docker Kubernetes GitHub Actions Azure Pipelines

Languages

Python Bash C# Java Go (learning) YAML ASP.NET SQL

Principles & Practices

Platform Engineering SRE GitOps Observability Strategy SLO/SLA Design FinOps AIOps
CNCF

All production observability work uses OpenTelemetry-native instrumentation — vendor neutrality, open standards, long-term portability.

Certs

GitHub Actions Certified · Microsoft Azure Fundamentals (AZ-900) · M.Tech Software Engineering, BITS Pilani

14 years of building things

Dec 2024 — Present

SRE & Observability Architect

McCain Foods · Global SRE Team · Remote, India

Current
  • Designed global observability platform (Grafana Cloud + OTel) — real-time visibility across 200+ workloads and multi-region
  • Observability-as-Code via Terraform + GitHub Actions: 100% reproducible, zero drift across all environments
  • Telemetry pipelines for Azure Container Apps, Function Apps, SQL MI, OT workloads — MTTA ↓ 50%, detection ↑ 40%
  • SLO/SLA governance with BigPanda AIOps and ServiceNow — alert noise ↓ 60%, actionable alerts ↑ 4×
  • Extended observability to factory systems and SAP RISE, bridging IT + OT telemetry
  • Mentored 8 SREs across India and Canada; co-led postmortem automation and observability maturity roadmap
✦ 100% coverage for critical services in 6 months · 6 product teams (200+ users) · Foundation for "Unified Telemetry by Design" 2026
Grafana CloudOpenTelemetryAKSTerraformBigPandaServiceNowSAP RISE
Mar 2021 — Dec 2024

Expert Site Reliability Engineer (SRE)

Finastra · Remote, India · [Global Financial Software]

  • Built IaC-based Grafana Cloud observability stack (Terraform, Helm, GitHub Actions) for 100+ microservices and 1,000+ containers — 30+ environments provisioned 95% faster
  • Developed OpenTelemetry tracing toggle utility for Java microservices — dynamic trace control, saving ~10 engineering hours per release
  • Implemented label-based RBAC and reusable monitoring templates across 12 product teams
  • Defined SLO framework and standard alerting taxonomy used org-wide
  • Collaborated with SRE Guild to automate incident taxonomy — postmortem efficiency ↑ 70%
Grafana CloudOpenTelemetryTerraformHelmGitHub ActionsJavaKubernetes
Feb 2018 — Mar 2021

Sr. Performance Engineer

Wolters Kluwer · Pune, India · [Tax & Accounting]

  • Designed ELK-based observability pipeline integrated with testing frameworks — analysis time ↓ 60%
  • Delivered capacity optimisation framework — $250K/year infrastructure cost savings
  • Implemented performance regression automation — Sev-1 incidents ↓ 35%
  • Established baseline SLIs and real-time performance telemetry with QA and DevOps teams
  • Winner — iLab Code Games 2018
ELK StackPerformance EngineeringSLI/SLOAutomation
May 2016 — Feb 2018

Infrastructure Technology Specialist

Cognizant Technology Solutions · Pune, India

  • Built Splunk dashboards for distributed systems — investigation time ↓ 90%
  • Developed .NET-based transaction tracing tool enabling cross-service visibility
  • JVM and GC tuning — throughput ↑ 25%
  • Automated provisioning with Ansible and shell scripts — QA setup time ↓ 50%
  • "Best Newcomer" award — Cognizant Infrastructure Services PACE Practice
SplunkAnsible.NETJVMDistributed Systems
Sep 2011 — May 2016

Sr. Project Engineer

Wipro Technologies · Pune, India

  • Invented and patented Global Performance Test Reporting Tool — reporting time 31 hours → 2 minutes, saving $1.2M+/year
  • Automated report ingestion and analytics using ASP.NET + SQL Server across 20+ global teams
  • Reduced manual effort from 8,000 → 2,000 hours/year through automation
  • Multiple innovation awards for automation-led delivery excellence
  • "Thanks, A Zillion" award — Citi US Client Appreciation
ASP.NETSQL ServerPerformance TestingAutomationInnovation

Sharing the hard parts

CTO Presentation March 2026

From Reactive to Resilient — McCain's Observability in Motion

McCain Foods

Presented the global observability transformation narrative to CTO leadership — framing the journey from reactive monitoring to resilient, autonomous reliability across hybrid cloud workloads. Covered key metrics, architecture decisions, and the roadmap to autonomous SRE.

ObservabilitySREPlatform EngineeringAIOps

On the radar

  • KubeCon / CNCF — observability at scale
  • SREcon — autonomous incident response with LLMs
  • Grafana ObservabilityCON — Alloy pipeline patterns

Let's connect

I'm actively exploring principal-level and staff engineering roles at MAANG and Tier-1 tech companies. If you're building something ambitious in the observability, reliability, or platform engineering space — let's talk.

Open to

Principal / Staff Engineer roles
Observability Architect positions
MAANG & Tier-1 tech companies
Conference talks & advisory
Open-source collaboration

Based in India (IST, UTC+5:30) · Open to remote globally