2 — Troubleshooting Case Studies
Purpose
[stub: troubleshooting-case-studies]
Metadata
| Author | Amit Singh |
| Scope | observability |
Local graph
Related notes
Observability Engineering
A book-shaped table of contents for observability engineering: foundations through architecture, metrics, logging, tracing, profiling, OpenTelemetry, instrumentation, Kubernetes/cloud, data platforms, visualization, alerting, SRE integration, cost, security, platform engineering, AI-driven operations, and MAANG interview preparation — cross-linking existing prometheus/grafana-cloud/kubernetes/sre/platform-engineering notes instead of duplicating them.
1 — Observability System Design Questions
Covers the recurring system-design prompt shape — 'design a metrics/logging/tracing platform at scale' — and the tradeoffs interviewers probe for.
3 — Telemetry Design Exercises
Covers exercises in designing the telemetry (metrics/logs/traces/labels) for a given service from scratch, a common interview format.
4 — Incident Walkthroughs
Covers narrating a real incident timeline and RCA in interview-answer form, structured for a behavioral or systems-thinking question.