Incident Response Playbook
Purpose
The step-by-step incident response flow for platform-impacting incidents.
Flow
- OBSERVE — confirm scope: what’s broken, who’s impacted, when it started, what changed.
- DECIDE — set severity (Severity Definitions); declare in IRM.
- ACT — stop the bleed with non-destructive steps first; name rollback paths before destructive ones.
- COMMUNICATE — use Communication Templates; update stakeholders.
- RESOLVE — confirm recovery against SLIs.
- LEARN — write a Post-Mortem (do not write it live).
Roles
| Role | Responsibility |
|---|---|
| Incident Commander | Owns the incident, makes calls |
| Comms lead | Stakeholder updates |
| Ops/SME | Hands-on diagnosis & remediation |
Tooling
Grafana IRM (routing/on-call), BigPanda (correlation), SNOW (tickets), Grafana Explore (metrics/logs/traces).
Related
Local graph
Linked from 7 notes
Runbook — ShipSolidApiGateway5xxHigh
- **Service:** api-gateway - **Owner Team:** Platform SRE
Incident Notification & Response
**Owner:** SRE Team **Last Updated:** 2025-05-01 **Applies to:** All production alerts routed
Notification & Alerting Strategy
A unified approach to delivering actionable alerts, minimizing noise, and ensuring timely
On-Call Handbook
Everything an on-call engineer needs for a shift on the observability platform.
04 — Operations & Incident Response
Running the platform: on-call, incident response, runbooks, and post-mortems.
Severity Definitions
Severity definitions so everyone agrees on what SEV-n means.
Rollback Procedures
How to roll back each component safely.
Related notes
Runbooks
Troubleshooting playbooks for every known Signal Forge failure mode, from missing traces to Grafana Cloud export errors.
Alert Runbooks
Runbooks invoked directly from paging alerts.
Communication Templates
Copy-paste communication templates for incidents.
Incident Trends & Themes
Cross-incident analysis — recurring themes, top contributors, and where to invest.