Notes / Projects / Platform Shipsolid / 04 Operations Incident Response

Incident Response Playbook

The step-by-step incident response flow for platform-impacting incidents.

Updated June 9, 2026 · §202606092046-29 ·

Incident Response Playbook

Purpose

The step-by-step incident response flow for platform-impacting incidents.

Flow

  1. OBSERVE — confirm scope: what’s broken, who’s impacted, when it started, what changed.
  2. DECIDE — set severity (Severity Definitions); declare in IRM.
  3. ACT — stop the bleed with non-destructive steps first; name rollback paths before destructive ones.
  4. COMMUNICATE — use Communication Templates; update stakeholders.
  5. RESOLVE — confirm recovery against SLIs.
  6. LEARN — write a Post-Mortem (do not write it live).

Roles

RoleResponsibility
Incident CommanderOwns the incident, makes calls
Comms leadStakeholder updates
Ops/SMEHands-on diagnosis & remediation

Tooling

Grafana IRM (routing/on-call), BigPanda (correlation), SNOW (tickets), Grafana Explore (metrics/logs/traces).

Local graph

Full graph →