Severity Definitions
Purpose
Severity definitions so everyone agrees on what SEV-n means.
| Severity | Definition | Example | Response |
|---|---|---|---|
| SEV1 | Critical, broad customer/business impact | Total ingest loss; prod down | Page IC immediately, all-hands |
| SEV2 | Significant impact, degraded | Major dashboard/alerting outage | Page on-call |
| SEV3 | Minor / contained | Single non-critical service degraded | Ticket, business hours |
| SEV4 | Negligible / informational | Cosmetic, no user impact | Backlog |
[stub: severity-thresholds-detail]— fill this in. Greppable doc-debt marker.
Related
Local graph
Linked from 4 notes
Runbook — ShipSolidApiGateway5xxHigh
- **Service:** api-gateway - **Owner Team:** Platform SRE
Incident Response Playbook
The step-by-step incident response flow for platform-impacting incidents.
On-Call Handbook
Everything an on-call engineer needs for a shift on the observability platform.
04 — Operations & Incident Response
Running the platform: on-call, incident response, runbooks, and post-mortems.
Related notes
Runbooks
Troubleshooting playbooks for every known Signal Forge failure mode, from missing traces to Grafana Cloud export errors.
Alert Runbooks
Runbooks invoked directly from paging alerts.
Communication Templates
Copy-paste communication templates for incidents.
Incident Response Playbook
The step-by-step incident response flow for platform-impacting incidents.