On-Call Handbook
Purpose
Everything an on-call engineer needs for a shift on the observability platform.
Before your shift
- Access to Grafana Cloud, IRM, BigPanda, SNOW confirmed.
- Paging device tested.
- Reviewed open incidents and recent changes.
During an incident
Follow the Incident Response Playbook. Default to non-destructive diagnosis first (read metrics/logs/traces) before any restart/rollback/config change. If a destructive action is needed, name the rollback path first.
Escalation
[stub: oncall-escalation-path]— fill this in. Greppable doc-debt marker.
Handoff
[stub: oncall-handoff-template]— fill this in. Greppable doc-debt marker.
Related
Local graph
Linked from 4 notes
Incident Response Playbook
The step-by-step incident response flow for platform-impacting incidents.
Engagement Model
How teams engage the platform team — intake, support tiers, and SLAs.
04 — Operations & Incident Response
Running the platform: on-call, incident response, runbooks, and post-mortems.
Rotation Schedule
Current on-call rotation and schedule.
Related notes
Runbooks
Troubleshooting playbooks for every known Signal Forge failure mode, from missing traces to Grafana Cloud export errors.
Alert Runbooks
Runbooks invoked directly from paging alerts.
Communication Templates
Copy-paste communication templates for incidents.
Incident Response Playbook
The step-by-step incident response flow for platform-impacting incidents.