# Agentic Ai Projects And Mastery
All Agentic Ai Projects And Mastery notes →# Observability
All Observability notes →1 — Dashboard Design
The three-question test for a vanity panel, why the same underlying data needs a different dashboard for different audiences, and the top-down layout that mirrors how an investigation actually drills down.
1 — Alerting & Alert Routing
Symptom-based vs. cause-based alerting, the noise-reduction problem (dedup, grouping, correlation), and routing/escalation — the design discipline standing between a real page and a bad one.
2 — SLOs & Error Budgets
SLI/SLO/SLA, the error budget as a spendable resource rather than a compliance number, burn rate as the mechanism connecting the two, and why multi-window multi-burn-rate alerting exists at all.
# Projects
All Projects notes →Runbooks
Troubleshooting playbooks for every known Signal Forge failure mode, from missing traces to Grafana Cloud export errors.
Alert Runbooks
Runbooks invoked directly from paging alerts.
Runbook — ShipSolidApiGateway5xxHigh
- **Service:** api-gateway - **Owner Team:** Platform SRE
Communication Templates
Copy-paste communication templates for incidents.
Incident Response Playbook
The step-by-step incident response flow for platform-impacting incidents.
Incident Trends & Themes
Cross-incident analysis — recurring themes, top contributors, and where to invest.
On-Call Handbook
Everything an on-call engineer needs for a shift on the observability platform.
04 — Operations & Incident Response
Running the platform: on-call, incident response, runbooks, and post-mortems.
2026-05-21 — billing-service elevated latency
- **Incident Commander:** On-call Engineer (Platform SRE)
Post-Mortems
Blameless post-mortems for resolved incidents.
Rotation Schedule
Current on-call rotation and schedule.
Runbook Index
Index of all operational runbooks.
Severity Definitions
Severity definitions so everyone agrees on what SEV-n means.
Incident Notification & Response
**Owner:** SRE Team **Last Updated:** 2025-05-01 **Applies to:** All production alerts routed
Notification & Alerting Strategy
A unified approach to delivering actionable alerts, minimizing noise, and ensuring timely
SRE Toolkit
The `srekit` CLI is a Python-based toolkit for SRE operations, located in
Visualization, Alerting & SLOs
An effective observability system translates raw telemetry into actionable insights through