Notes / tag / leadership-org

#leadership-org

9 notes

1 — Building an SRE Organization

The team-topology decisions — embedded vs. centralized, how many SREs per service — that determine whether SRE scales with the org or becomes its bottleneck.

sre leadership-org book

2 — Defining Reliability Strategy

Setting reliability targets and investment priorities at an org level, not per-service — the strategy layer above any individual team's SLOs.

sre leadership-org book

3 — Reliability Reviews

The recurring cadence that keeps SLO attainment, error-budget burn, and toil trends visible to leadership before they become a crisis.

sre leadership-org book

4 — Executive Reliability Metrics

Translating error budgets and burn rate into the handful of numbers an executive actually needs to make a reliability-vs-velocity call.

sre leadership-org book

5 — Engineering Culture

Why blameless postmortems and error budgets only work if the surrounding culture actually rewards surfacing problems instead of hiding them.

sre leadership-org book

6 — Hiring SREs

What to actually screen for in an SRE hire — the systems-thinking and incident judgment that don't show up in a standard coding interview.

sre leadership-org book

7 — Mentoring Engineers

Building the next generation of on-call-capable engineers deliberately, instead of letting incident experience be the only teacher.

sre leadership-org book

8 — Technical Leadership

Driving a reliability initiative across teams that don't report to you — the influence-without-authority skill every staff-plus SRE role actually runs on.

sre leadership-org book

9 — Organizational Scaling

How SRE practices that work at 10 services and one team break at 1,000 services and thirty teams, and what has to change structurally to keep up.

sre leadership-org book