Notes / Projects / Platform Shipsolid / 08 Strategy Planning

Platform & Cloud Maturity Model (L1 → L5)

This model defines maturity progression across core platform and cloud architecture pillars.

Updated May 1, 2026 · §202605011919 ·

Platform & Cloud Maturity Model (L1 → L5)

Overview

This model defines maturity progression across core platform and cloud architecture pillars. It is designed for practical adoption in Azure-centric environments with evolving observability and SRE practices.

LevelNameDescription
L1Ad HocReactive, fragmented, tool-driven
L2DefinedBasic standards, partial adoption
L3StandardizedOrganization-wide consistency
L4MeasuredSLO-driven, data-informed
L5OptimizedAutonomous, predictive systems

Governance

  • L1: No standards, teams operate independently
  • L2: Basic policies (naming, tagging), inconsistent enforcement
  • L3: Central governance model, enforced via pipelines
  • L4: Policy-as-code (Azure Policy/OPA), audit + drift detection
  • L5: Continuous compliance, auto-remediation

Security

  • L1: Static credentials, minimal controls
  • L2: RBAC, secrets stored in Key Vault
  • L3: Managed identities, network isolation
  • L4: Integrated security (Defender, SIEM), vulnerability management
  • L5: Zero-trust, automated threat response

Platform (Infrastructure + Runtime)

  • L1: Manual provisioning, snowflake environments
  • L2: Infrastructure as Code introduced
  • L3: Standardized templates (ACA, Functions, DBs)
  • L4: Self-service platform (golden paths)
  • L5: Fully abstracted, policy-driven provisioning

Services

  • L1: Monolithic systems
  • L2: Basic service separation
  • L3: Clear service boundaries, API contracts
  • L4: Resilience patterns (retry, circuit breaker)
  • L5: Adaptive, self-optimizing services

Delivery (CI/CD)

  • L1: Manual deployments
  • L2: Basic pipelines
  • L3: Standardized CI/CD workflows
  • L4: Progressive delivery (canary, blue/green)
  • L5: Autonomous delivery (policy + metrics gating)

Observability

  • L1: Logs only
  • L2: Metrics + logs, basic dashboards
  • L3: Unified observability (metrics, logs, traces)
  • L4: SLO-based alerts, actionable insights
  • L5: Context-aware observability feeding AIops

Reliability (SRE)

  • L1: Reactive firefighting
  • L2: Basic SLAs
  • L3: SLOs and error budgets
  • L4: Error budgets influence releases
  • L5: Chaos engineering, auto-recovery

AIops

  • L1: Alert noise
  • L2: Alert correlation
  • L3: Anomaly detection
  • L4: Root cause suggestions
  • L5: Autonomous remediation

Developer Experience (DX)

  • L1: High friction, tribal knowledge
  • L2: Basic documentation
  • L3: Templates and onboarding guides
  • L4: Developer portal, self-service
  • L5: AI-assisted workflows

FinOps

  • L1: No visibility
  • L2: Basic tracking
  • L3: Cost allocation and dashboards
  • L4: Optimization practices enforced
  • L5: Automated cost optimization

Documentation

  • L1: Scattered
  • L2: Centralized repository
  • L3: Structured (runbooks, standards)
  • L4: Versioned and system-linked
  • L5: AI-queryable knowledge base

Labs / Innovation

  • L1: No experimentation
  • L2: Ad hoc POCs
  • L3: Structured experimentation
  • L4: Roadmap-aligned innovation
  • L5: Continuous innovation pipeline

Focus on achieving L3 maturity across core pillars:

  • Platform
  • Delivery
  • Observability
  • Reliability

This enables:

  • Consistent system behavior
  • Scalable operations
  • Foundation for AIops

Notes

  • Progression should be incremental; avoid skipping levels
  • Observability (L3+) is a prerequisite for meaningful SRE and AIops
  • Security and governance should evolve in parallel, not as afterthoughts

Scoring Model (0–5 per pillar)

Each pillar is scored on an integer scale of 0–5. The scale extends the L1 → L5 rubric above with a 0 for “absent / not started”, which the level-based rubric does not cover.

ScoreAnchorDefinition
0AbsentCapability does not exist; no ownership, no artefacts, no signal
1Ad HocReactive, fragmented, tool-driven (per L1)
2DefinedBasic standards in writing, partial adoption (per L2)
3StandardizedOrg-wide consistency, enforced via templates / pipelines / policy (per L3)
4MeasuredSLO-driven, metric-informed, drift detected automatically (per L4)
5OptimizedAutonomous, predictive, self-healing or self-correcting (per L5)

Scoring rules

  • Integer only. No half-points — round down on partial coverage to keep gap analysis honest.
  • Score the weakest link. A pillar is at the level of its least-mature workload, not its best one. If 90% of services are L3 and 10% are L1, the pillar score is 1 — those laggards are where the next investment goes.
  • Evidence-based. Every score above 0 must point to a concrete artefact: a policy file, a dashboard, a runbook, an SLO definition, a CI gate, an incident retro that exercised the capability. If you can’t link to evidence, the score is one level lower.
  • Re-score on cadence. Pillars age. A score recorded six months ago is a hypothesis, not a fact. Default cadence: quarterly for active pillars, semi-annually for steady-state ones.

Assessment table template

Use this table to capture a point-in-time snapshot. Keep a dated copy per assessment — the delta over time is the actual signal.

PillarCurrentTargetGapOwnerEvidenceNotes
Governance
Security
Platform
Services
Delivery
Observability
Reliability
AIops
DX
FinOps
Documentation
Labs

Gap = Target − Current. Any pillar with Gap ≥ 2 is a candidate for the next quarter’s investment plan.


Radar Chart Structure

The 12 pillar scores form a natural radar / spider plot — one axis per pillar, two overlays (current and target) showing the shape of the gap.

Axes and dataset shape

  • Axes (12): Governance, Security, Platform, Services, Delivery, Observability, Reliability, AIops, DX, FinOps, Documentation, Labs
  • Range: 0–5 per axis (must match scoring scale exactly)
  • Datasets: at minimum two — current and target. Optional third — peer-benchmark (where available) or previous-quarter for trendlines.

Portable data block

Capture scores in this structured form alongside the assessment table — it’s what feeds any chart renderer (mermaid, Grafana, matplotlib, Excel).

maturity_assessment:
  date: 2026-05-01
  scope: lab          # lab | platform | both
  axes:
    - Governance
    - Security
    - Platform
    - Services
    - Delivery
    - Observability
    - Reliability
    - AIops
    - DX
    - FinOps
    - Documentation
    - Labs
  datasets:
    current:  [0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0]
    target:   [3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3]

Array order must match axes order — the chart breaks silently otherwise.

Mermaid example

Mermaid radar charts (radar-beta) are supported in Mermaid v11+. MkDocs Material renders this when the mermaid2 plugin is enabled. Beta syntax may shift; treat the YAML block above as the canonical source.

radar-beta
  title Platform & Cloud Maturity — Current vs Target
  axis G["Governance"], S["Security"], P["Platform"], Sv["Services"], D["Delivery"], O["Observability"], R["Reliability"], A["AIops"], DX["DX"], F["FinOps"], Doc["Docs"], L["Labs"]
  curve current["Current"]{0,0,0,0,0,0,0,0,0,0,0,0}
  curve target["Target"]{3,3,3,3,3,3,3,3,3,3,3,3}

  max 5
  min 0

Reading the chart

  • Symmetrical shape, low score — practice is consistent but immature; invest broadly.
  • Spiky shape — one or two pillars are far ahead of the rest; either over-invested there or under-invested everywhere else. Observability spiking above Reliability and AIops is a common pattern and usually means the data is collected but not yet acted on.
  • Target line inside current line on any axis — over-invested. Rare, but worth flagging; usually a sign of vendor lock-in or sunk-cost gold-plating.
  • Inverted L-shape (Governance + Security low, everything else high) — fragile. Compliance debt will pull the rest of the chart down on first audit or incident.

Local graph

Full graph →

Linked from 8 notes

8 — Case Study: Reactive → Resilient → Autonomous

An illustrative three-act arc — built on the ShipSolid platform maturity model, not any single real deployment — showing why the disciplined middle act is what actually earns the reliability, and why the autonomous act doesn't work without it.

2 — Driving Adoption

A paved road nobody travels on didn't help anyone — onboarding time as the leading indicator, self-service as the mechanism that actually moves it, and why migrating an existing service is a harder adoption problem than a greenfield one.

AIOps Overview

The **h-aiops** pillar is an experimental sandbox for the in-house SRE Agent and related AIOps

Future-Readiness & Extensibility

The observability framework is architected with scalability, flexibility, and longevity in mind.

6 — IDP Capability Maturity Model

A capability maturity model for scoring an IDP's coverage across self-service, golden paths, catalog, and DevEx.

3 — Platform Maturity Models

Covers the crawl/walk/run maturity stages and continuous evolution of a platform engineering practice.

Platform Engineering Fundamentals

A book-shaped table of contents for platform engineering fundamentals: evolution and organizational foundations, platform-as-a-product thinking, core principles (self-service, golden paths, automation, APIs), design principles, the platform lifecycle, DORA/SPACE metrics, anti-patterns, enterprise governance, and MAANG interview prep — cross-linking existing sre/patterns/observability/projects notes instead of duplicating them.

08 — Strategy & Planning

Direction: roadmap, OKRs, and the initiative portfolio.