RFC: Clear Descriptive Title
- RFC ID: rfc-YYYY-MM-
- Authors: [Name(s), Role(s)]
- Status: Draft | Proposed | Approved | Implemented | Rejected | Superseded
- Created: YYYY-MM-DD
- Last Updated: YYYY-MM-DD
- Target Release: [Version or Date]
- Supersedes: [Link to old RFC, if any]
- Related Docs: [ADR links (see ADR Template), GitHub issues, diagrams, dashboards, PRDs]
1. Summary
A concise, executive overview of the proposal. State what you’re proposing and why it matters.
Example: This RFC proposes onboarding Grafana Cloud as the unified observability backend for metrics, logs, and traces across Azure, On-Prem, and containerized environments.
2. Background & Motivation
Explain the current situation, gaps, incidents, pain points, or strategic drivers that led to this proposal.
- What problems exist today?
- Why is change needed now?
- Which users or systems are impacted?
3. Goals & Non-Goals
Goals
- What this RFC aims to accomplish
Non-Goals
- What this RFC explicitly does not cover
4. Scope
What environments, systems, teams, or services are included?
- Cloud platforms (e.g., Azure Functions, AKS, Cosmos DB)
- On-Prem (e.g., FactoryTalk, VMs)
- SAP RISE (if relevant)
- Tools involved (e.g., Grafana Cloud, Prometheus, Tempo, Loki, FluentBit)
5. Proposed Solution
5.1 Overview
High-level description of the proposed solution.
5.2 Architecture
- System diagrams (link or embed)
- Data flow (metrics/logs/traces)
- Collector design (agent vs sidecar vs daemonset)
5.3 Instrumentation & Pipelines
- Tools/SDKs to be used
- Manual vs auto-instrumentation
- Exporters and formats (e.g., OTLP, Prometheus Remote Write)
5.4 Dashboards & Alerts
- Which metrics/logs/traces will be visualized?
- Alerting strategy and integrations
6. Security & Compliance
See Security & Compliance for the platform’s general treatment of this topic; fill in service-specific detail below.
- PII/PCI/GxP data handling
- Obfuscation, access control, retention
- Regulatory alignment (GDPR, HIPAA, etc.)
7. Testing & Validation Plan
- How the solution will be validated
- Rollout environments
- Metrics for success/failure
8. Rollout Plan
8.1 Phases
Phased adoption by teams or environments.
8.2 Rollback Plan
Backout or contingency plan if the rollout fails.
9. Success Criteria
Measurable indicators of success:
- Trace coverage (%)
- Dashboard usage stats
- MTTR reduction
- Alert noise reduction
10. Alternatives Considered
- Option 1: <summary + why it was rejected>
- Option 2: <summary + why it was rejected>
- Status quo: <why we’re not keeping things as-is>
11. Risks & Mitigations
- Risk 1: [description] → Mitigation: [strategy]
- Risk 2: …
12. Open Questions
- What needs stakeholder alignment?
- What’s still uncertain?
13. Stakeholders & Reviewers
| Name | Role | Responsibility |
|---|---|---|
| Jane Doe | Platform Lead | Review & Approval |
| John Smith | Observability Champion | Technical Validation |
| Team X | Service Owner | Implementation Feedback |
14. References
- Grafana Cloud Docs
- OpenTelemetry Spec
- Incident Postmortem #123 (see Incident Post-Mortem Template)
- Link to Architecture Diagram
Local graph
Linked from 4 notes
ADR Template
- **Status**: Proposed | Accepted | Rejected | Superseded
Confluence Content Templates — SRE / Observability / Platform
**Space:** `OBS` or `PLT` **Parent page:** `Decision Log (ADRs)`
Incident Post-Mortem Template
- **Incident Commander**: [FILL: name] - **Severity**: SEV1 | SEV2 | SEV3
10 — Templates
Reusable authoring templates for ADRs, RFCs, runbooks, post-mortems, and Confluence pages across the ShipSolid platform.
Related notes
10 — Templates
Reusable authoring templates for ADRs, RFCs, runbooks, post-mortems, and Confluence pages across the ShipSolid platform.
Communication Templates
Copy-paste communication templates for incidents.
ADR Template
- **Status**: Proposed | Accepted | Rejected | Superseded
Confluence Content Templates — SRE / Observability / Platform
**Space:** `OBS` or `PLT` **Parent page:** `Decision Log (ADRs)`