Appears in: Telemetry Ingestion Pipeline — this is §7 of the full design, split into its own file so the root stays a table of contents.
7. Component Map (What Exists in the Wild)
| Layer | OSS option | Managed/SaaS option | Your experience |
|---|---|---|---|
| Agent | OTel Collector, Grafana Alloy | Datadog Agent, New Relic | Alloy at ShipSolid (production) |
| Ingestion gateway | OTel Collector (gateway mode) | Grafana Cloud ingest | Alloy gateway mode |
| Buffer | Kafka (Apache), Pulsar, Kinesis | Confluent Cloud, MSK | — |
| Metric processor | OTel Collector processors | Custom Flink/Spark job | Alloy pipelines |
| Trace processor | OTel Collector (tail sampler) | Jaeger, Tempo with tail samp | Tempo at ShipSolid |
| Metric store | Prometheus, Mimir, Thanos, Cortex | Grafana Cloud Mimir | Mimir (production) |
| Log store | Loki, Elasticsearch, ClickHouse | Grafana Cloud Loki | Loki (production) |
| Trace store | Tempo, Jaeger, Zipkin | Grafana Cloud Tempo | Tempo (production) |
Local graph
Related notes
Chapter 1 — Telemetry Ingestion Pipeline
Principal/Staff-level design of a high-throughput telemetry ingestion pipeline — requirements, architecture, deep dives, and trade-offs at 10x scale.
Q1: 500M Samples/Sec, Zero Drop on Rolling Deploy
Full principal-level solution: design a telemetry ingestion pipeline for 500M metric samples/sec from 100K services globally with a zero-drop guarantee during rolling deployment of the ingestion tier.
Q10: Self-Service Tenant Onboarding With Zero Platform-Team Involvement
Full principal-level solution: design a self-service tenant onboarding API for a telemetry pipeline that protects shared infrastructure from a misbehaving new tenant on day one.
Q8: Counters Resetting to Zero After an OTel SDK Upgrade
Full principal-level solution: diagnose and fix a tenant's dashboards showing counters reset to zero every few minutes after an OTel SDK upgrade, without requiring instrumentation changes.