Appears in: Telemetry Ingestion Pipeline §2 (high-level architecture — ingestion frontier).
The diagram in the parent doc shows an ingestion layer where different telemetry protocols enter the observability platform. Each gateway is responsible for a different ingestion protocol. The OTLP Gateway is just one of them.
Here’s what each gateway does.
| Gateway | Accepts | Telemetry Type | Typical Clients |
|---|---|---|---|
| OTLP Gateway | OTLP/gRPC, OTLP/HTTP | Metrics, Logs, Traces | OpenTelemetry SDKs, Grafana Alloy, OTel Collector |
| Prometheus Remote Write (RW) Gateway | Prometheus Remote Write | Metrics | Prometheus, VictoriaMetrics Agent, Grafana Alloy |
| Syslog Gateway | Syslog (UDP/TCP/TLS) | Logs | Linux servers, routers, firewalls, switches |
| HTTP Log Gateway | HTTP/REST | Logs | Fluent Bit, Vector, custom applications |
| Kafka Gateway | Kafka protocol | Logs, Metrics, Events | Kafka producers |
| StatsD Gateway | UDP StatsD | Metrics | Legacy applications |
| Jaeger Gateway | Jaeger gRPC/Thrift | Traces | Jaeger clients |
| Zipkin Gateway | Zipkin HTTP | Traces | Zipkin clients |
| Fluentd/Fluent Bit Gateway | Fluent Forward protocol | Logs | Fluentd, Fluent Bit |
| OpenMetrics Gateway | HTTP scrape | Metrics | Prometheus exporters |
1. OTLP Gateway
Purpose: universal receiver for OpenTelemetry — metrics, logs, and traces over a single protocol.
Ports:
4317 gRPC
4318 HTTP
flowchart LR
A["Java Application\n(OTel SDK)"] -->|OTLP| B["OTLP Gateway"]
This is the industry standard for new instrumentation — prefer it over every protocol below unless a producer can’t emit it.
2. Prometheus Remote Write Gateway
Prometheus is a pull-based monitoring system. After scraping, it (or an agent) pushes the scraped series onward via remote-write:
flowchart LR
N["Node Exporter"] -->|scrape| P["Prometheus"]
P -->|Remote Write| GW["RW Gateway"]
The gateway accepts the Prometheus Remote Write protocol.
Typical responsibilities:
- Authentication
- Rate limiting
- Compression
- Tenant routing
- Forward to Mimir
Used by:
- Prometheus
- Grafana Alloy
- Prometheus Agent
- VictoriaMetrics Agent
3. Syslog Gateway
Designed for traditional infrastructure. Receives logs from:
- Linux / Unix
- Routers, switches, firewalls
- Load balancers
flowchart LR
C["Cisco Switch"] -->|Syslog| GW["Syslog Gateway"]
GW --> L["Loki"]
Usually supports UDP, TCP, and TLS transport.
4. HTTP Log Gateway
Some applications simply POST logs.
flowchart LR
A["Application"] -->|HTTP POST| GW["HTTP Gateway"]
GW --> L["Log Backend"]
Common clients: Fluent Bit, Vector, custom agents. Useful when Syslog isn’t available.
5. Kafka Gateway
Many enterprises already use Kafka as their transport backbone.
flowchart LR
A["Applications"] -->|Kafka| GW["Kafka Gateway"]
GW --> S["Loki / Tempo / Mimir"]
Advantages:
- Buffering
- Replay
- High throughput
- Decoupling producers and consumers
Common in very large deployments — the trade-off is operational cost: a consumer group, schema registry, and offset-management story to run alongside it.
6. StatsD Gateway
Legacy metrics protocol, still common for older Java, Python, and Ruby applications.
flowchart LR
A["Application"] -->|StatsD UDP| GW["StatsD Gateway"]
GW --> P["Prometheus"]
The gateway converts StatsD metrics into Prometheus/OpenTelemetry format.
7. Jaeger Gateway
Before OpenTelemetry, Jaeger was a popular tracing system.
flowchart LR
A["Application"] -->|Jaeger| GW["Jaeger Gateway"]
GW --> T["Tempo"]
Purpose: receive Jaeger traces and convert them if needed.
8. Zipkin Gateway
Another legacy tracing protocol — many Spring Boot applications historically emitted Zipkin traces.
flowchart LR
A["Application"] -->|Zipkin| GW["Zipkin Gateway"]
GW --> T["Tempo"]
9. Fluent Bit / Fluentd Gateway
Log collectors often use the Fluent Forward protocol — common in Kubernetes.
flowchart LR
A["Fluent Bit"] -->|Forward| GW["Gateway"]
GW --> L["Loki"]
10. OpenMetrics Gateway
Some exporters expose metrics directly over HTTP scrape.
flowchart LR
E["Exporter"] -->|HTTP| GW["OpenMetrics Gateway"]
This converts scraped metrics into the internal format for storage.
Why have multiple gateways?
An enterprise observability platform needs to support many telemetry producers, not just OpenTelemetry.
flowchart TD
APPS["Applications"]
APPS --> OTLP["OTLP"]
APPS --> PRW["Prom Remote Write"]
APPS --> SYS["Syslog"]
APPS --> KAF["Kafka"]
OTLP --> GW1["OTLP GW"]
PRW --> GW2["RW GW"]
SYS --> GW3["Syslog GW"]
KAF --> GW4["Kafka GW"]
GW1 --> ING["Ingestion Layer\n(Auth · Rate Limiting · Schema Validation ·\nTenant Routing · Metadata Enrichment)"]
GW2 --> ING
GW3 --> ING
GW4 --> ING
ING --> STORE[("Observability Storage\n(Mimir, Loki, Tempo, etc.)")]
In a Grafana Cloud deployment
For the Azure environment you described (Azure Container Apps, Azure Functions, AKS), you would typically use:
- OTLP Gateway for modern applications instrumented with OpenTelemetry.
- Prometheus Remote Write Gateway for Prometheus metrics collected from Kubernetes clusters.
- Syslog/HTTP Gateway only if you have network devices, Linux hosts, or legacy systems producing syslog or HTTP-based logs.
- StatsD, Jaeger, and Zipkin gateways only during migrations from older observability stacks. For new deployments, OTLP is generally preferred because it provides a single protocol for metrics, logs, and traces.
Local graph
Linked from 6 notes
2. High-Level Architecture
The producers → ingestion gateway → Kafka → processors → storage diagram for the telemetry ingestion pipeline, plus the push-over-pull key insight to state early in the interview.
What is Fluent Bit
CNCF-graduated, C-written log/metrics/trace forwarder — Fluentd's lightweight sibling, the de facto node-level log-collection DaemonSet in most Kubernetes clusters, and Grafana Alloy's main incumbent competitor for that slot.
What is Jaeger
CNCF-graduated distributed tracing system built at Uber in 2015, Dapper-lineage like Zipkin before it — and, since Jaeger v2, rebuilt on top of the OpenTelemetry Collector rather than bespoke ingestion code.
What is StatsD
Etsy's 2011 UDP-based metrics protocol and daemon — the simplest possible fire-and-forget instrumentation format, superseded as a client API by OTel/Prometheus but still alive everywhere as a compatibility ingestion shim.
Protocol Termination at the Ingestion Frontier
What actually happens where the wire protocol ends — TCP/TLS handoff, HTTP/2 frame demux, gRPC message decode, protobuf deserialization — and why L4 vs L7 termination and connection-lifecycle tuning are the load-bearing decisions here, not the crypto itself.
System Design
Principal/Staff-level system design reference collection for MAANG interview preparation — observability pipelines, distributed systems, reliability engineering, and beyond.
Related notes
Chapter 1 — Telemetry Ingestion Pipeline
Principal/Staff-level design of a high-throughput telemetry ingestion pipeline — requirements, architecture, deep dives, and trade-offs at 10x scale.
Q1: 500M Samples/Sec, Zero Drop on Rolling Deploy
Full principal-level solution: design a telemetry ingestion pipeline for 500M metric samples/sec from 100K services globally with a zero-drop guarantee during rolling deployment of the ingestion tier.
Q10: Self-Service Tenant Onboarding With Zero Platform-Team Involvement
Full principal-level solution: design a self-service tenant onboarding API for a telemetry pipeline that protects shared infrastructure from a misbehaving new tenant on day one.
Q8: Counters Resetting to Zero After an OTel SDK Upgrade
Full principal-level solution: diagnose and fix a tenant's dashboards showing counters reset to zero every few minutes after an OTel SDK upgrade, without requiring instrumentation changes.