Notes / System Design / 15 Complete Case Studies / 01 Telemetry Ingestion Pipeline

Tenant Identification and Routing at the Ingestion Frontier

How the gateway decides whose data this is — cert/API-key-derived tenant identity, propagation as a Kafka header and X-Scope-OrgID, and the shared-topic-with-filter vs per-tenant-topic vs shuffle-sharded-pool routing trade-off.

Appears in: Telemetry Ingestion Pipeline §3.1 (ingestion frontier — tenant identification and routing).

The reason this is its own design question, not a one-liner: every downstream decision in the pipeline — which cardinality budget applies, which Kafka partition/topic a message lands on, which X-Scope-OrgID gets written to storage — depends on the platform having correctly answered “whose data is this?” before the message leaves the gateway. Get tenant identification wrong and every other isolation mechanism in §3.6 (Multi-Tenancy) is enforcing the wrong boundary.


Where the tenant ID actually comes from

There are four common mechanisms for identifying the tenant on an inbound request. They are not interchangeable — the trust properties differ sharply:

MechanismHow it worksTrust level
mTLS certificate SAN/CNGateway extracts tenant ID from the client cert’s Subject Alternative Name, verified by the TLS handshake itselfStrongest — cannot be forged without the private key
API key → tenant mappingGateway looks up the presented API key in an identity store; the store, not the key itself, says which tenant it belongs toStrong — as good as key issuance/rotation hygiene
JWT claim (tenant_id in a signed token)Gateway verifies the token signature, reads the claimStrong if signature verification is enforced; a common bug is trusting the claim without checking the signature covers it
Path/subdomain/header (/v1/{tenant}, X-Tenant-ID)Extracted directly from the request, no cryptographic bindingUntrusted on its own — trivially spoofable by anything that can reach the gateway

The rule that matters most in an interview: tenant identity must be derived from the authenticated credential (the cert or the key-to-tenant lookup), never read directly from a client-controlled field in the path, header, or payload body. A path-based or header-based tenant hint is fine as a routing convenience (e.g., load-balancer-level sharding before termination), but it must be cross-checked against the authenticated identity and the request rejected on mismatch — this is exactly the “defense 1” fix in Q11 (compromised agent threat model), which covers the spoofing failure mode in depth. This note assumes that defense is already in place and focuses on the identification-and-routing mechanics themselves.

flowchart TD
    A["Request arrives"] --> B["TLS handshake / API key lookup"]
    B --> C["Tenant ID resolved from authenticated identity"]
    C --> D{"Request also carries a\ntenant hint (path/header)?"}
    D -->|"Matches authenticated tenant"| E["Accept — proceed to routing"]
    D -->|"Mismatch"| F["Reject — telemetry_gateway_tenant_id_mismatch_total"]
    D -->|"No hint present"| E

Propagating tenant identity downstream

Once resolved, the tenant ID has to survive every hop without being re-derived (re-deriving it at each layer duplicates trust logic and multiplies the places a bug can leak isolation):

Gateway:    tenant ID resolved from cert/API-key → attached as a message header, not re-embedded in the body
Kafka:      tenant ID carried as a Kafka message header (or encoded in the partition/topic choice — see below)
Processor:  reads the header, not the payload, to apply per-tenant cardinality budget (§3.3) and enrichment
Storage:    tenant ID becomes the X-Scope-OrgID header on the Mimir/Loki write — this is what actually
            partitions data at rest

A common mistake worth naming explicitly: relabeling the tenant ID into a Prometheus label (tenant="acme") instead of using the storage layer’s dedicated tenant header. Labels are just another series dimension — putting the tenant in a label multiplies series cardinality by tenant count and makes cross-tenant queries structurally possible (a permissions bug away from a data leak). X-Scope-OrgID partitions tenants at the storage/query-routing level, entirely separate from the label space.


Routing patterns: how tenant ID becomes a physical routing decision

Knowing the tenant doesn’t by itself decide isolation — that’s a separate choice about how much physical separation the data gets as it moves through the buffer and into storage:

| Pattern | Isolation | Operational cost | When to use | | ------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------ | ----------------------------------------------- | | Shared topic + tenant_id field + consumer filter | Weakest — one noisy tenant can still cause topic-level lag for everyone | Lowest — one topic, no per-tenant broker-side bookkeeping | Default for the long tail of tenants with unremarkable volume | | Per-tenant Kafka topic | Strong at the buffer layer — one tenant’s backlog doesn’t touch another’s topic | High — each topic is a directory on broker disk; doesn’t scale to thousands of tenants | Contractual hard-isolation requirements, or a small number of very large tenants | | Shuffle-sharded downstream pool (ingesters/compactors) | Bounded — a tenant is pinned to a subset of the pool (e.g., 24 of 500 ingesters), so a spike inflates memory on only that subset | Moderate — requires a shard-assignment scheme, but no per-tenant infra to provision | The general answer at MAANG scale — see Q2 (cardinality storm) for the full mechanism |

Answer, stated directly: shared topic with a tenant_id header for routing/filtering at the buffer layer, combined with shuffle-sharding at the storage tier (ingesters, compactors). Per-tenant topics are reserved for the small set of tenants where a topic-per-tenant is contractually required — they don’t scale as a default because Kafka topic count is itself an operational ceiling (partition and file-handle overhead per topic).


Regional / data-residency routing

Tenant identity also drives a coarser routing decision made before the request ever reaches a gateway: which region does this tenant’s traffic land in at all. This composes with the regional deployment topology in §3.8 — a tenant with an EU-data-residency requirement is pinned to the EU region’s gateway endpoint (via DNS routing, or agent-side config pointing only at that region), and that pinning has to be enforced at the edge, not left as an assumption the agent’s config file happens to satisfy. Losing that pinning silently (e.g., a DNS failover routing an EU tenant’s agents to the US region during an outage) is a compliance incident, not just an availability one — worth calling out explicitly if data residency comes up in the interview.


Failure mode: missing or unresolvable tenant ID

Per the Layer 1 responsibilities in §3.1 (“schema validation and rejection — fail fast before the buffer”), a request that fails to resolve to a tenant should be rejected at the gateway, not passed downstream with a placeholder or default tenant. Letting an unidentified request reach Kafka means the processor or storage layer inherits a decision the gateway was in the best position to make immediately, with full context on the auth failure reason (expired cert, unknown API key, revoked credential) that gets lost by the time the message is a few hops downstream.


Local graph

Full graph →

Linked from 8 notes

7 — Multi-Tenancy

Two separate guarantees hiding under one name — data isolation and performance fairness — and the tenant identification, quota enforcement, and selective backpressure that make both hold under shared infrastructure.

3.1 Layer 1: Ingestion Frontier

Layer 1 of the telemetry ingestion pipeline: the ingestion frontier — responsibilities, fan-in at 100K+ agents, protocol negotiation, batching, backpressure, and rate limiting.

3.6 Multi-Tenancy

Multi-tenancy isolation layers and quota enforcement points across the telemetry ingestion pipeline, from network inbound to the storage write path.

Authentication at the Ingestion Frontier: mTLS, Bearer Tokens, API Keys

How the gateway proves an agent's credential is valid — mTLS handshake validation, JWT bearer token signature checks, and API key lookups — with the revocation-speed vs operational-complexity trade-off between them, and credential rotation at 10M-agent scale.

Q11: Redesigning the Ingestion Frontier for a Compromised-Agent Threat Model

Full principal-level solution: redesign the telemetry ingestion frontier assuming a compromised agent sends malformed and adversarial payloads — oversized batches, spoofed tenant IDs, garbage label values.

Q2: Cardinality Storm — Detect and Mitigate Without Affecting Other Tenants

Full principal-level solution: a tenant sends 50M unique label combinations/minute causing TSDB compaction storms — design detection and mitigation that isolates the blast radius to that tenant.

Rate Limiting Architecture: Token Bucket, Gossip, and Envoy Global Limits

Three ways to enforce rate limits across a replicated gateway fleet — centralized Redis token bucket, decentralized gossip-based estimation, and Envoy's sidecar-plus-global-service pattern — with the precision/SPOF/latency trade-offs between them.

Protocol Termination at the Ingestion Frontier

What actually happens where the wire protocol ends — TCP/TLS handoff, HTTP/2 frame demux, gRPC message decode, protobuf deserialization — and why L4 vs L7 termination and connection-lifecycle tuning are the load-bearing decisions here, not the crypto itself.