Notes / Observability / 06 Opentelemetry / 1 Opentelemetry Architecture

1 — OpenTelemetry SDKs & Semantic Conventions

OpenTelemetry is a specification and an API/SDK, not a backend — the pieces that make it up, and the semantic-convention vocabulary that lets two unrelated teams' telemetry be queried the same way.

Updated July 17, 2026 · §202607161840 ·

1 — OpenTelemetry SDKs & Semantic Conventions

The single most common confusion about OpenTelemetry is treating it as a backend, the way Prometheus or Jaeger are. It isn’t one. OTel is a vendor-neutral specification for how telemetry gets created, shaped, and exported — the API application code calls, the SDK that implements that API, and a shared attribute vocabulary — with the actual storage and querying left entirely to whatever backend you export to (Prometheus/Mimir for metrics, Tempo/Jaeger for traces, Loki for logs).


The pieces, and where each one runs

Application code


  OTel API        ← what your code calls: tracer.start_span(), meter.create_counter()


  OTel SDK        ← the implementation: samplers, processors, exporters (runs in-process)


  OTLP export


OTel Collector    ← a separate process: receives, transforms, routes (out of scope here)


   Backend        ← Mimir / Tempo / Loki / vendor platform

This chapter is about the top two boxes — what happens inside the application process, before a single byte leaves it. What the Collector does with that data next is a pipeline-design question, covered in OTel Collector Pipeline Design.

API vs. SDK is a deliberate split, not an implementation detail: application code and instrumented libraries depend only on the API. If no SDK is registered, every API call is a documented no-op — a library can call tracer.start_span() unconditionally, and it costs nothing in a process that never configured OpenTelemetry at all. This is why third-party libraries can ship OTel instrumentation built in without forcing every consumer to take on tracing as a hard dependency.


The three signal APIs, and how they map to what you already know

OTel APICore typeMaps to
TracingTracerSpanTrace — one span per unit of work
MetricsMeterCounter / Gauge / Histogram / UpDownCounterMetric — the instrument kind decides how it aggregates
LoggingLoggerLogRecordLog — structured, with the same Resource attached

The instrument kind matters more than it looks: a Counter only ever goes up (requests served, so it composes with a plain sum); a Histogram buckets observations so a percentile can be computed correctly after merging across instances — see 3 — Aggregation Composability — Why You Can't Average Percentiles for exactly why that distinction exists and what breaks if you fake a percentile with a Gauge instead.


Resource: identifying who is talking

Every span, metric, and log point gets stamped with a Resource — a fixed set of attributes identifying the process/host/pod that produced it (service.name, service.version, k8s.pod.name, cloud.region, …), set once at SDK startup and attached to everything that SDK emits. This is the attribute set every dashboard and alert filters or groups by, and it’s the reason service.name is the one label that’s never optional — everything downstream (routing, dashboards, cost attribution) assumes it’s there and correctly set.


Semantic conventions: a shared vocabulary, not a suggestion

Two services, two teams, both instrumenting an HTTP call. Without a shared standard, one emits http_method and the other emits httpVerb, and no dashboard, alert, or vendor tool can query both the same way. Semantic conventions are OTel’s namespaced, versioned specification for exactly this — http.request.method, db.system.name, k8s.pod.name — so that any two conformant instrumentations produce attributes a query or a Grafana panel can rely on by name, regardless of which team or which language wrote the code.

This is also the mechanism that makes Label & Attribute Schema Design tractable: semantic conventions cover the well-known dimensions (HTTP, database, messaging, k8s); that chapter is about the naming discipline for everything semconv doesn’t already define for you — your own business/domain attributes.

Semantic conventions are versioned and evolve — an attribute can move from experimental to stable, get renamed, or get deprecated between spec releases. Pinning a specific semconv version as a platform baseline, and treating an upgrade as a deliberate, reviewed change rather than something that happens silently on the next SDK bump, is a real governance decision — see ADR-006: Pin OpenTelemetry Semantic Conventions to a Platform Baseline for what that looks like as an actual platform decision, not just a specification detail.


Manual vs. automatic: who writes the spans

The SDK gives you the primitives; it doesn’t decide whether a human writes tracer.start_span(...) by hand or whether an auto-instrumentation agent generates it for you at the framework boundary. That trade-off — and the eBPF and service-mesh alternatives that need no SDK in the application at all — is its own chapter: Auto vs. Manual Instrumentation.


Context propagation, briefly

Every span carries a trace_id/span_id pair that has to survive every hop across process boundaries for a trace to hold together at all — the W3C Trace Context (traceparent header) is the wire format OTel uses to carry it. This chapter’s sibling note, 8 — Deadline Propagation, covers the same propagation problem for a different value (a request’s remaining time budget, not its identity); 3 — Cross-Signal Correlation covers why that identity is what makes metrics, logs, and traces usable together at all, rather than three siloed tools.


What this looks like in a real service

SignalForge’s Instrumentation Reference and OTel Signal Contracts document every instrumentation decision for a real (lab) service end to end — Resource attributes, span naming, which attributes are custom vs. semconv-standard.


Why this matters for an Observability Architect

Semantic-convention discipline is what turns “everyone uses OpenTelemetry” into “every team’s telemetry is actually interoperable.” Two teams can both be fully OTel-compliant and still produce data a shared dashboard can’t cleanly query, if one used a custom attribute where a semconv one already existed. Reviewing a new service’s instrumentation for semconv adherence — not just “does it emit spans” — is what keeps the platform’s tooling generic instead of accumulating a per-service dashboard for every service that rolled its own attribute names.

Metadata

DimensionDetail
AuthorAmit Singh
Scopeobservability

Local graph

Full graph →

Linked from 12 notes

4 — Auto vs. Manual Instrumentation

Four ways a span gets created — hand-written, framework-level auto-instrumentation, eBPF, and service-mesh sidecar capture — and the trade-off between code changes and business context each one makes.

9 — OTel Collector Pipeline Design

Receivers, processors, and exporters chained into a pipeline; why a platform runs more than one; and the agent/gateway topology that tail sampling specifically forces on that design.

Chapter 2 — Telemetry Pipelines

OpenTelemetry, OTLP, collectors, sampling, and aggregation as the pipeline that gets a signal from emission to storage without becoming the outage itself.

5 — Label & Attribute Schema Design

Cardinality budget, naming conventions, and the high-churn label traps that turn a cheap metric into a production incident — the design discipline for the labels semantic conventions don't already cover for you.

8 — Deadline Propagation

How a client deadline must be inherited by every downstream goroutine or service call so that cancelled work stops consuming resources rather than running to completion unobserved.

ADR-006: Pin OpenTelemetry Semantic Conventions to v1.26 as Platform Baseline

- **Status**: Proposed - **Date**: 2026-05-07

Grafana Cloud

A book-shaped table of contents for Grafana Cloud: platform foundations through telemetry collection, Mimir/Loki/Tempo/Pyroscope, visualization, application observability, reliability tooling, developer experience, governance, and enterprise reference architectures — cross-linking existing notes instead of duplicating them.

Kubernetes

A book-shaped table of contents for Kubernetes: cloud-native foundations, the CKAD/CKA/CKS certification tracks, control-plane internals, platform tooling, multi-cluster architecture, and MAANG-level system design and interview prep — cross-linking the existing Prometheus, Observability, and Platform Engineering chapters instead of duplicating them.

7 — Distributed Tracing Backend

How spans that arrive out of order, from different services, get assembled into one trace — and the two competing storage models (indexed search vs. object storage plus a trace-ID lookup) that trade query flexibility for cost.

1 — AIOps / Agentic RCA

What's actually new versus a static runbook — an investigation loop, not a fixed trigger-action mapping — why it depends on everything earlier in this book already being solid, and the read-vs-write safety line most real deployments draw.

8 — Case Study: Reactive → Resilient → Autonomous

An illustrative three-act arc — built on the ShipSolid platform maturity model, not any single real deployment — showing why the disciplined middle act is what actually earns the reliability, and why the autonomous act doesn't work without it.

Observability Engineering

A book-shaped table of contents for observability engineering: foundations through architecture, metrics, logging, tracing, profiling, OpenTelemetry, instrumentation, Kubernetes/cloud, data platforms, visualization, alerting, SRE integration, cost, security, platform engineering, AI-driven operations, and MAANG interview preparation — cross-linking existing prometheus/grafana-cloud/kubernetes/sre/platform-engineering notes instead of duplicating them.