OpenTelemetry fundamentals
OpenTelemetry (OTel) is the open standard for producing telemetry: an API your code calls,
an SDK that processes the data, a wire protocol (OTLP), and a Collector that
routes it to any backend. Instrument once and send to Grafana, Langfuse, Datadog or all of them. Every
example here runs with the real opentelemetry-sdk, using in-memory exporters so you can see
exactly what gets recorded.
A parcel tracking number
A parcel gets one tracking number (the trace id) at the first depot. Every depot, truck and sorting machine scans it and records what it did and how long it took (a span). When the parcel crosses to a partner courier, the number travels on the label (context propagation), so the whole journey can be reassembled later, whichever company handled each leg.
1. The data model
| Term | Meaning |
|---|---|
| Resource | who is emitting: service.name, version, environment, host. Attached to everything. |
| Tracer / Meter / Logger | the API objects your code uses to create spans, metric instruments and log records |
| Span kind | SERVER (handles an incoming request), CLIENT (makes an outgoing call), INTERNAL, PRODUCER, CONSUMER |
| Processor / Exporter | the SDK pipeline: processors batch or filter; exporters send (OTLP, console, in-memory) |
| Semantic conventions | agreed attribute names (http.request.method, gen_ai.usage.input_tokens), so every backend understands your data |
2. Traces in Python
import time
from opentelemetry import trace
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import SimpleSpanProcessor
from opentelemetry.sdk.trace.export.in_memory_span_exporter import InMemorySpanExporter
from opentelemetry.trace import Status, StatusCode
exporter = InMemorySpanExporter() # in production: OTLPSpanExporter + BatchSpanProcessor
provider = TracerProvider(resource=Resource.create({"service.name": "docs-copilot", "deployment.environment": "dev"}))
provider.add_span_processor(SimpleSpanProcessor(exporter))
trace.set_tracer_provider(provider)
tracer = trace.get_tracer("docs_copilot.rag")
@tracer.start_as_current_span("retrieve") # decorator form
def retrieve(question):
span = trace.get_current_span()
time.sleep(0.02)
span.set_attribute("retrieval.top_k", 4)
span.add_event("hits", {"urls": ["sql/06-order-limit.html#pagination", "sql/27-explain.html"]})
return ["chunk-1", "chunk-2"]
def answer(question):
with tracer.start_as_current_span("POST /ask", kind=trace.SpanKind.SERVER) as root:
root.set_attribute("app.question.length", len(question))
chunks = retrieve(question)
with tracer.start_as_current_span("chat gpt-5.6-luna", kind=trace.SpanKind.CLIENT) as llm:
time.sleep(0.05)
llm.set_attribute("gen_ai.usage.input_tokens", 1_310)
with tracer.start_as_current_span("verify_citations") as check:
try:
raise ValueError("answer cites [3] but only 2 sources were sent")
except ValueError as e:
check.record_exception(e) # an event with the stack trace
check.set_status(Status(StatusCode.ERROR, "ungrounded answer"))
return "…"
answer("How do I paginate a big table?")
spans = exporter.get_finished_spans()
children = {}
for s in spans:
children.setdefault(s.parent.span_id if s.parent else None, []).append(s)
def show(parent_id=None, depth=0):
for s in sorted(children.get(parent_id, []), key=lambda s: s.start_time):
ms = (s.end_time - s.start_time) / 1e6
extra = {k: v for k, v in s.attributes.items()} or ""
print(f"{' ' * depth}{s.name:<{28 - 3 * depth}} {ms:5.0f} ms {s.status.status_code.name:<5} {extra}")
for event in s.events:
print(f"{' ' * depth} · event {event.name}: {str(dict(event.attributes))[:60]}")
show(s.context.span_id, depth + 1)
show()
print("one trace id for all:", {format(s.context.trace_id, "032x") for s in spans} == {format(spans[0].context.trace_id, "032x")})
print("resource:", dict(spans[0].resource.attributes)["service.name"])
Context is implicit
start_as_current_span makes the new span the parent of anything started inside it, including in called functions. You never pass span objects around.
Errors are data
record_exception adds an event with type, message and stack. set_status(ERROR) makes the span count as failed. (Exceptions that escape a with block are recorded automatically.)
Attributes vs events
Attributes describe the whole span and are what you filter and group by. Events are timestamped moments within it.
3. Context propagation across services
When service A calls service B, the trace must continue in B. OTel injects the current context into
the outgoing request's headers as a W3C traceparent, and B extracts it. HTTP
instrumentation libraries do this for you. Here it is by hand, so you can see the header:
from opentelemetry import trace
from opentelemetry.propagate import extract, inject
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import SimpleSpanProcessor
from opentelemetry.sdk.trace.export.in_memory_span_exporter import InMemorySpanExporter
exporter = InMemorySpanExporter()
provider = TracerProvider()
provider.add_span_processor(SimpleSpanProcessor(exporter))
trace.set_tracer_provider(provider)
tracer = trace.get_tracer("demo")
# ---- service A: the agent, calling an MCP server over HTTP
with tracer.start_as_current_span("agent turn", kind=trace.SpanKind.CLIENT):
headers = {"content-type": "application/json"}
inject(headers) # adds traceparent (and tracestate / baggage if set)
print("outgoing headers:", headers)
# ---- service B: the MCP server, receiving the request
incoming_context = extract(headers)
with tracer.start_as_current_span("tools/call search_docs", context=incoming_context,
kind=trace.SpanKind.SERVER):
pass
a, b = exporter.get_finished_spans()
version, trace_id, parent_id, flags = headers["traceparent"].split("-")
print("same trace:", a.context.trace_id == b.context.trace_id)
print("B's parent is A:", b.parent.span_id == a.context.span_id, "| traceparent parent id:", parent_id)
print("trace flags:", flags, "(bit 01 = sampled; bit 02 = the trace id is random, W3C Trace Context level 2)")
The MCP spec reserves traceparent, tracestate and baggage in
params._meta for exactly this. Over stdio there are no HTTP headers, so the context rides in
the JSON-RPC message instead, and host and server spans join one trace.
4. Metrics
Spans describe single requests. Metrics are cheap aggregates you keep for months and alert on. The core instruments are a Counter (only goes up: requests, tokens), an UpDownCounter (in-flight requests), a Histogram (distributions: latency, tokens per request) and Gauges (a current value). Views customise aggregation, such as histogram bucket boundaries:
import random
from opentelemetry.sdk.metrics import MeterProvider
from opentelemetry.sdk.metrics.export import InMemoryMetricReader
from opentelemetry.sdk.metrics.view import ExplicitBucketHistogramAggregation, View
reader = InMemoryMetricReader()
latency_buckets = View(instrument_name="app.request.duration", # seconds
aggregation=ExplicitBucketHistogramAggregation([0.5, 1, 2, 4, 8, 16]))
meter = MeterProvider(metric_readers=[reader], views=[latency_buckets]).get_meter("docs_copilot")
requests = meter.create_counter("app.requests", unit="{request}", description="Answered questions")
tokens = meter.create_counter("app.tokens", unit="{token}")
duration = meter.create_histogram("app.request.duration", unit="s")
random.seed(1)
for _ in range(500):
route = random.choice(["docs", "docs", "docs", "stats"])
requests.add(1, {"route": route})
tokens.add(random.randint(800, 3000), {"type": "input"})
tokens.add(random.randint(50, 400), {"type": "output"})
duration.record(random.lognormvariate(0.6, 0.6), {"route": route})
for rm in reader.get_metrics_data().resource_metrics:
for m in rm.scope_metrics[0].metrics:
for point in m.data.data_points:
labels = dict(point.attributes)
if hasattr(point, "bucket_counts"):
print(f"{m.name} {labels}: count={point.count} sum={point.sum:.0f}s "
f"buckets(≤0.5,1,2,4,8,16,+Inf)={list(point.bucket_counts)}")
else:
print(f"{m.name} {labels}: {point.value:,}")
Every distinct combination of attribute values is a separate time series. route has 2
values, which is fine. user_id or the question text as a metric attribute creates millions
of series and takes down your metrics backend. Put high-cardinality values on spans, and keep
metric attributes to small, fixed sets.
5. Logs that know their trace
Keep using Python's logging, but stamp every record with the current trace and span id.
Then Grafana can jump from a log line to its trace and back. The OTel logging instrumentation does this
automatically. The mechanism is just a filter:
import logging
import sys
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
trace.set_tracer_provider(TracerProvider())
tracer = trace.get_tracer("demo")
class TraceIdFilter(logging.Filter):
def filter(self, record):
ctx = trace.get_current_span().get_span_context()
record.trace_id = format(ctx.trace_id, "032x") if ctx.is_valid else "-"
record.span_id = format(ctx.span_id, "016x") if ctx.is_valid else "-"
return True
handler = logging.StreamHandler(sys.stdout)
handler.addFilter(TraceIdFilter())
handler.setFormatter(logging.Formatter("%(levelname)s trace=%(trace_id)s span=%(span_id)s %(message)s"))
log = logging.getLogger("docs_copilot")
log.addHandler(handler)
log.setLevel(logging.INFO)
log.info("service started") # outside any span
with tracer.start_as_current_span("POST /ask"):
log.info("retrieved 4 chunks")
with tracer.start_as_current_span("chat gpt-5.6-luna"):
log.warning("answer had no valid citation, returning fallback")
6. Sampling
At high traffic you may not want every trace. A head sampler decides when a trace starts.
ParentBased(TraceIdRatioBased(0.1)) keeps 10% of new traces and always follows the parent's
decision, so traces are never half-recorded across services. Tail sampling in the Collector
decides after the trace finishes, so it can keep every error and every slow request.
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import SimpleSpanProcessor
from opentelemetry.sdk.trace.export.in_memory_span_exporter import InMemorySpanExporter
from opentelemetry.sdk.trace.sampling import ParentBased, TraceIdRatioBased
exporter = InMemorySpanExporter()
provider = TracerProvider(sampler=ParentBased(TraceIdRatioBased(0.10)))
provider.add_span_processor(SimpleSpanProcessor(exporter))
tracer = provider.get_tracer("demo")
for i in range(2_000):
with tracer.start_as_current_span("request"):
with tracer.start_as_current_span("llm call"): # follows its parent's decision
pass
spans = exporter.get_finished_spans()
roots = [s for s in spans if s.parent is None]
print(f"kept {len(roots)} of 2000 traces ({len(roots) / 2000:.1%}); spans kept: {len(spans)}")
print("every kept child has its parent kept too:",
all(s.parent.span_id in {r.context.span_id for r in roots} for s in spans if s.parent))
7. Exporting for real: OTLP and the Collector
In production, swap the in-memory exporter for OTLP with a BatchSpanProcessor (it sends in the background, never on the request path), and point it at a Collector. Configuration usually comes from standard environment variables, so the code does not change between environments:
from opentelemetry import metrics, trace
from opentelemetry.exporter.otlp.proto.grpc.metric_exporter import OTLPMetricExporter
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry.sdk.metrics import MeterProvider
from opentelemetry.sdk.metrics.export import PeriodicExportingMetricReader
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
def setup_telemetry(service_name: str) -> None:
resource = Resource.create({"service.name": service_name}) # OTEL_RESOURCE_ATTRIBUTES adds more
tracer_provider = TracerProvider(resource=resource)
tracer_provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter())) # endpoint from env
trace.set_tracer_provider(tracer_provider)
reader = PeriodicExportingMetricReader(OTLPMetricExporter(), export_interval_millis=15_000)
metrics.set_meter_provider(MeterProvider(resource=resource, metric_readers=[reader]))
export OTEL_SERVICE_NAME=docs-copilot export OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4317 # the Collector (gRPC); 4318 for HTTP export OTEL_TRACES_SAMPLER=parentbased_traceidratio export OTEL_TRACES_SAMPLER_ARG=0.25 # zero-code instrumentation for common libraries (FastAPI, httpx, requests, SQLAlchemy …): pip install opentelemetry-distro opentelemetry-exporter-otlp && opentelemetry-bootstrap -a install opentelemetry-instrument python app.py
Batch processors hold spans in memory. Short scripts and serverless functions must call
provider.shutdown() (or force_flush()) before exiting, or the last spans are
lost. Long-running servers flush on shutdown from their lifespan hook.
Recap
- A trace is a tree of spans sharing a trace id. Spans carry attributes, events and status.
- Context propagates implicitly in-process and via
traceparentacross services (and MCP's_meta). - Metrics are cheap aggregates: counters and histograms with low-cardinality attributes.
- Export OTLP to a Collector with batch processors; configure via
OTEL_*variables; sample deliberately; flush on exit.
Checkpoint
user_id as an attribute on the request counter. Why is that dangerous?