Module 4 · Observability & operations

OpenTelemetry fundamentals

Advanced 24 min read Traces, metrics, context propagation, export, in Python

OpenTelemetry (OTel) is the open standard for producing telemetry: an API your code calls, an SDK that processes the data, a wire protocol (OTLP), and a Collector that routes it to any backend. Instrument once and send to Grafana, Langfuse, Datadog or all of them. Every example here runs with the real opentelemetry-sdk, using in-memory exporters so you can see exactly what gets recorded.

📦

A parcel tracking number

A parcel gets one tracking number (the trace id) at the first depot. Every depot, truck and sorting machine scans it and records what it did and how long it took (a span). When the parcel crosses to a partner courier, the number travels on the label (context propagation), so the whole journey can be reassembled later, whichever company handled each leg.

1. The data model

trace_id 4bf92f3577b34da6a3ce929d0e0e4736 (shared by every span) POST /ask (SERVER span, root) retrieve embed query chat gpt-5.6-luna (CLIENT span) verify 0 ms2,400 ms every span: name · span_id · parent_span_id · start/end · kind · status (OK/ERROR) attributes (key → value) · events (timestamped notes, exceptions) · links
TermMeaning
Resourcewho is emitting: service.name, version, environment, host. Attached to everything.
Tracer / Meter / Loggerthe API objects your code uses to create spans, metric instruments and log records
Span kindSERVER (handles an incoming request), CLIENT (makes an outgoing call), INTERNAL, PRODUCER, CONSUMER
Processor / Exporterthe SDK pipeline: processors batch or filter; exporters send (OTLP, console, in-memory)
Semantic conventionsagreed attribute names (http.request.method, gen_ai.usage.input_tokens), so every backend understands your data

2. Traces in Python

import time

from opentelemetry import trace
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import SimpleSpanProcessor
from opentelemetry.sdk.trace.export.in_memory_span_exporter import InMemorySpanExporter
from opentelemetry.trace import Status, StatusCode

exporter = InMemorySpanExporter()                        # in production: OTLPSpanExporter + BatchSpanProcessor
provider = TracerProvider(resource=Resource.create({"service.name": "docs-copilot", "deployment.environment": "dev"}))
provider.add_span_processor(SimpleSpanProcessor(exporter))
trace.set_tracer_provider(provider)
tracer = trace.get_tracer("docs_copilot.rag")

@tracer.start_as_current_span("retrieve")                # decorator form
def retrieve(question):
    span = trace.get_current_span()
    time.sleep(0.02)
    span.set_attribute("retrieval.top_k", 4)
    span.add_event("hits", {"urls": ["sql/06-order-limit.html#pagination", "sql/27-explain.html"]})
    return ["chunk-1", "chunk-2"]

def answer(question):
    with tracer.start_as_current_span("POST /ask", kind=trace.SpanKind.SERVER) as root:
        root.set_attribute("app.question.length", len(question))
        chunks = retrieve(question)
        with tracer.start_as_current_span("chat gpt-5.6-luna", kind=trace.SpanKind.CLIENT) as llm:
            time.sleep(0.05)
            llm.set_attribute("gen_ai.usage.input_tokens", 1_310)
        with tracer.start_as_current_span("verify_citations") as check:
            try:
                raise ValueError("answer cites [3] but only 2 sources were sent")
            except ValueError as e:
                check.record_exception(e)                    # an event with the stack trace
                check.set_status(Status(StatusCode.ERROR, "ungrounded answer"))
        return "…"

answer("How do I paginate a big table?")

spans = exporter.get_finished_spans()
children = {}
for s in spans:
    children.setdefault(s.parent.span_id if s.parent else None, []).append(s)

def show(parent_id=None, depth=0):
    for s in sorted(children.get(parent_id, []), key=lambda s: s.start_time):
        ms = (s.end_time - s.start_time) / 1e6
        extra = {k: v for k, v in s.attributes.items()} or ""
        print(f"{'   ' * depth}{s.name:<{28 - 3 * depth}} {ms:5.0f} ms  {s.status.status_code.name:<5} {extra}")
        for event in s.events:
            print(f"{'   ' * depth}   · event {event.name}: {str(dict(event.attributes))[:60]}")
        show(s.context.span_id, depth + 1)

show()
print("one trace id for all:", {format(s.context.trace_id, "032x") for s in spans} == {format(spans[0].context.trace_id, "032x")})
print("resource:", dict(spans[0].resource.attributes)["service.name"])
POST /ask 74 ms UNSET {'app.question.length': 30} retrieve 20 ms UNSET {'retrieval.top_k': 4} · event hits: {'urls': ('sql/06-order-limit.html#pagination', 'sql/27-expl chat gpt-5.6-luna 50 ms UNSET {'gen_ai.usage.input_tokens': 1310} verify_citations 2 ms ERROR · event exception: {'exception.type': 'ValueError', 'exception.message': 'answe one trace id for all: True resource: docs-copilot

Context is implicit

start_as_current_span makes the new span the parent of anything started inside it, including in called functions. You never pass span objects around.

Errors are data

record_exception adds an event with type, message and stack. set_status(ERROR) makes the span count as failed. (Exceptions that escape a with block are recorded automatically.)

Attributes vs events

Attributes describe the whole span and are what you filter and group by. Events are timestamped moments within it.

3. Context propagation across services

When service A calls service B, the trace must continue in B. OTel injects the current context into the outgoing request's headers as a W3C traceparent, and B extracts it. HTTP instrumentation libraries do this for you. Here it is by hand, so you can see the header:

from opentelemetry import trace
from opentelemetry.propagate import extract, inject
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import SimpleSpanProcessor
from opentelemetry.sdk.trace.export.in_memory_span_exporter import InMemorySpanExporter

exporter = InMemorySpanExporter()
provider = TracerProvider()
provider.add_span_processor(SimpleSpanProcessor(exporter))
trace.set_tracer_provider(provider)
tracer = trace.get_tracer("demo")

# ---- service A: the agent, calling an MCP server over HTTP
with tracer.start_as_current_span("agent turn", kind=trace.SpanKind.CLIENT):
    headers = {"content-type": "application/json"}
    inject(headers)                               # adds traceparent (and tracestate / baggage if set)
    print("outgoing headers:", headers)

# ---- service B: the MCP server, receiving the request
incoming_context = extract(headers)
with tracer.start_as_current_span("tools/call search_docs", context=incoming_context,
                                  kind=trace.SpanKind.SERVER):
    pass

a, b = exporter.get_finished_spans()
version, trace_id, parent_id, flags = headers["traceparent"].split("-")
print("same trace:", a.context.trace_id == b.context.trace_id)
print("B's parent is A:", b.parent.span_id == a.context.span_id, "| traceparent parent id:", parent_id)
print("trace flags:", flags, "(bit 01 = sampled; bit 02 = the trace id is random, W3C Trace Context level 2)")
outgoing headers: {'content-type': 'application/json', 'traceparent': '00-7fdc05c8d8374754a7ecabea43ec9160-0e8a0b708a4e2459-03'} same trace: True B's parent is A: True | traceparent parent id: 0e8a0b708a4e2459 trace flags: 03 (bit 01 = sampled; bit 02 = the trace id is random, W3C Trace Context level 2)

The MCP spec reserves traceparent, tracestate and baggage in params._meta for exactly this. Over stdio there are no HTTP headers, so the context rides in the JSON-RPC message instead, and host and server spans join one trace.

4. Metrics

Spans describe single requests. Metrics are cheap aggregates you keep for months and alert on. The core instruments are a Counter (only goes up: requests, tokens), an UpDownCounter (in-flight requests), a Histogram (distributions: latency, tokens per request) and Gauges (a current value). Views customise aggregation, such as histogram bucket boundaries:

import random

from opentelemetry.sdk.metrics import MeterProvider
from opentelemetry.sdk.metrics.export import InMemoryMetricReader
from opentelemetry.sdk.metrics.view import ExplicitBucketHistogramAggregation, View

reader = InMemoryMetricReader()
latency_buckets = View(instrument_name="app.request.duration",       # seconds
                       aggregation=ExplicitBucketHistogramAggregation([0.5, 1, 2, 4, 8, 16]))
meter = MeterProvider(metric_readers=[reader], views=[latency_buckets]).get_meter("docs_copilot")

requests = meter.create_counter("app.requests", unit="{request}", description="Answered questions")
tokens = meter.create_counter("app.tokens", unit="{token}")
duration = meter.create_histogram("app.request.duration", unit="s")

random.seed(1)
for _ in range(500):
    route = random.choice(["docs", "docs", "docs", "stats"])
    requests.add(1, {"route": route})
    tokens.add(random.randint(800, 3000), {"type": "input"})
    tokens.add(random.randint(50, 400), {"type": "output"})
    duration.record(random.lognormvariate(0.6, 0.6), {"route": route})

for rm in reader.get_metrics_data().resource_metrics:
    for m in rm.scope_metrics[0].metrics:
        for point in m.data.data_points:
            labels = dict(point.attributes)
            if hasattr(point, "bucket_counts"):
                print(f"{m.name} {labels}: count={point.count} sum={point.sum:.0f}s "
                      f"buckets(≤0.5,1,2,4,8,16,+Inf)={list(point.bucket_counts)}")
            else:
                print(f"{m.name} {labels}: {point.value:,}")
app.requests {'route': 'docs'}: 378 app.requests {'route': 'stats'}: 122 app.tokens {'type': 'input'}: 942,898 app.tokens {'type': 'output'}: 116,053 app.request.duration {'route': 'docs'}: count=378 sum=819s buckets(≤0.5,1,2,4,8,16,+Inf)=[5, 64, 136, 140, 31, 2, 0] app.request.duration {'route': 'stats'}: count=122 sum=237s buckets(≤0.5,1,2,4,8,16,+Inf)=[1, 28, 48, 34, 11, 0, 0]
Cardinality

Every distinct combination of attribute values is a separate time series. route has 2 values, which is fine. user_id or the question text as a metric attribute creates millions of series and takes down your metrics backend. Put high-cardinality values on spans, and keep metric attributes to small, fixed sets.

5. Logs that know their trace

Keep using Python's logging, but stamp every record with the current trace and span id. Then Grafana can jump from a log line to its trace and back. The OTel logging instrumentation does this automatically. The mechanism is just a filter:

import logging
import sys

from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider

trace.set_tracer_provider(TracerProvider())
tracer = trace.get_tracer("demo")

class TraceIdFilter(logging.Filter):
    def filter(self, record):
        ctx = trace.get_current_span().get_span_context()
        record.trace_id = format(ctx.trace_id, "032x") if ctx.is_valid else "-"
        record.span_id = format(ctx.span_id, "016x") if ctx.is_valid else "-"
        return True

handler = logging.StreamHandler(sys.stdout)
handler.addFilter(TraceIdFilter())
handler.setFormatter(logging.Formatter("%(levelname)s trace=%(trace_id)s span=%(span_id)s %(message)s"))
log = logging.getLogger("docs_copilot")
log.addHandler(handler)
log.setLevel(logging.INFO)

log.info("service started")                                    # outside any span
with tracer.start_as_current_span("POST /ask"):
    log.info("retrieved 4 chunks")
    with tracer.start_as_current_span("chat gpt-5.6-luna"):
        log.warning("answer had no valid citation, returning fallback")
INFO trace=- span=- service started INFO trace=aa5f78e0efff84fe0477212f54ab4ec8 span=25991b768e2c2b46 retrieved 4 chunks WARNING trace=aa5f78e0efff84fe0477212f54ab4ec8 span=54d1d76dd0b95422 answer had no valid citation, returning fallback

6. Sampling

At high traffic you may not want every trace. A head sampler decides when a trace starts. ParentBased(TraceIdRatioBased(0.1)) keeps 10% of new traces and always follows the parent's decision, so traces are never half-recorded across services. Tail sampling in the Collector decides after the trace finishes, so it can keep every error and every slow request.

from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import SimpleSpanProcessor
from opentelemetry.sdk.trace.export.in_memory_span_exporter import InMemorySpanExporter
from opentelemetry.sdk.trace.sampling import ParentBased, TraceIdRatioBased

exporter = InMemorySpanExporter()
provider = TracerProvider(sampler=ParentBased(TraceIdRatioBased(0.10)))
provider.add_span_processor(SimpleSpanProcessor(exporter))
tracer = provider.get_tracer("demo")

for i in range(2_000):
    with tracer.start_as_current_span("request"):
        with tracer.start_as_current_span("llm call"):        # follows its parent's decision
            pass

spans = exporter.get_finished_spans()
roots = [s for s in spans if s.parent is None]
print(f"kept {len(roots)} of 2000 traces ({len(roots) / 2000:.1%}); spans kept: {len(spans)}")
print("every kept child has its parent kept too:",
      all(s.parent.span_id in {r.context.span_id for r in roots} for s in spans if s.parent))
kept 200 of 2000 traces (10.0%); spans kept: 400 every kept child has its parent kept too: True

7. Exporting for real: OTLP and the Collector

In production, swap the in-memory exporter for OTLP with a BatchSpanProcessor (it sends in the background, never on the request path), and point it at a Collector. Configuration usually comes from standard environment variables, so the code does not change between environments:

from opentelemetry import metrics, trace
from opentelemetry.exporter.otlp.proto.grpc.metric_exporter import OTLPMetricExporter
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry.sdk.metrics import MeterProvider
from opentelemetry.sdk.metrics.export import PeriodicExportingMetricReader
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor

def setup_telemetry(service_name: str) -> None:
    resource = Resource.create({"service.name": service_name})          # OTEL_RESOURCE_ATTRIBUTES adds more
    tracer_provider = TracerProvider(resource=resource)
    tracer_provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter()))    # endpoint from env
    trace.set_tracer_provider(tracer_provider)
    reader = PeriodicExportingMetricReader(OTLPMetricExporter(), export_interval_millis=15_000)
    metrics.set_meter_provider(MeterProvider(resource=resource, metric_readers=[reader]))
export OTEL_SERVICE_NAME=docs-copilot
export OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4317       # the Collector (gRPC); 4318 for HTTP
export OTEL_TRACES_SAMPLER=parentbased_traceidratio
export OTEL_TRACES_SAMPLER_ARG=0.25
# zero-code instrumentation for common libraries (FastAPI, httpx, requests, SQLAlchemy …):
pip install opentelemetry-distro opentelemetry-exporter-otlp && opentelemetry-bootstrap -a install
opentelemetry-instrument python app.py
Flush before you exit

Batch processors hold spans in memory. Short scripts and serverless functions must call provider.shutdown() (or force_flush()) before exiting, or the last spans are lost. Long-running servers flush on shutdown from their lifespan hook.

Recap

  • A trace is a tree of spans sharing a trace id. Spans carry attributes, events and status.
  • Context propagates implicitly in-process and via traceparent across services (and MCP's _meta).
  • Metrics are cheap aggregates: counters and histograms with low-cardinality attributes.
  • Export OTLP to a Collector with batch processors; configure via OTEL_* variables; sample deliberately; flush on exit.

Checkpoint

1 · You add user_id as an attribute on the request counter. Why is that dangerous?
Metric backends store one series per attribute combination. Spans are designed for high-cardinality detail.
2 · Traces from the API and the MCP server show up as two separate traces. What is missing?
Without the incoming context, the server starts a new root span with a new trace id.
3 · A nightly batch script's last spans never reach the backend. Fix?
BatchSpanProcessor exports in the background. Process exit loses whatever is still buffered unless you flush.