Module 4 · Observability & operations

Tracing LLM calls with OpenTelemetry

Advanced 18 min read GenAI semantic conventions, by hand and automatic

OpenTelemetry has semantic conventions for generative AI: agreed span names, attribute names and metrics for model calls, retrieval, tool execution and agents. Follow them and any backend (Grafana, Langfuse, Datadog, Phoenix) understands your traces without custom mapping. This lesson instruments a model call by hand to learn exactly what the conventions ask for, traces a whole RAG request, and then shows the libraries that do it automatically.

🏷️

Standard shipping labels

Couriers can scan any parcel because the label layout is standard: the postcode is always in the same box. Semantic conventions are the label layout for telemetry. When every tool writes gen_ai.usage.input_tokens, every dashboard can add it up.

1. The conventions you will use most

SpanNameKey attributes
Model call (kind CLIENT){operation} {model}, e.g. chat gpt-5.6-lunagen_ai.operation.name (chat, embeddings…), gen_ai.provider.name (openai), gen_ai.request.model, gen_ai.response.model, gen_ai.response.id, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, gen_ai.response.finish_reasons, error.type
Retrievalretrieval {data_source}gen_ai.operation.name = retrieval, gen_ai.data_source.id
Tool executionexecute_tool {tool}gen_ai.tool.name, gen_ai.tool.call.id
Agentinvoke_agent {agent}gen_ai.agent.name, gen_ai.conversation.id
Content (opt-in only)—gen_ai.input.messages, gen_ai.output.messages, gen_ai.system_instructions
MetricInstrumentNotes
gen_ai.client.token.usagehistogram, unit {token}one record per call per gen_ai.token.type (input, output); buckets 1, 4, 16 … 67,108,864
gen_ai.client.operation.durationhistogram, unit sbuckets 0.01, 0.02, 0.04 … 81.92
gen_ai.client.operation.time_to_first_chunkhistogram, unit sfor streaming: the TTFT of lesson 03
Still in development

The GenAI conventions (now maintained in their own semantic-conventions-genai repository) are marked Development, and names have changed before: gen_ai.system became gen_ai.provider.name. Keep your instrumentation in one module so a rename is a one-line change, and check the current spec when you upgrade.

2. Instrument a model call by hand

A wrapper around responses.create that emits a convention-compliant span and both standard metrics, and captures message content only when explicitly enabled. It runs through the real SDK against the course's fake server, including a rate-limit failure:

import json
import os
import time

import openai
from opentelemetry import metrics, trace
from opentelemetry.sdk.metrics import MeterProvider
from opentelemetry.sdk.metrics.export import InMemoryMetricReader
from opentelemetry.sdk.metrics.view import ExplicitBucketHistogramAggregation, View
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import SimpleSpanProcessor
from opentelemetry.sdk.trace.export.in_memory_span_exporter import InMemorySpanExporter
from opentelemetry.trace import SpanKind
from fake_openai import fake_client

spans = InMemorySpanExporter()
tp = TracerProvider()
tp.add_span_processor(SimpleSpanProcessor(spans))
reader = InMemoryMetricReader()
mp = MeterProvider(metric_readers=[reader], views=[
    View(instrument_name="gen_ai.client.token.usage", aggregation=ExplicitBucketHistogramAggregation(
        [1, 4, 16, 64, 256, 1024, 4096, 16384, 65536, 262144, 1048576, 4194304, 16777216, 67108864])),
    View(instrument_name="gen_ai.client.operation.duration", aggregation=ExplicitBucketHistogramAggregation(
        [0.01, 0.02, 0.04, 0.08, 0.16, 0.32, 0.64, 1.28, 2.56, 5.12, 10.24, 20.48, 40.96, 81.92]))])
tracer, meter = tp.get_tracer("genai.demo"), mp.get_meter("genai.demo")
token_usage = meter.create_histogram("gen_ai.client.token.usage", unit="{token}")
duration = meter.create_histogram("gen_ai.client.operation.duration", unit="s")

CAPTURE = os.environ.get("OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT", "").lower() in (
    "true", "span_only", "span_and_event")                                      # opt-in, off by default

def traced_response(client, model, input, instructions=None, **kwargs):
    attrs = {"gen_ai.operation.name": "chat", "gen_ai.provider.name": "openai", "gen_ai.request.model": model}
    start = time.perf_counter()
    with tracer.start_as_current_span(f"chat {model}", kind=SpanKind.CLIENT, attributes=attrs) as span:
        if CAPTURE:
            span.set_attribute("gen_ai.input.messages", json.dumps([{"role": "user", "content": input}]))
        try:
            r = client.responses.create(model=model, input=input, instructions=instructions, **kwargs)
        except openai.APIError as e:
            error = type(e).__name__
            span.set_attribute("error.type", error)                   # the exception itself is recorded
            duration.record(time.perf_counter() - start, {**attrs, "error.type": error})
            raise
        finish = "stop" if r.status == "completed" else "length"
        span.set_attributes({"gen_ai.response.model": r.model, "gen_ai.response.id": r.id,
                             "gen_ai.usage.input_tokens": r.usage.input_tokens,
                             "gen_ai.usage.output_tokens": r.usage.output_tokens,
                             "gen_ai.response.finish_reasons": [finish]})
        labels = {**attrs, "gen_ai.response.model": r.model}
        token_usage.record(r.usage.input_tokens, {**labels, "gen_ai.token.type": "input"})
        token_usage.record(r.usage.output_tokens, {**labels, "gen_ai.token.type": "output"})
        duration.record(time.perf_counter() - start, labels)
        return r

client = fake_client(script=["warm-up", {"text": "Use keyset pagination.", "usage": (1310, 42)},
                             {"status": 429}], max_retries=0)
client.responses.create(model="gpt-5.6-luna", input="hi")   # first call pays one-time client setup
traced_response(client, "gpt-5.6-luna", "Why is OFFSET slow for deep pages?")
try:
    traced_response(client, "gpt-5.6-luna", "Second question")
except openai.RateLimitError:
    pass

for s in spans.get_finished_spans():
    print(f"{s.name} [{s.kind.name}] status={s.status.status_code.name}")
    for k, v in s.attributes.items():
        print(f"    {k} = {v}")
for m in reader.get_metrics_data().resource_metrics[0].scope_metrics[0].metrics:
    for p in m.data.data_points:
        tag = p.attributes.get("gen_ai.token.type") or p.attributes.get("error.type") or "ok"
        print(f"{m.name} ({tag}): count={p.count} sum={p.sum:.3g}")
chat gpt-5.6-luna [CLIENT] status=UNSET gen_ai.operation.name = chat gen_ai.provider.name = openai gen_ai.request.model = gpt-5.6-luna gen_ai.response.model = gpt-5.6-luna gen_ai.response.id = resp_5 gen_ai.usage.input_tokens = 1310 gen_ai.usage.output_tokens = 42 gen_ai.response.finish_reasons = ('stop',) chat gpt-5.6-luna [CLIENT] status=ERROR gen_ai.operation.name = chat gen_ai.provider.name = openai gen_ai.request.model = gpt-5.6-luna error.type = RateLimitError gen_ai.client.token.usage (input): count=1 sum=1.31e+03 gen_ai.client.token.usage (output): count=1 sum=42 gen_ai.client.operation.duration (ok): count=1 sum=0.00407 gen_ai.client.operation.duration (RateLimitError): count=1 sum=0.00212

3. A whole RAG request as one trace

The value appears when every step of a request is in one trace. Here is the lesson 09 pipeline with a span per step. Retrieval is real (mini_rag) and the model is the offline stand-in. Custom details that have no convention use your own app.* namespace:

import re

from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import SimpleSpanProcessor
from opentelemetry.sdk.trace.export.in_memory_span_exporter import InMemorySpanExporter
from opentelemetry.trace import SpanKind, Status, StatusCode
from fake_openai import fake_client
from mini_rag import HybridIndex
from site_docs import load_sections

exporter = InMemorySpanExporter()
provider = TracerProvider()
provider.add_span_processor(SimpleSpanProcessor(exporter))
tracer = provider.get_tracer("docs_copilot")
index = HybridIndex(load_sections())
client = fake_client(script=["warm-up", "Keep the last key you showed and ask for rows after it [1]."])
client.responses.create(model="gpt-5.6-luna", input="hi")   # keep one-time client setup out of the trace
PRICE_IN, PRICE_OUT = 0.20 / 1e6, 1.20 / 1e6

def ask(question):
    with tracer.start_as_current_span("POST /ask", kind=SpanKind.SERVER) as root:
        root.set_attributes({"app.prompt.version": "answer-v3", "app.index.version": "2026-09-19"})
        with tracer.start_as_current_span("retrieval site-lessons", attributes={
                "gen_ai.operation.name": "retrieval", "gen_ai.data_source.id": "site-lessons"}) as r:
            hits = index.search(question, k=3)
            r.set_attributes({"app.retrieval.urls": [h.doc["url"] for h in hits], "app.retrieval.k": 3})
        with tracer.start_as_current_span("chat gpt-5.6-luna", kind=SpanKind.CLIENT, attributes={
                "gen_ai.operation.name": "chat", "gen_ai.provider.name": "openai",
                "gen_ai.request.model": "gpt-5.6-luna"}) as llm:
            sources = "\n".join(f"[{i}] {h.doc['text'][:600]}" for i, h in enumerate(hits, 1))
            resp = client.responses.create(model="gpt-5.6-luna", input=f"{sources}\nQuestion: {question}")
            u = resp.usage
            llm.set_attributes({"gen_ai.usage.input_tokens": u.input_tokens,
                                "gen_ai.usage.output_tokens": u.output_tokens,
                                "app.cost_usd": round(u.input_tokens * PRICE_IN + u.output_tokens * PRICE_OUT, 7)})
        with tracer.start_as_current_span("verify_citations") as v:
            cited = [int(n) for n in re.findall(r"\[(\d+)\]", resp.output_text)]
            grounded = bool(cited) and all(1 <= n <= len(hits) for n in cited)
            v.set_attribute("app.grounded", grounded)
            if not grounded:
                v.set_status(Status(StatusCode.ERROR, "ungrounded answer"))
        root.set_attribute("app.grounded", grounded)
        return resp.output_text

ask("Why is OFFSET slow for deep pages?")
spans = sorted(exporter.get_finished_spans(), key=lambda s: s.start_time)
root = next(s for s in spans if s.parent is None)
for s in spans:
    depth = 0 if s.parent is None else 1
    ms = (s.end_time - s.start_time) / 1e6
    offset = (s.start_time - root.start_time) / 1e6
    shown = {k: v for k, v in s.attributes.items() if not k.startswith("gen_ai.provider")}
    print(f"{'   ' * depth}{s.name:<26} +{offset:5.0f} ms {ms:6.1f} ms  {shown}")
POST /ask + 0 ms 13.3 ms {'app.prompt.version': 'answer-v3', 'app.index.version': '2026-09-19', 'app.grounded': True} retrieval site-lessons + 0 ms 10.5 ms {'gen_ai.operation.name': 'retrieval', 'gen_ai.data_source.id': 'site-lessons', 'app.retrieval.urls': ('sql/06-order-limit.html#pagination', 'mysql/14-query-optimisation.html#recap', 'sql/06-order-limit.html#recap'), 'app.retrieval.k': 3} chat gpt-5.6-luna + 11 ms 2.2 ms {'gen_ai.operation.name': 'chat', 'gen_ai.request.model': 'gpt-5.6-luna', 'gen_ai.usage.input_tokens': 344, 'gen_ai.usage.output_tokens': 15, 'app.cost_usd': 8.68e-05} verify_citations + 13 ms 0.2 ms {'app.grounded': True}

One look at this trace answers the questions from lesson 15: which prompt and index versions ran, which sections were retrieved, how many tokens were used, what it cost, and whether the answer was grounded. In Grafana Tempo or Langfuse it appears as a waterfall you can click through.

4. Agents and MCP in the same trace

invoke_agent docs-agent chat gpt-5.6-luna execute_tool search_docs tools/call search_docs (MCP client) tools/call search_docs (MCP server process) chat gpt-5.6-luna The dashed span runs in another process: the client puts traceparent in the request's _meta, and the server extracts it (lesson 16), so the MCP server's work appears inside the agent's trace.

5. Automatic instrumentation

You rarely hand-write the model-call span. Instrumentation packages patch the client library for you. The OpenTelemetry project maintains one for the OpenAI SDK (not installed here, so not run):

from opentelemetry.instrumentation.openai_v2 import OpenAIInstrumentor

OpenAIInstrumentor().instrument()        # after setting up the tracer and meter providers (lesson 16)
# every client.chat.completions.create / embeddings.create call now emits GenAI spans and metrics

# prompts and completions are NOT captured unless you opt in, e.g.:
#   OTEL_SEMCONV_STABILITY_OPT_IN=gen_ai_latest_experimental
#   OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=span_only
OptionWhat it isNote
opentelemetry-instrumentation-openai-v2official OTel contrib packagefollows the GenAI conventions closely; check its docs for which OpenAI APIs it covers in your version
OpenLLMetry (Traceloop)instrumentations for many LLM SDKs, vector DBs and frameworksOTel-based; broad coverage
OpenInference (Arize)instrumentations plus its own attribute conventionsused by Phoenix; also OTel-based
Langfuse SDK / integrations@observe, a drop-in OpenAI wrapper, framework callbacksbuilt on OTel; covered next lesson
Your own wrapper (section 2)full control, no dependencyright for custom steps (retrieval, verification) in any case

A sensible split: auto-instrument the libraries (HTTP, database, OpenAI), and hand-instrument your own pipeline steps (retrieval, verification, guardrails) with the conventions where they exist and app.* attributes where they do not.

Recap

  • GenAI conventions: chat {model} CLIENT spans with gen_ai.* attributes; gen_ai.client.token.usage and …operation.duration histograms.
  • Content capture is opt-in; metadata (model, tokens, latency, errors) is always safe to record.
  • Trace the whole request: retrieval, model, tools, verification, with versions and cost on the spans.
  • Auto-instrument libraries, hand-instrument your pipeline, and keep it in one module because the conventions still evolve.

Checkpoint

1 · What should a model-call span be named under the GenAI conventions?
Span names must be low-cardinality. The operation and model identify the kind of call; ids and text go into attributes (or nowhere).
2 · Legal asks whether prompts are stored in your tracing backend. With default OTel GenAI instrumentation…
Content attributes are opt-in in the conventions and the instrumentations. Enabling them is a deliberate, reviewable decision.