Tracing LLM calls with OpenTelemetry
OpenTelemetry has semantic conventions for generative AI: agreed span names, attribute names and metrics for model calls, retrieval, tool execution and agents. Follow them and any backend (Grafana, Langfuse, Datadog, Phoenix) understands your traces without custom mapping. This lesson instruments a model call by hand to learn exactly what the conventions ask for, traces a whole RAG request, and then shows the libraries that do it automatically.
Standard shipping labels
Couriers can scan any parcel because the label layout is standard: the postcode is always in the same
box. Semantic conventions are the label layout for telemetry. When every tool writes
gen_ai.usage.input_tokens, every dashboard can add it up.
1. The conventions you will use most
| Span | Name | Key attributes |
|---|---|---|
| Model call (kind CLIENT) | {operation} {model}, e.g. chat gpt-5.6-luna | gen_ai.operation.name (chat, embeddings…), gen_ai.provider.name (openai), gen_ai.request.model, gen_ai.response.model, gen_ai.response.id, gen_ai.usage.input_tokens, gen_ai.usage.output_tokens, gen_ai.response.finish_reasons, error.type |
| Retrieval | retrieval {data_source} | gen_ai.operation.name = retrieval, gen_ai.data_source.id |
| Tool execution | execute_tool {tool} | gen_ai.tool.name, gen_ai.tool.call.id |
| Agent | invoke_agent {agent} | gen_ai.agent.name, gen_ai.conversation.id |
| Content (opt-in only) | — | gen_ai.input.messages, gen_ai.output.messages, gen_ai.system_instructions |
| Metric | Instrument | Notes |
|---|---|---|
gen_ai.client.token.usage | histogram, unit {token} | one record per call per gen_ai.token.type (input, output); buckets 1, 4, 16 … 67,108,864 |
gen_ai.client.operation.duration | histogram, unit s | buckets 0.01, 0.02, 0.04 … 81.92 |
gen_ai.client.operation.time_to_first_chunk | histogram, unit s | for streaming: the TTFT of lesson 03 |
The GenAI conventions (now maintained in their own semantic-conventions-genai repository) are
marked Development, and names have changed before: gen_ai.system became
gen_ai.provider.name. Keep your instrumentation in one module so a rename is a one-line
change, and check the current spec when you upgrade.
2. Instrument a model call by hand
A wrapper around responses.create that emits a convention-compliant span and both standard
metrics, and captures message content only when explicitly enabled. It runs through the real SDK against
the course's fake server, including a rate-limit failure:
import json
import os
import time
import openai
from opentelemetry import metrics, trace
from opentelemetry.sdk.metrics import MeterProvider
from opentelemetry.sdk.metrics.export import InMemoryMetricReader
from opentelemetry.sdk.metrics.view import ExplicitBucketHistogramAggregation, View
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import SimpleSpanProcessor
from opentelemetry.sdk.trace.export.in_memory_span_exporter import InMemorySpanExporter
from opentelemetry.trace import SpanKind
from fake_openai import fake_client
spans = InMemorySpanExporter()
tp = TracerProvider()
tp.add_span_processor(SimpleSpanProcessor(spans))
reader = InMemoryMetricReader()
mp = MeterProvider(metric_readers=[reader], views=[
View(instrument_name="gen_ai.client.token.usage", aggregation=ExplicitBucketHistogramAggregation(
[1, 4, 16, 64, 256, 1024, 4096, 16384, 65536, 262144, 1048576, 4194304, 16777216, 67108864])),
View(instrument_name="gen_ai.client.operation.duration", aggregation=ExplicitBucketHistogramAggregation(
[0.01, 0.02, 0.04, 0.08, 0.16, 0.32, 0.64, 1.28, 2.56, 5.12, 10.24, 20.48, 40.96, 81.92]))])
tracer, meter = tp.get_tracer("genai.demo"), mp.get_meter("genai.demo")
token_usage = meter.create_histogram("gen_ai.client.token.usage", unit="{token}")
duration = meter.create_histogram("gen_ai.client.operation.duration", unit="s")
CAPTURE = os.environ.get("OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT", "").lower() in (
"true", "span_only", "span_and_event") # opt-in, off by default
def traced_response(client, model, input, instructions=None, **kwargs):
attrs = {"gen_ai.operation.name": "chat", "gen_ai.provider.name": "openai", "gen_ai.request.model": model}
start = time.perf_counter()
with tracer.start_as_current_span(f"chat {model}", kind=SpanKind.CLIENT, attributes=attrs) as span:
if CAPTURE:
span.set_attribute("gen_ai.input.messages", json.dumps([{"role": "user", "content": input}]))
try:
r = client.responses.create(model=model, input=input, instructions=instructions, **kwargs)
except openai.APIError as e:
error = type(e).__name__
span.set_attribute("error.type", error) # the exception itself is recorded
duration.record(time.perf_counter() - start, {**attrs, "error.type": error})
raise
finish = "stop" if r.status == "completed" else "length"
span.set_attributes({"gen_ai.response.model": r.model, "gen_ai.response.id": r.id,
"gen_ai.usage.input_tokens": r.usage.input_tokens,
"gen_ai.usage.output_tokens": r.usage.output_tokens,
"gen_ai.response.finish_reasons": [finish]})
labels = {**attrs, "gen_ai.response.model": r.model}
token_usage.record(r.usage.input_tokens, {**labels, "gen_ai.token.type": "input"})
token_usage.record(r.usage.output_tokens, {**labels, "gen_ai.token.type": "output"})
duration.record(time.perf_counter() - start, labels)
return r
client = fake_client(script=["warm-up", {"text": "Use keyset pagination.", "usage": (1310, 42)},
{"status": 429}], max_retries=0)
client.responses.create(model="gpt-5.6-luna", input="hi") # first call pays one-time client setup
traced_response(client, "gpt-5.6-luna", "Why is OFFSET slow for deep pages?")
try:
traced_response(client, "gpt-5.6-luna", "Second question")
except openai.RateLimitError:
pass
for s in spans.get_finished_spans():
print(f"{s.name} [{s.kind.name}] status={s.status.status_code.name}")
for k, v in s.attributes.items():
print(f" {k} = {v}")
for m in reader.get_metrics_data().resource_metrics[0].scope_metrics[0].metrics:
for p in m.data.data_points:
tag = p.attributes.get("gen_ai.token.type") or p.attributes.get("error.type") or "ok"
print(f"{m.name} ({tag}): count={p.count} sum={p.sum:.3g}")
3. A whole RAG request as one trace
The value appears when every step of a request is in one trace. Here is the lesson 09 pipeline with a
span per step. Retrieval is real (mini_rag) and the model is the offline stand-in. Custom
details that have no convention use your own app.* namespace:
import re
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import SimpleSpanProcessor
from opentelemetry.sdk.trace.export.in_memory_span_exporter import InMemorySpanExporter
from opentelemetry.trace import SpanKind, Status, StatusCode
from fake_openai import fake_client
from mini_rag import HybridIndex
from site_docs import load_sections
exporter = InMemorySpanExporter()
provider = TracerProvider()
provider.add_span_processor(SimpleSpanProcessor(exporter))
tracer = provider.get_tracer("docs_copilot")
index = HybridIndex(load_sections())
client = fake_client(script=["warm-up", "Keep the last key you showed and ask for rows after it [1]."])
client.responses.create(model="gpt-5.6-luna", input="hi") # keep one-time client setup out of the trace
PRICE_IN, PRICE_OUT = 0.20 / 1e6, 1.20 / 1e6
def ask(question):
with tracer.start_as_current_span("POST /ask", kind=SpanKind.SERVER) as root:
root.set_attributes({"app.prompt.version": "answer-v3", "app.index.version": "2026-09-19"})
with tracer.start_as_current_span("retrieval site-lessons", attributes={
"gen_ai.operation.name": "retrieval", "gen_ai.data_source.id": "site-lessons"}) as r:
hits = index.search(question, k=3)
r.set_attributes({"app.retrieval.urls": [h.doc["url"] for h in hits], "app.retrieval.k": 3})
with tracer.start_as_current_span("chat gpt-5.6-luna", kind=SpanKind.CLIENT, attributes={
"gen_ai.operation.name": "chat", "gen_ai.provider.name": "openai",
"gen_ai.request.model": "gpt-5.6-luna"}) as llm:
sources = "\n".join(f"[{i}] {h.doc['text'][:600]}" for i, h in enumerate(hits, 1))
resp = client.responses.create(model="gpt-5.6-luna", input=f"{sources}\nQuestion: {question}")
u = resp.usage
llm.set_attributes({"gen_ai.usage.input_tokens": u.input_tokens,
"gen_ai.usage.output_tokens": u.output_tokens,
"app.cost_usd": round(u.input_tokens * PRICE_IN + u.output_tokens * PRICE_OUT, 7)})
with tracer.start_as_current_span("verify_citations") as v:
cited = [int(n) for n in re.findall(r"\[(\d+)\]", resp.output_text)]
grounded = bool(cited) and all(1 <= n <= len(hits) for n in cited)
v.set_attribute("app.grounded", grounded)
if not grounded:
v.set_status(Status(StatusCode.ERROR, "ungrounded answer"))
root.set_attribute("app.grounded", grounded)
return resp.output_text
ask("Why is OFFSET slow for deep pages?")
spans = sorted(exporter.get_finished_spans(), key=lambda s: s.start_time)
root = next(s for s in spans if s.parent is None)
for s in spans:
depth = 0 if s.parent is None else 1
ms = (s.end_time - s.start_time) / 1e6
offset = (s.start_time - root.start_time) / 1e6
shown = {k: v for k, v in s.attributes.items() if not k.startswith("gen_ai.provider")}
print(f"{' ' * depth}{s.name:<26} +{offset:5.0f} ms {ms:6.1f} ms {shown}")
One look at this trace answers the questions from lesson 15: which prompt and index versions ran, which sections were retrieved, how many tokens were used, what it cost, and whether the answer was grounded. In Grafana Tempo or Langfuse it appears as a waterfall you can click through.
4. Agents and MCP in the same trace
5. Automatic instrumentation
You rarely hand-write the model-call span. Instrumentation packages patch the client library for you. The OpenTelemetry project maintains one for the OpenAI SDK (not installed here, so not run):
from opentelemetry.instrumentation.openai_v2 import OpenAIInstrumentor OpenAIInstrumentor().instrument() # after setting up the tracer and meter providers (lesson 16) # every client.chat.completions.create / embeddings.create call now emits GenAI spans and metrics # prompts and completions are NOT captured unless you opt in, e.g.: # OTEL_SEMCONV_STABILITY_OPT_IN=gen_ai_latest_experimental # OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=span_only
| Option | What it is | Note |
|---|---|---|
opentelemetry-instrumentation-openai-v2 | official OTel contrib package | follows the GenAI conventions closely; check its docs for which OpenAI APIs it covers in your version |
| OpenLLMetry (Traceloop) | instrumentations for many LLM SDKs, vector DBs and frameworks | OTel-based; broad coverage |
| OpenInference (Arize) | instrumentations plus its own attribute conventions | used by Phoenix; also OTel-based |
| Langfuse SDK / integrations | @observe, a drop-in OpenAI wrapper, framework callbacks | built on OTel; covered next lesson |
| Your own wrapper (section 2) | full control, no dependency | right for custom steps (retrieval, verification) in any case |
A sensible split: auto-instrument the libraries (HTTP, database, OpenAI), and hand-instrument your own
pipeline steps (retrieval, verification, guardrails) with the conventions where they exist and
app.* attributes where they do not.
Recap
- GenAI conventions:
chat {model}CLIENT spans withgen_ai.*attributes;gen_ai.client.token.usageand…operation.durationhistograms. - Content capture is opt-in; metadata (model, tokens, latency, errors) is always safe to record.
- Trace the whole request: retrieval, model, tools, verification, with versions and cost on the spans.
- Auto-instrument libraries, hand-instrument your pipeline, and keep it in one module because the conventions still evolve.