Module 4 · Observability & operations

Why LLM apps need observability

Intermediate 14 min read HTTP 200 is not "working"

A classic web service is healthy if it is fast and returns no errors. An LLM application can be fast, error-free and wrong: it answers from the wrong document, ignores a new prompt rule, or quietly costs five times more after a change. Observability means being able to answer "what happened on this request, and is the system getting better or worse?" from the data your system emits. This lesson explains what to capture and why. The next four build it with OpenTelemetry, Langfuse and Grafana.

✈️

A flight recorder, not just an arrival board

The arrival board says the plane landed on time. The flight recorder says what every instrument read, which decisions the crew made, and when. When something goes subtly wrong, only the recorder helps. Uptime and latency dashboards are the arrival board. Traces with prompts, retrieval hits, tokens and quality scores are the recorder.

1. What is different about LLM applications

PropertyConsequence for monitoring
Non-deterministic outputyou cannot replay a bug. You need the exact prompt, context and parameters that were used.
Multi-step (retrieve, rewrite, call tools, generate)"slow" or "wrong" could be any step. You need the step tree: a trace.
Cost varies by requesta few long agent runs can dominate the bill. Track tokens and cost per request, user and feature.
Quality failures return HTTP 200error rates stay green while answers degrade. You need quality signals: evals, groundedness, feedback.
Behaviour changes without a code deployprompt edits, model version updates and index refreshes all change outputs. Record versions on every call.
Inputs contain personal datacapturing prompts is useful and risky. Redaction, sampling and retention rules are part of the design.

2. The signals

traces one request as a tree "why was THIS request slow, costly or wrong?" metrics numbers over time p95 latency, tokens/min, cost/day, error rate cheap to keep for years logs & events details, as records prompt v7 rendered retrieved 4 chunks guardrail: pii redacted linked to traces by trace_id scores quality, attached to traces groundedness 0.92 user feedback 👍 judge: faithful the LLM-specific one

3. Averages lie: look at distributions

LLM latency and cost have long tails: a few requests with huge contexts or long agent loops. A simulation of 2,000 requests of a RAG feature, where 2% are multi-step agent runs, shows why dashboards use percentiles and why cost needs a per-request view:

import numpy as np

rng = np.random.default_rng(7)
n = 2_000
input_tokens = rng.lognormal(mean=7.6, sigma=0.5, size=n).astype(int)     # median ≈ 2,000
output_tokens = rng.lognormal(mean=5.3, sigma=0.7, size=n).astype(int)    # median ≈ 200
agent = rng.random(n) < 0.02                                             # 2% are agent runs
input_tokens[agent] *= 8
output_tokens[agent] *= 5

latency = 0.35 + input_tokens / 5_000 + output_tokens / 90 + rng.exponential(0.2, n)
cost = (input_tokens * 0.20 + output_tokens * 1.20) / 1e6                  # gpt-5.6-luna prices

p50, p95, p99 = np.percentile(latency, [50, 95, 99])
print(f"latency  mean {latency.mean():.2f}s | p50 {p50:.2f}s  p95 {p95:.2f}s  p99 {p99:.2f}s  max {latency.max():.1f}s")
top = np.sort(cost)[::-1]
print(f"cost     total ${cost.sum():.2f} | the most expensive 1% of requests = {top[:n // 100].sum() / cost.sum():.0%} of spend")
print(f"agent runs are {agent.mean():.0%} of requests and {cost[agent].sum() / cost.sum():.0%} of spend")
latency mean 4.21s | p50 3.26s p95 9.52s p99 20.78s max 63.5s cost total $1.71 | the most expensive 1% of requests = 8% of spend agent runs are 2% of requests and 14% of spend

The mean hides the users waiting at p99 (six times the median here), and a total cost figure hides that 2% of requests, the agent runs, account for 14% of the spend. Both show up only when every request records its tokens, cost and latency with labels such as feature, model and user tier that let you slice them.

4. Silent quality drift

Here is the failure observability exists for. On day 19 a prompt edit ships. Nothing errors. Latency is unchanged. But the share of answers that cite a valid source drops. With a quality metric and a simple baseline rule, you catch it the same day:

import numpy as np

rng = np.random.default_rng(3)
days = 30
http_errors = rng.normal(0.004, 0.001, days).clip(0)          # 0.4% errors, unchanged all month
grounded = rng.normal(0.93, 0.008, days)                      # share of answers with valid citations
grounded[18:] -= 0.07                                         # day 19: prompt v7 goes live

for day in range(7, days):
    window = grounded[day - 7:day]
    threshold = window.mean() - 3 * window.std()
    if grounded[day] < threshold:
        print(f"day {day + 1}: grounded {grounded[day]:.1%} < 7-day baseline {window.mean():.1%} - 3σ  → ALERT")
        break
print(f"HTTP error rate the whole month: {http_errors.min():.2%}–{http_errors.max():.2%} (all green)")
day 19: grounded 86.2% < 7-day baseline 93.1% - 3σ → ALERT HTTP error rate the whole month: 0.14%–0.73% (all green)
The fix is in the trace

The alert tells you that quality dropped. Traces tagged with prompt.version tell you why: group the grounded rate by prompt version and v7 stands out immediately. That is why lesson 04 insisted on logging versions with every call.

5. What to capture on every request

CategoryFields
Identitytrace id, request id, session/conversation id, user or tenant id (hashed if needed), feature name
Versionsmodel (requested and returned), prompt name and version, index version, app version
Retrievalquery (or its hash), rewritten query, top-k ids and scores, filters applied
Model callinput, output, cached and reasoning tokens; cost; TTFT and total latency; finish reason; retries
Toolseach call: name, arguments (redacted), result size, duration, error
Safetyguardrail triggers, redactions, refusals
Qualitycitation validity, groundedness score, judge verdict (sampled), user feedback
Content (opt-in)prompts and outputs: sampled, redacted, with a retention limit

6. The tool landscape, and how it fits together

your app OpenTelemetry SDK lessons 16–17 OTLP OTel Collector batch, redact, route Tempo (traces) Prometheus Loki (logs) Grafana dashboards, alerts · 19 Langfuse LLM traces, evals · 18

OpenTelemetry

The vendor-neutral standard for emitting traces, metrics and logs. Instrument once, send anywhere.

Langfuse

Open-source LLM engineering platform: traces of prompts and generations, costs, scores, prompt management, datasets. Its Python SDK is built on OpenTelemetry.

Grafana stack

Prometheus (metrics), Tempo (traces), Loki (logs) and Grafana (dashboards and alerts): the operations view.

Alternatives exist in every box: Arize Phoenix, LangSmith and Helicone for LLM tracing; Datadog, Honeycomb and New Relic as all-in-one platforms. Because they all accept OpenTelemetry, instrumenting with OTel keeps your options open.

7. Privacy is part of the design

Content is opt-in

The OpenTelemetry GenAI conventions mark prompt and response content as opt-in attributes. Metadata (model, tokens, latency) is safe by default; text is not.

Redact, sample, expire

Redact PII before export (in code or in the Collector), store full content for a sample only, and set retention limits on traces that contain text.

Recap

  • LLM apps fail quietly: wrong answers, drifting quality and runaway cost all return HTTP 200.
  • Traces, metrics, logs and scores: the step tree, the aggregates, the details, and the quality.
  • Use percentiles and per-request cost; tails and outliers drive both user pain and spend.
  • Record versions and quality signals on every request, and treat captured content as sensitive data.

Checkpoint

1 · Error rate 0.3%, p95 latency steady, yet support tickets about wrong answers doubled this week. What signal was missing?
Wrong answers are successful HTTP responses. Only quality measurements, sliced by what changed, reveal them.
2 · Why track p95 and p99 latency rather than the mean?
A few huge-context or agent requests take many times longer. Percentiles describe what a given share of users actually wait.