Why LLM apps need observability
A classic web service is healthy if it is fast and returns no errors. An LLM application can be fast, error-free and wrong: it answers from the wrong document, ignores a new prompt rule, or quietly costs five times more after a change. Observability means being able to answer "what happened on this request, and is the system getting better or worse?" from the data your system emits. This lesson explains what to capture and why. The next four build it with OpenTelemetry, Langfuse and Grafana.
A flight recorder, not just an arrival board
The arrival board says the plane landed on time. The flight recorder says what every instrument read, which decisions the crew made, and when. When something goes subtly wrong, only the recorder helps. Uptime and latency dashboards are the arrival board. Traces with prompts, retrieval hits, tokens and quality scores are the recorder.
1. What is different about LLM applications
| Property | Consequence for monitoring |
|---|---|
| Non-deterministic output | you cannot replay a bug. You need the exact prompt, context and parameters that were used. |
| Multi-step (retrieve, rewrite, call tools, generate) | "slow" or "wrong" could be any step. You need the step tree: a trace. |
| Cost varies by request | a few long agent runs can dominate the bill. Track tokens and cost per request, user and feature. |
| Quality failures return HTTP 200 | error rates stay green while answers degrade. You need quality signals: evals, groundedness, feedback. |
| Behaviour changes without a code deploy | prompt edits, model version updates and index refreshes all change outputs. Record versions on every call. |
| Inputs contain personal data | capturing prompts is useful and risky. Redaction, sampling and retention rules are part of the design. |
2. The signals
3. Averages lie: look at distributions
LLM latency and cost have long tails: a few requests with huge contexts or long agent loops. A simulation of 2,000 requests of a RAG feature, where 2% are multi-step agent runs, shows why dashboards use percentiles and why cost needs a per-request view:
import numpy as np
rng = np.random.default_rng(7)
n = 2_000
input_tokens = rng.lognormal(mean=7.6, sigma=0.5, size=n).astype(int) # median ≈ 2,000
output_tokens = rng.lognormal(mean=5.3, sigma=0.7, size=n).astype(int) # median ≈ 200
agent = rng.random(n) < 0.02 # 2% are agent runs
input_tokens[agent] *= 8
output_tokens[agent] *= 5
latency = 0.35 + input_tokens / 5_000 + output_tokens / 90 + rng.exponential(0.2, n)
cost = (input_tokens * 0.20 + output_tokens * 1.20) / 1e6 # gpt-5.6-luna prices
p50, p95, p99 = np.percentile(latency, [50, 95, 99])
print(f"latency mean {latency.mean():.2f}s | p50 {p50:.2f}s p95 {p95:.2f}s p99 {p99:.2f}s max {latency.max():.1f}s")
top = np.sort(cost)[::-1]
print(f"cost total ${cost.sum():.2f} | the most expensive 1% of requests = {top[:n // 100].sum() / cost.sum():.0%} of spend")
print(f"agent runs are {agent.mean():.0%} of requests and {cost[agent].sum() / cost.sum():.0%} of spend")
The mean hides the users waiting at p99 (six times the median here), and a total cost figure hides that 2% of requests, the agent runs, account for 14% of the spend. Both show up only when every request records its tokens, cost and latency with labels such as feature, model and user tier that let you slice them.
4. Silent quality drift
Here is the failure observability exists for. On day 19 a prompt edit ships. Nothing errors. Latency is unchanged. But the share of answers that cite a valid source drops. With a quality metric and a simple baseline rule, you catch it the same day:
import numpy as np
rng = np.random.default_rng(3)
days = 30
http_errors = rng.normal(0.004, 0.001, days).clip(0) # 0.4% errors, unchanged all month
grounded = rng.normal(0.93, 0.008, days) # share of answers with valid citations
grounded[18:] -= 0.07 # day 19: prompt v7 goes live
for day in range(7, days):
window = grounded[day - 7:day]
threshold = window.mean() - 3 * window.std()
if grounded[day] < threshold:
print(f"day {day + 1}: grounded {grounded[day]:.1%} < 7-day baseline {window.mean():.1%} - 3σ → ALERT")
break
print(f"HTTP error rate the whole month: {http_errors.min():.2%}–{http_errors.max():.2%} (all green)")
The alert tells you that quality dropped. Traces tagged with prompt.version tell you
why: group the grounded rate by prompt version and v7 stands out immediately. That is why lesson
04 insisted on logging versions with every call.
5. What to capture on every request
| Category | Fields |
|---|---|
| Identity | trace id, request id, session/conversation id, user or tenant id (hashed if needed), feature name |
| Versions | model (requested and returned), prompt name and version, index version, app version |
| Retrieval | query (or its hash), rewritten query, top-k ids and scores, filters applied |
| Model call | input, output, cached and reasoning tokens; cost; TTFT and total latency; finish reason; retries |
| Tools | each call: name, arguments (redacted), result size, duration, error |
| Safety | guardrail triggers, redactions, refusals |
| Quality | citation validity, groundedness score, judge verdict (sampled), user feedback |
| Content (opt-in) | prompts and outputs: sampled, redacted, with a retention limit |
6. The tool landscape, and how it fits together
OpenTelemetry
The vendor-neutral standard for emitting traces, metrics and logs. Instrument once, send anywhere.
Langfuse
Open-source LLM engineering platform: traces of prompts and generations, costs, scores, prompt management, datasets. Its Python SDK is built on OpenTelemetry.
Grafana stack
Prometheus (metrics), Tempo (traces), Loki (logs) and Grafana (dashboards and alerts): the operations view.
Alternatives exist in every box: Arize Phoenix, LangSmith and Helicone for LLM tracing; Datadog, Honeycomb and New Relic as all-in-one platforms. Because they all accept OpenTelemetry, instrumenting with OTel keeps your options open.
7. Privacy is part of the design
Content is opt-in
The OpenTelemetry GenAI conventions mark prompt and response content as opt-in attributes. Metadata (model, tokens, latency) is safe by default; text is not.
Redact, sample, expire
Redact PII before export (in code or in the Collector), store full content for a sample only, and set retention limits on traces that contain text.
Recap
- LLM apps fail quietly: wrong answers, drifting quality and runaway cost all return HTTP 200.
- Traces, metrics, logs and scores: the step tree, the aggregates, the details, and the quality.
- Use percentiles and per-request cost; tails and outliers drive both user pain and spend.
- Record versions and quality signals on every request, and treat captured content as sensitive data.