Grafana, Prometheus, Tempo & Loki
The Grafana stack is the operations view of your LLM system. Prometheus stores metrics, Tempo stores traces, Loki stores logs, and Grafana puts them on one screen with links between them: from a latency spike, to an example slow trace, to that trace's log lines. An OpenTelemetry Collector in front receives everything your app emits and routes it. This lesson covers the setup, the queries (PromQL, TraceQL, LogQL), a dashboard, and alerts that fire on the things that matter for LLM apps.
A control room
Big gauges on the wall (metrics) tell you something is off. The incident log (logs) says what was written at the time. The security camera recording (traces) shows exactly what happened to one visitor. A good control room lets you click from the gauge to the recording of the moment it spiked.
1. Who does what
| Component | Stores | Query language | LLM questions it answers |
|---|---|---|---|
| OTel Collector | nothing: receives, processes, routes | YAML config | redact prompts before storage; fan out to Tempo, Prometheus, Loki and Langfuse |
| Prometheus | metrics (time series) | PromQL | p95 latency, tokens/min, cost/hour, error and grounded rates |
| Tempo | traces | TraceQL | show me slow, expensive or ungrounded requests, step by step |
| Loki | logs (indexed by labels only) | LogQL | what did the app log for this trace id? |
| Grafana | dashboards, alerts | all of the above | one screen, cross-links, alerting and on-call routing |
2. Run the whole stack locally
For learning and development, Grafana publishes grafana/otel-lgtm: one container with the
Collector, Prometheus, Tempo, Loki (and Pyroscope) and a pre-wired Grafana. Point your app's OTLP exporter
at it and everything from lessons 16β17 shows up:
services:
lgtm:
image: grafana/otel-lgtm
ports:
- "3000:3000" # Grafana: http://localhost:3000 (default login admin / admin; change it)
- "4317:4317" # OTLP gRPC
- "4318:4318" # OTLP HTTP
volumes:
- ./grafana/llm-dashboard.json:/otel-lgtm/grafana/conf/provisioning/dashboards/custom/llm-dashboard.json:ro
- ./grafana/dashboards.yaml:/otel-lgtm/grafana/conf/provisioning/dashboards/custom.yaml:ro
docker compose up -d export OTEL_SERVICE_NAME=docs-copilot OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4317 python app.py # the setup_telemetry() from lesson 16
In production each component runs separately, scaled and with retention configured, and the Collector config becomes yours to own.
3. The Collector config
A Collector pipeline is receivers β processors β exporters, one pipeline per signal. Two processors earn their place in every LLM deployment: redaction of message content (defence in depth, even if the app should not send it), and spanmetrics, which derives request, error and duration metrics from spans, so traces alone give you dashboards.
receivers:
otlp:
protocols:
grpc: { endpoint: 0.0.0.0:4317 }
http: { endpoint: 0.0.0.0:4318 }
processors:
memory_limiter: { check_interval: 1s, limit_percentage: 80 }
batch: {}
attributes/redact: # drop prompt and response text if anything sends it
actions:
- { key: gen_ai.input.messages, action: delete }
- { key: gen_ai.output.messages, action: delete }
- { key: gen_ai.system_instructions, action: delete }
connectors:
spanmetrics: # traces β RED metrics, with model as a dimension
dimensions: [{ name: gen_ai.request.model }, { name: gen_ai.operation.name }]
exporters:
otlp/tempo: { endpoint: tempo:4317, tls: { insecure: true } }
otlphttp/prom: { endpoint: http://prometheus:9090/api/v1/otlp } # Prometheus's native OTLP receiver
otlphttp/loki: { endpoint: http://loki:3100/otlp } # Loki's native OTLP endpoint
otlphttp/langfuse:
endpoint: https://cloud.langfuse.com/api/public/otel
headers: { Authorization: "Basic ${env:LANGFUSE_AUTH}" } # base64("pk:sk"), from a secret
service:
pipelines:
traces: { receivers: [otlp], processors: [memory_limiter, attributes/redact, batch],
exporters: [otlp/tempo, spanmetrics, otlphttp/langfuse] }
metrics: { receivers: [otlp, spanmetrics], processors: [memory_limiter, batch], exporters: [otlphttp/prom] }
logs: { receivers: [otlp], processors: [memory_limiter, batch], exporters: [otlphttp/loki] }
This config follows the Collector's documented components, but it was not run here, since Docker is not running on this machine. Endpoint paths depend on your Prometheus, Loki and Langfuse versions (Prometheus needs its OTLP receiver enabled), so check each backend's OTLP ingestion docs when you deploy.
4. PromQL for LLM metrics
OTel metric names are translated for Prometheus: dots become underscores and the unit becomes a suffix,
so gen_ai.client.operation.duration (seconds) appears as
gen_ai_client_operation_duration_seconds_bucket, _sum and _count.
Check the exact names in Grafana's metric browser, because translation settings vary.
| Panel | PromQL |
|---|---|
| p95 model latency, by model | histogram_quantile(0.95, sum by (le, gen_ai_request_model) (rate(gen_ai_client_operation_duration_seconds_bucket[5m]))) |
| Tokens per minute, input vs output | 60 * sum by (gen_ai_token_type) (rate(gen_ai_client_token_usage_sum[5m])) |
| Model error ratio | sum(rate(gen_ai_client_operation_duration_seconds_count{error_type!=""}[5m])) / sum(rate(gen_ai_client_operation_duration_seconds_count[5m])) |
| Spend per hour (an app counter) | sum(increase(app_cost_usd_total[1h])) |
| Grounded answer rate | sum(rate(app_answers_total{grounded="true"}[1h])) / sum(rate(app_answers_total[1h])) |
histogram_quantile does not see raw latencies, only bucket counts, so it estimates
the percentile by interpolating inside the bucket that contains it. Bucket boundaries therefore decide
how accurate your p95 is. Here is the same algorithm in Python, against the true value:
import numpy as np
def histogram_quantile(q, bounds, cumulative_counts):
"""Prometheus's estimate: find the bucket holding rank q*N, interpolate linearly inside it."""
total = cumulative_counts[-1]
rank = q * total
lower, prev = 0.0, 0
for upper, count in zip(bounds, cumulative_counts):
if count >= rank:
if upper == float("inf"):
return lower # the +Inf bucket: return its lower bound
return lower + (upper - lower) * (rank - prev) / max(count - prev, 1)
lower, prev = upper, count
return lower
rng = np.random.default_rng(11)
latencies = rng.lognormal(mean=1.0, sigma=0.6, size=20_000) # seconds, long-tailed
true_p95 = np.percentile(latencies, 95)
layouts = {
"GenAI spec (0.01 β¦ 81.92, doubling)": [0.01, 0.02, 0.04, 0.08, 0.16, 0.32, 0.64, 1.28, 2.56,
5.12, 10.24, 20.48, 40.96, 81.92],
"coarse (1, 5, 30)": [1, 5, 30],
"tuned for this SLO (2, 4, 5, 6, 7, 8, 10)": [2, 4, 5, 6, 7, 8, 10],
}
print(f"true p95: {true_p95:.2f}s")
for name, edges in layouts.items():
bounds = edges + [float("inf")]
cumulative = [int((latencies <= b).sum()) for b in bounds]
est = histogram_quantile(0.95, bounds, cumulative)
print(f"{name:44} estimate {est:5.2f}s (error {100 * (est - true_p95) / true_p95:+.0f}%)")
Coarse buckets can be badly wrong (+202% here). The GenAI conventions' doubling boundaries are a reasonable default that works across very different models, but because each bucket is twice the width of the one before, they are still coarse at any single point (+22% here). If you alert on a threshold, add boundaries close together around it. The tuned layout lands within 1%.
5. TraceQL and LogQL: from "something is wrong" to "this request"
| Find⦠| Query |
|---|---|
| slow model calls (TraceQL) | { span.gen_ai.request.model = "gpt-5.6-luna" && duration > 5s } |
| ungrounded answers | { resource.service.name = "docs-copilot" && span.app.grounded = false } |
| huge prompts | { span.gen_ai.usage.input_tokens > 20000 } |
| failed tool calls | { span.gen_ai.operation.name = "execute_tool" && status = error } |
| log lines for one trace (LogQL) | {service_name="docs-copilot"} |= "4bf92f3577b34da6a3ce929d0e0e4736" |
| warning rate by message (LogQL) | sum by (detected_level) (count_over_time({service_name="docs-copilot"} | detected_level="warn" [5m])) |
Configure the Prometheus data source with exemplars (a trace id attached to histogram samples) and the Tempo data source with trace-to-logs. Then a dot on the p95 panel opens the exact slow trace, and the trace view opens its log lines. That three-click path is the whole point of the stack.
6. A dashboard that earns its screen space
7. Alerts that wake people up for the right reasons
| Alert | Condition (example) | Why |
|---|---|---|
| Latency SLO burn | p95 > 8 s for 10 min | users are waiting; often a provider incident or a prompt-size regression |
| Model error ratio | > 5% for 5 min | rate limits, outages, or a bad deploy |
| Spend anomaly | spend/hour > 2Γ the same hour last week | runaway agent loop, abuse, or a caching bug |
| Quality drop | grounded rate below its 7-day baseline β 3Ο (lesson 15) | the silent failure: a prompt, model or index change |
| Pipeline freshness | no successful index sync for 24 h | answers going stale without any error |
groups:
- name: docs-copilot
rules:
- alert: LLMLatencyHigh
expr: histogram_quantile(0.95, sum by (le) (rate(gen_ai_client_operation_duration_seconds_bucket[5m]))) > 8
for: 10m
labels: { severity: page }
annotations: { summary: "p95 model latency {{ $value | humanizeDuration }} (SLO 8s)" }
- alert: LLMSpendAnomaly
expr: sum(increase(app_cost_usd_total[1h])) > 2 * sum(increase(app_cost_usd_total[1h] offset 1w))
for: 15m
labels: { severity: ticket }
Alert on symptoms users feel (latency, errors, quality, spend), not on every cause. Route "page" alerts to on-call and "ticket" alerts to a queue. Every alert should link to a dashboard and a runbook.
Recap
- Collector β Prometheus / Tempo / Loki β Grafana;
grafana/otel-lgtmgives you all of it in one container for development. - Redact in the Collector as defence in depth; derive RED metrics from spans with spanmetrics.
- histogram_quantile estimates from buckets, so choose boundaries near your SLOs.
- Dashboards: traffic, economics, quality. Alerts on symptoms, with links from metric to trace to logs.