Module 4 Β· Observability & operations

Grafana, Prometheus, Tempo & Loki

Advanced 22 min read Dashboards and alerts for latency, tokens, cost and quality

The Grafana stack is the operations view of your LLM system. Prometheus stores metrics, Tempo stores traces, Loki stores logs, and Grafana puts them on one screen with links between them: from a latency spike, to an example slow trace, to that trace's log lines. An OpenTelemetry Collector in front receives everything your app emits and routes it. This lesson covers the setup, the queries (PromQL, TraceQL, LogQL), a dashboard, and alerts that fire on the things that matter for LLM apps.

πŸŽ›οΈ

A control room

Big gauges on the wall (metrics) tell you something is off. The incident log (logs) says what was written at the time. The security camera recording (traces) shows exactly what happened to one visitor. A good control room lets you click from the gauge to the recording of the moment it spiked.

1. Who does what

ComponentStoresQuery languageLLM questions it answers
OTel Collectornothing: receives, processes, routesYAML configredact prompts before storage; fan out to Tempo, Prometheus, Loki and Langfuse
Prometheusmetrics (time series)PromQLp95 latency, tokens/min, cost/hour, error and grounded rates
TempotracesTraceQLshow me slow, expensive or ungrounded requests, step by step
Lokilogs (indexed by labels only)LogQLwhat did the app log for this trace id?
Grafanadashboards, alertsall of the aboveone screen, cross-links, alerting and on-call routing

2. Run the whole stack locally

For learning and development, Grafana publishes grafana/otel-lgtm: one container with the Collector, Prometheus, Tempo, Loki (and Pyroscope) and a pre-wired Grafana. Point your app's OTLP exporter at it and everything from lessons 16–17 shows up:

services:
  lgtm:
    image: grafana/otel-lgtm
    ports:
      - "3000:3000"      # Grafana: http://localhost:3000  (default login admin / admin; change it)
      - "4317:4317"      # OTLP gRPC
      - "4318:4318"      # OTLP HTTP
    volumes:
      - ./grafana/llm-dashboard.json:/otel-lgtm/grafana/conf/provisioning/dashboards/custom/llm-dashboard.json:ro
      - ./grafana/dashboards.yaml:/otel-lgtm/grafana/conf/provisioning/dashboards/custom.yaml:ro
docker compose up -d
export OTEL_SERVICE_NAME=docs-copilot OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4317
python app.py          # the setup_telemetry() from lesson 16

In production each component runs separately, scaled and with retention configured, and the Collector config becomes yours to own.

3. The Collector config

A Collector pipeline is receivers β†’ processors β†’ exporters, one pipeline per signal. Two processors earn their place in every LLM deployment: redaction of message content (defence in depth, even if the app should not send it), and spanmetrics, which derives request, error and duration metrics from spans, so traces alone give you dashboards.

receivers:
  otlp:
    protocols:
      grpc: { endpoint: 0.0.0.0:4317 }
      http: { endpoint: 0.0.0.0:4318 }

processors:
  memory_limiter: { check_interval: 1s, limit_percentage: 80 }
  batch: {}
  attributes/redact:                      # drop prompt and response text if anything sends it
    actions:
      - { key: gen_ai.input.messages,     action: delete }
      - { key: gen_ai.output.messages,    action: delete }
      - { key: gen_ai.system_instructions, action: delete }

connectors:
  spanmetrics:                            # traces β†’ RED metrics, with model as a dimension
    dimensions: [{ name: gen_ai.request.model }, { name: gen_ai.operation.name }]

exporters:
  otlp/tempo:       { endpoint: tempo:4317, tls: { insecure: true } }
  otlphttp/prom:    { endpoint: http://prometheus:9090/api/v1/otlp }   # Prometheus's native OTLP receiver
  otlphttp/loki:    { endpoint: http://loki:3100/otlp }                # Loki's native OTLP endpoint
  otlphttp/langfuse:
    endpoint: https://cloud.langfuse.com/api/public/otel
    headers: { Authorization: "Basic ${env:LANGFUSE_AUTH}" }            # base64("pk:sk"), from a secret

service:
  pipelines:
    traces:  { receivers: [otlp], processors: [memory_limiter, attributes/redact, batch],
               exporters: [otlp/tempo, spanmetrics, otlphttp/langfuse] }
    metrics: { receivers: [otlp, spanmetrics], processors: [memory_limiter, batch], exporters: [otlphttp/prom] }
    logs:    { receivers: [otlp], processors: [memory_limiter, batch], exporters: [otlphttp/loki] }

This config follows the Collector's documented components, but it was not run here, since Docker is not running on this machine. Endpoint paths depend on your Prometheus, Loki and Langfuse versions (Prometheus needs its OTLP receiver enabled), so check each backend's OTLP ingestion docs when you deploy.

4. PromQL for LLM metrics

OTel metric names are translated for Prometheus: dots become underscores and the unit becomes a suffix, so gen_ai.client.operation.duration (seconds) appears as gen_ai_client_operation_duration_seconds_bucket, _sum and _count. Check the exact names in Grafana's metric browser, because translation settings vary.

PanelPromQL
p95 model latency, by modelhistogram_quantile(0.95, sum by (le, gen_ai_request_model) (rate(gen_ai_client_operation_duration_seconds_bucket[5m])))
Tokens per minute, input vs output60 * sum by (gen_ai_token_type) (rate(gen_ai_client_token_usage_sum[5m]))
Model error ratiosum(rate(gen_ai_client_operation_duration_seconds_count{error_type!=""}[5m])) / sum(rate(gen_ai_client_operation_duration_seconds_count[5m]))
Spend per hour (an app counter)sum(increase(app_cost_usd_total[1h]))
Grounded answer ratesum(rate(app_answers_total{grounded="true"}[1h])) / sum(rate(app_answers_total[1h]))

histogram_quantile does not see raw latencies, only bucket counts, so it estimates the percentile by interpolating inside the bucket that contains it. Bucket boundaries therefore decide how accurate your p95 is. Here is the same algorithm in Python, against the true value:

import numpy as np

def histogram_quantile(q, bounds, cumulative_counts):
    """Prometheus's estimate: find the bucket holding rank q*N, interpolate linearly inside it."""
    total = cumulative_counts[-1]
    rank = q * total
    lower, prev = 0.0, 0
    for upper, count in zip(bounds, cumulative_counts):
        if count >= rank:
            if upper == float("inf"):
                return lower                         # the +Inf bucket: return its lower bound
            return lower + (upper - lower) * (rank - prev) / max(count - prev, 1)
        lower, prev = upper, count
    return lower

rng = np.random.default_rng(11)
latencies = rng.lognormal(mean=1.0, sigma=0.6, size=20_000)              # seconds, long-tailed
true_p95 = np.percentile(latencies, 95)

layouts = {
    "GenAI spec (0.01 … 81.92, doubling)": [0.01, 0.02, 0.04, 0.08, 0.16, 0.32, 0.64, 1.28, 2.56,
                                           5.12, 10.24, 20.48, 40.96, 81.92],
    "coarse (1, 5, 30)": [1, 5, 30],
    "tuned for this SLO (2, 4, 5, 6, 7, 8, 10)": [2, 4, 5, 6, 7, 8, 10],
}
print(f"true p95: {true_p95:.2f}s")
for name, edges in layouts.items():
    bounds = edges + [float("inf")]
    cumulative = [int((latencies <= b).sum()) for b in bounds]
    est = histogram_quantile(0.95, bounds, cumulative)
    print(f"{name:44} estimate {est:5.2f}s  (error {100 * (est - true_p95) / true_p95:+.0f}%)")
true p95: 7.25s GenAI spec (0.01 … 81.92, doubling) estimate 8.84s (error +22%) coarse (1, 5, 30) estimate 21.91s (error +202%) tuned for this SLO (2, 4, 5, 6, 7, 8, 10) estimate 7.29s (error +1%)

Coarse buckets can be badly wrong (+202% here). The GenAI conventions' doubling boundaries are a reasonable default that works across very different models, but because each bucket is twice the width of the one before, they are still coarse at any single point (+22% here). If you alert on a threshold, add boundaries close together around it. The tuned layout lands within 1%.

5. TraceQL and LogQL: from "something is wrong" to "this request"

Find…Query
slow model calls (TraceQL){ span.gen_ai.request.model = "gpt-5.6-luna" && duration > 5s }
ungrounded answers{ resource.service.name = "docs-copilot" && span.app.grounded = false }
huge prompts{ span.gen_ai.usage.input_tokens > 20000 }
failed tool calls{ span.gen_ai.operation.name = "execute_tool" && status = error }
log lines for one trace (LogQL){service_name="docs-copilot"} |= "4bf92f3577b34da6a3ce929d0e0e4736"
warning rate by message (LogQL)sum by (detected_level) (count_over_time({service_name="docs-copilot"} | detected_level="warn" [5m]))
Exemplars and links

Configure the Prometheus data source with exemplars (a trace id attached to histogram samples) and the Tempo data source with trace-to-logs. Then a dot on the p95 panel opens the exact slow trace, and the trace view opens its log lines. That three-click path is the whole point of the stack.

6. A dashboard that earns its screen space

traffic & health requests / minby route error ratiomodel + tool errors p50 / p95 latencyexemplars β†’ traces p95 time to first tokenwhat users feel economics tokens / mininput Β· output Β· cached spend / hourvs budget line cost / request p95catches runaway agents cache hit ratioprompt + answer caches quality grounded rateby prompt version user feedbackπŸ‘ ratio, daily judge scoresampled, from Langfuse guardrail triggersinjection, PII, refusals

7. Alerts that wake people up for the right reasons

AlertCondition (example)Why
Latency SLO burnp95 > 8 s for 10 minusers are waiting; often a provider incident or a prompt-size regression
Model error ratio> 5% for 5 minrate limits, outages, or a bad deploy
Spend anomalyspend/hour > 2Γ— the same hour last weekrunaway agent loop, abuse, or a caching bug
Quality dropgrounded rate below its 7-day baseline βˆ’ 3Οƒ (lesson 15)the silent failure: a prompt, model or index change
Pipeline freshnessno successful index sync for 24 hanswers going stale without any error
groups:
  - name: docs-copilot
    rules:
      - alert: LLMLatencyHigh
        expr: histogram_quantile(0.95, sum by (le) (rate(gen_ai_client_operation_duration_seconds_bucket[5m]))) > 8
        for: 10m
        labels: { severity: page }
        annotations: { summary: "p95 model latency {{ $value | humanizeDuration }} (SLO 8s)" }
      - alert: LLMSpendAnomaly
        expr: sum(increase(app_cost_usd_total[1h])) > 2 * sum(increase(app_cost_usd_total[1h] offset 1w))
        for: 15m
        labels: { severity: ticket }

Alert on symptoms users feel (latency, errors, quality, spend), not on every cause. Route "page" alerts to on-call and "ticket" alerts to a queue. Every alert should link to a dashboard and a runbook.

Recap

  • Collector β†’ Prometheus / Tempo / Loki β†’ Grafana; grafana/otel-lgtm gives you all of it in one container for development.
  • Redact in the Collector as defence in depth; derive RED metrics from spans with spanmetrics.
  • histogram_quantile estimates from buckets, so choose boundaries near your SLOs.
  • Dashboards: traffic, economics, quality. Alerts on symptoms, with links from metric to trace to logs.

Checkpoint

1 Β· Your latency histogram buckets are 1 s, 5 s and 30 s, and the SLO is p95 < 4 s. Problem?
Quantile estimates are only as precise as the bucket containing them. Add boundaries around 3–5 s.
2 Β· Where is the best place to guarantee that prompt text never reaches the tracing backend, even if an app is misconfigured?
The Collector sits on every path to storage, so deleting content attributes there protects against any single app getting it wrong.