Module 4 · Observability & operations

Langfuse: LLM tracing & evals

Advanced 22 min read Traces, scores, prompts and experiments in one place

Langfuse is an open-source LLM engineering platform (cloud or self-hosted). Where Grafana answers "is the system healthy?", Langfuse answers "what did the model see and say, what did it cost, and was it any good?" It stores traces of prompts and generations, attaches scores (user feedback, automated checks, LLM judges), versions prompts, and runs experiments over datasets. Its Python SDK (v4) is built on OpenTelemetry, so it fits the instrumentation from the last two lessons.

What was and was not run

The Langfuse SDK calls follow Langfuse's documentation (checked September 2026) but were not executed: they need the langfuse package, which is not installed here, plus an account or a self-hosted server. The experiment in section 7 was run locally, using the same function signatures Langfuse's experiment runner calls, so the same functions work unchanged against Langfuse.

📓

A lab notebook for your model

Every run is written down: exact inputs, exact outputs, which recipe version, how long it took, what it cost, and a grade in the margin. When a result looks odd you can find the page. When you change the recipe, you run the standard test set again and compare the grades side by side.

1. The data model

session (one conversation, via session_id) trace one request · user_id · tags · metadata retriever: search_docs (observation) generation: chat gpt-5.6-lunamodel · input · output · usage · cost · prompt "answer" v3 span: verify_citations scores grounded = 1 (code) user_feedback = 👍 · judge = 0.9 prompts & datasets versioned, labelled "production" golden items → experiment runs

2. Set up

pip install langfuse
export LANGFUSE_PUBLIC_KEY=pk-lf-...
export LANGFUSE_SECRET_KEY=sk-lf-...
export LANGFUSE_BASE_URL=https://cloud.langfuse.com       # or your self-hosted URL
from langfuse import get_client

langfuse = get_client()             # reads the LANGFUSE_* environment variables
assert langfuse.auth_check()        # fail fast on bad keys (development only; it makes a network call)

3. Trace with the @observe decorator

Decorate functions and their nesting becomes the trace tree. Inputs and outputs are captured automatically. Mark the model call as a generation and record model and usage, and Langfuse computes cost from its model price table:

from langfuse import get_client, observe, propagate_attributes
from openai import OpenAI

langfuse = get_client()
openai_client = OpenAI()

@observe(as_type="retriever")                                # also: span (default), tool, agent, …
def retrieve(question: str) -> list[dict]:
    return [h.doc for h in index.search(question, k=4)]

@observe(as_type="generation", name="chat")
def generate(question: str, docs: list[dict]) -> str:
    r = openai_client.responses.create(model="gpt-5.6-luna", instructions=INSTRUCTIONS,
                                       input=build_input(question, docs))
    langfuse.update_current_generation(
        model="gpt-5.6-luna",
        usage_details={"input_tokens": r.usage.input_tokens, "output_tokens": r.usage.output_tokens})
    return r.output_text

@observe()
def answer(question: str, user_id: str, session_id: str) -> str:
    with propagate_attributes(user_id=user_id, session_id=session_id,
                              tags=["docs-copilot", "prod"], metadata={"prompt_version": "answer-v3"}):
        docs = retrieve(question)
        text = generate(question, docs)
        langfuse.score_current_trace(name="grounded", value=1.0 if has_valid_citations(text, docs) else 0.0)
        return text

propagate_attributes

Sets user, session, tags and metadata for every observation created inside it, so you can filter by user or group a conversation.

Nesting is automatic

Decorated functions called inside decorated functions become children. It works with async functions too.

Scores from code

score_current_trace attaches a check your code already makes (citations valid, JSON parsed) to the trace.

4. Context managers and the OpenAI drop-in

with langfuse.start_as_current_observation(as_type="span", name="answer") as span:
    docs = retrieve(question)
    with langfuse.start_as_current_observation(as_type="generation", name="chat",
                                               model="gpt-5.6-luna") as gen:
        r = openai_client.responses.create(model="gpt-5.6-luna", input=build_input(question, docs))
        gen.update(output=r.output_text,
                   usage_details={"input_tokens": r.usage.input_tokens,
                                  "output_tokens": r.usage.output_tokens})
    span.update(output=r.output_text)
    span.score_trace(name="grounded", value=1.0)
from langfuse.openai import OpenAI          # instead of: from openai import OpenAI

client = OpenAI()
completion = client.chat.completions.create(
    model="gpt-5.6-luna",
    messages=[{"role": "user", "content": "What is keyset pagination?"}],
    name="keyset-question",                                          # Langfuse-only extras
    metadata={"langfuse_session_id": "s-42", "langfuse_user_id": "u-7", "langfuse_tags": ["faq"]},
)
Flush in short-lived processes

Like any OTel batch exporter, the SDK sends in the background. Scripts, notebooks and serverless handlers must call langfuse.flush() before exiting.

5. Scores: turning traces into quality data

SourceHowExample
Your codescore_current_trace / span.score(...)grounded, valid_json, refused
Userslangfuse.create_score(trace_id=..., name="user_feedback", value=1) from your feedback endpointthumbs up and down, "was this helpful?"
LLM judgesevaluators configured in Langfuse, run on a sample of new tracesfaithfulness, relevance, tone (validate them against humans, lesson 10)
Humansannotation queues in the UIexpert review of sampled or flagged traces
# when answering: return the trace id to the frontend
trace_id = langfuse.get_current_trace_id()

# later, when the user clicks 👍 or 👎:
langfuse.create_score(trace_id=trace_id, name="user_feedback", value=1 if thumbs_up else 0,
                      data_type="BOOLEAN", comment=user_comment)

6. Prompt management

Keep prompts in Langfuse instead of in code: versioned, labelled (production, staging), editable without a deploy, and linked to the generations that used them, so scores and costs can be compared per prompt version. This is the hosted version of lesson 04's registry:

langfuse.create_prompt(
    name="answer", type="text", labels=["production"],
    prompt="Answer using only the sources; cite [n]; say I don't know if unsure.\n"
           "{{sources}}\n<question>{{question}}</question>",
    config={"model": "gpt-5.6-luna"})

prompt = langfuse.get_prompt("answer")                 # the "production" version by default; cached client-side
text = prompt.compile(sources=sources_block, question=question)

with langfuse.start_as_current_observation(as_type="generation", name="chat",
                                           model=prompt.config["model"], prompt=prompt) as gen:
    r = openai_client.responses.create(model=prompt.config["model"], input=text)
    gen.update(output=r.output_text)

Rolling out a new prompt then means labelling version 4 as production, watching its grounded and feedback scores next to version 3, and moving the label back if they drop. Because the SDK caches prompts, keep a fallback so a Langfuse outage cannot take your app down.

7. Datasets and experiments

A dataset is your golden set stored in Langfuse. An experiment runs a task over every item, applies evaluators, and records the run so versions can be compared in the UI. You write two kinds of function, and Langfuse calls them with keyword arguments:

langfuse.create_dataset(name="docs-golden")
for question, lessons in GOLDEN:
    langfuse.create_dataset_item(dataset_name="docs-golden", input=question, expected_output=sorted(lessons))

dataset = langfuse.get_dataset("docs-golden")
result = dataset.run_experiment(name="hybrid-k5", task=task, evaluators=[hit_at_5, reciprocal_rank],
                                run_evaluators=[average("hit@5")], max_concurrency=4)
print(result.format())

The task and evaluators are ordinary Python functions, so you can develop and unit-test them locally. Here they are, run over the golden set with a tiny local runner that calls them exactly as Langfuse does. Evaluation is a local stand-in for langfuse.Evaluation:

from dataclasses import dataclass
from statistics import mean

from mini_rag import GOLDEN, HybridIndex, lessons_in
from site_docs import load_sections

@dataclass
class Evaluation:                                  # same fields as langfuse.Evaluation
    name: str
    value: float
    comment: str | None = None

index = HybridIndex(load_sections())

def task(*, item, **kwargs):                        # Langfuse passes the dataset item
    return lessons_in(index.rank(item["input"]), index.docs, k=5)

def hit_at_5(*, input, output, expected_output, **kwargs):
    hit = bool(set(expected_output) & set(output))
    return Evaluation("hit@5", float(hit), None if hit else f"got {output[0]}")

def reciprocal_rank(*, input, output, expected_output, **kwargs):
    rank = next((i for i, url in enumerate(output, 1) if url in expected_output), None)
    return Evaluation("reciprocal_rank", 1 / rank if rank else 0.0)

def average(name):                                  # a run-level evaluator
    def run_evaluator(*, item_results, **kwargs):
        return Evaluation(f"avg_{name}", mean(e.value for r in item_results for e in r["evaluations"]
                                              if e.name == name))
    return run_evaluator

def run_experiment_locally(name, data, task, evaluators, run_evaluators=()):
    item_results = []
    for item in data:
        output = task(item=item)
        evals = [ev(input=item["input"], output=output, expected_output=item["expected_output"])
                 for ev in evaluators]
        item_results.append({"item": item, "output": output, "evaluations": evals})
    print(f"experiment {name!r}: {len(item_results)} items")
    for rev in run_evaluators:
        e = rev(item_results=item_results)
        print(f"  {e.name} = {e.value:.2f}")
    misses = [r for r in item_results if r["evaluations"][0].value == 0]
    for r in misses:
        print(f"  miss: {r['item']['input'][:52]!r:56} {r['evaluations'][0].comment}")

data = [{"input": q, "expected_output": sorted(lessons)} for q, lessons in GOLDEN]
run_experiment_locally("hybrid-k5", data, task, [hit_at_5, reciprocal_rank],
                       run_evaluators=[average("hit@5"), average("reciprocal_rank")])
experiment 'hybrid-k5': 32 items avg_hit@5 = 0.84 avg_reciprocal_rank = 0.77 miss: "Why doesn't my Spark code do anything until I call c" got openai-sdk/07-function-calling.html miss: 'One key has most of the rows and a single task runs ' got databricks/12-jobs.html miss: 'running total per customer ordered by date' got course/17-spark-sql.html miss: 'find customers who never placed an order' got sql/01-relational-model.html miss: 'show tokens to the user as they are generated' got openai-sdk/13-cost-latency.html

8. Langfuse and OpenTelemetry together

Same spans, two destinations

The v4 SDK is an OpenTelemetry tracer provider underneath, and Langfuse also accepts OTLP from any OTel instrumentation. In the course project, the Collector sends traces to Tempo for operations and to Langfuse for LLM analysis.

Who looks where

On-call engineers: Grafana (latency, errors, saturation, alerts). AI engineers and product: Langfuse (prompts, outputs, scores, experiments, cost per feature).

DecisionGuidance
Cloud or self-hostCloud to start. Self-host (Docker or Kubernetes; it runs on Postgres, ClickHouse, Redis and S3-compatible storage) when data must stay in your network.
What content to sendRedact PII before it leaves your process; sample full content if volume is high; set retention.
EnvironmentsSeparate projects (or environment tags) for dev, staging and prod, so experiments never pollute production data.

Recap

  • Traces → observations (spans, generations, tools, retrievers), grouped by session and user.
  • Instrument with @observe, context managers, or the OpenAI drop-in; propagate_attributes for user, session and tags; flush in scripts.
  • Scores from code, users, judges and humans turn traces into quality data.
  • Prompts are versioned and linked to generations; experiments run tasks and evaluators over datasets. Develop the evaluators locally.

Checkpoint

1 · You want to compare user feedback between prompt versions 3 and 4. What must be true?
Linking prompt → generation and score → trace is what makes "score by prompt version" a filter instead of a data-engineering project.
2 · A nightly evaluation script finishes, but its traces never appear in Langfuse. Likely fix?
The SDK batches in the background. A process that exits immediately drops what is still queued.