Langfuse: LLM tracing & evals
Langfuse is an open-source LLM engineering platform (cloud or self-hosted). Where Grafana answers "is the system healthy?", Langfuse answers "what did the model see and say, what did it cost, and was it any good?" It stores traces of prompts and generations, attaches scores (user feedback, automated checks, LLM judges), versions prompts, and runs experiments over datasets. Its Python SDK (v4) is built on OpenTelemetry, so it fits the instrumentation from the last two lessons.
The Langfuse SDK calls follow Langfuse's documentation (checked September 2026) but were not
executed: they need the langfuse package, which is not installed here, plus an account
or a self-hosted server. The experiment in section 7 was run locally, using the same
function signatures Langfuse's experiment runner calls, so the same functions work unchanged against
Langfuse.
A lab notebook for your model
Every run is written down: exact inputs, exact outputs, which recipe version, how long it took, what it cost, and a grade in the margin. When a result looks odd you can find the page. When you change the recipe, you run the standard test set again and compare the grades side by side.
1. The data model
2. Set up
pip install langfuse export LANGFUSE_PUBLIC_KEY=pk-lf-... export LANGFUSE_SECRET_KEY=sk-lf-... export LANGFUSE_BASE_URL=https://cloud.langfuse.com # or your self-hosted URL
from langfuse import get_client langfuse = get_client() # reads the LANGFUSE_* environment variables assert langfuse.auth_check() # fail fast on bad keys (development only; it makes a network call)
3. Trace with the @observe decorator
Decorate functions and their nesting becomes the trace tree. Inputs and outputs are captured automatically. Mark the model call as a generation and record model and usage, and Langfuse computes cost from its model price table:
from langfuse import get_client, observe, propagate_attributes
from openai import OpenAI
langfuse = get_client()
openai_client = OpenAI()
@observe(as_type="retriever") # also: span (default), tool, agent, …
def retrieve(question: str) -> list[dict]:
return [h.doc for h in index.search(question, k=4)]
@observe(as_type="generation", name="chat")
def generate(question: str, docs: list[dict]) -> str:
r = openai_client.responses.create(model="gpt-5.6-luna", instructions=INSTRUCTIONS,
input=build_input(question, docs))
langfuse.update_current_generation(
model="gpt-5.6-luna",
usage_details={"input_tokens": r.usage.input_tokens, "output_tokens": r.usage.output_tokens})
return r.output_text
@observe()
def answer(question: str, user_id: str, session_id: str) -> str:
with propagate_attributes(user_id=user_id, session_id=session_id,
tags=["docs-copilot", "prod"], metadata={"prompt_version": "answer-v3"}):
docs = retrieve(question)
text = generate(question, docs)
langfuse.score_current_trace(name="grounded", value=1.0 if has_valid_citations(text, docs) else 0.0)
return text
propagate_attributes
Sets user, session, tags and metadata for every observation created inside it, so you can filter by user or group a conversation.
Nesting is automatic
Decorated functions called inside decorated functions become children. It works with async functions too.
Scores from code
score_current_trace attaches a check your code already makes (citations valid, JSON parsed) to the trace.
4. Context managers and the OpenAI drop-in
with langfuse.start_as_current_observation(as_type="span", name="answer") as span:
docs = retrieve(question)
with langfuse.start_as_current_observation(as_type="generation", name="chat",
model="gpt-5.6-luna") as gen:
r = openai_client.responses.create(model="gpt-5.6-luna", input=build_input(question, docs))
gen.update(output=r.output_text,
usage_details={"input_tokens": r.usage.input_tokens,
"output_tokens": r.usage.output_tokens})
span.update(output=r.output_text)
span.score_trace(name="grounded", value=1.0)
from langfuse.openai import OpenAI # instead of: from openai import OpenAI
client = OpenAI()
completion = client.chat.completions.create(
model="gpt-5.6-luna",
messages=[{"role": "user", "content": "What is keyset pagination?"}],
name="keyset-question", # Langfuse-only extras
metadata={"langfuse_session_id": "s-42", "langfuse_user_id": "u-7", "langfuse_tags": ["faq"]},
)
Like any OTel batch exporter, the SDK sends in the background. Scripts, notebooks and serverless
handlers must call langfuse.flush() before exiting.
5. Scores: turning traces into quality data
| Source | How | Example |
|---|---|---|
| Your code | score_current_trace / span.score(...) | grounded, valid_json, refused |
| Users | langfuse.create_score(trace_id=..., name="user_feedback", value=1) from your feedback endpoint | thumbs up and down, "was this helpful?" |
| LLM judges | evaluators configured in Langfuse, run on a sample of new traces | faithfulness, relevance, tone (validate them against humans, lesson 10) |
| Humans | annotation queues in the UI | expert review of sampled or flagged traces |
# when answering: return the trace id to the frontend
trace_id = langfuse.get_current_trace_id()
# later, when the user clicks 👍 or 👎:
langfuse.create_score(trace_id=trace_id, name="user_feedback", value=1 if thumbs_up else 0,
data_type="BOOLEAN", comment=user_comment)
6. Prompt management
Keep prompts in Langfuse instead of in code: versioned, labelled (production,
staging), editable without a deploy, and linked to the generations that used them, so
scores and costs can be compared per prompt version. This is the hosted version of lesson 04's
registry:
langfuse.create_prompt(
name="answer", type="text", labels=["production"],
prompt="Answer using only the sources; cite [n]; say I don't know if unsure.\n"
"{{sources}}\n<question>{{question}}</question>",
config={"model": "gpt-5.6-luna"})
prompt = langfuse.get_prompt("answer") # the "production" version by default; cached client-side
text = prompt.compile(sources=sources_block, question=question)
with langfuse.start_as_current_observation(as_type="generation", name="chat",
model=prompt.config["model"], prompt=prompt) as gen:
r = openai_client.responses.create(model=prompt.config["model"], input=text)
gen.update(output=r.output_text)
Rolling out a new prompt then means labelling version 4 as production, watching its grounded
and feedback scores next to version 3, and moving the label back if they drop. Because the SDK caches
prompts, keep a fallback so a Langfuse outage cannot take your app down.
7. Datasets and experiments
A dataset is your golden set stored in Langfuse. An experiment runs a task over every item, applies evaluators, and records the run so versions can be compared in the UI. You write two kinds of function, and Langfuse calls them with keyword arguments:
langfuse.create_dataset(name="docs-golden")
for question, lessons in GOLDEN:
langfuse.create_dataset_item(dataset_name="docs-golden", input=question, expected_output=sorted(lessons))
dataset = langfuse.get_dataset("docs-golden")
result = dataset.run_experiment(name="hybrid-k5", task=task, evaluators=[hit_at_5, reciprocal_rank],
run_evaluators=[average("hit@5")], max_concurrency=4)
print(result.format())
The task and evaluators are ordinary Python functions, so you can develop and unit-test them locally.
Here they are, run over the golden set with a tiny local runner that calls them exactly as Langfuse
does. Evaluation is a local stand-in for langfuse.Evaluation:
from dataclasses import dataclass
from statistics import mean
from mini_rag import GOLDEN, HybridIndex, lessons_in
from site_docs import load_sections
@dataclass
class Evaluation: # same fields as langfuse.Evaluation
name: str
value: float
comment: str | None = None
index = HybridIndex(load_sections())
def task(*, item, **kwargs): # Langfuse passes the dataset item
return lessons_in(index.rank(item["input"]), index.docs, k=5)
def hit_at_5(*, input, output, expected_output, **kwargs):
hit = bool(set(expected_output) & set(output))
return Evaluation("hit@5", float(hit), None if hit else f"got {output[0]}")
def reciprocal_rank(*, input, output, expected_output, **kwargs):
rank = next((i for i, url in enumerate(output, 1) if url in expected_output), None)
return Evaluation("reciprocal_rank", 1 / rank if rank else 0.0)
def average(name): # a run-level evaluator
def run_evaluator(*, item_results, **kwargs):
return Evaluation(f"avg_{name}", mean(e.value for r in item_results for e in r["evaluations"]
if e.name == name))
return run_evaluator
def run_experiment_locally(name, data, task, evaluators, run_evaluators=()):
item_results = []
for item in data:
output = task(item=item)
evals = [ev(input=item["input"], output=output, expected_output=item["expected_output"])
for ev in evaluators]
item_results.append({"item": item, "output": output, "evaluations": evals})
print(f"experiment {name!r}: {len(item_results)} items")
for rev in run_evaluators:
e = rev(item_results=item_results)
print(f" {e.name} = {e.value:.2f}")
misses = [r for r in item_results if r["evaluations"][0].value == 0]
for r in misses:
print(f" miss: {r['item']['input'][:52]!r:56} {r['evaluations'][0].comment}")
data = [{"input": q, "expected_output": sorted(lessons)} for q, lessons in GOLDEN]
run_experiment_locally("hybrid-k5", data, task, [hit_at_5, reciprocal_rank],
run_evaluators=[average("hit@5"), average("reciprocal_rank")])
8. Langfuse and OpenTelemetry together
Same spans, two destinations
The v4 SDK is an OpenTelemetry tracer provider underneath, and Langfuse also accepts OTLP from any OTel instrumentation. In the course project, the Collector sends traces to Tempo for operations and to Langfuse for LLM analysis.
Who looks where
On-call engineers: Grafana (latency, errors, saturation, alerts). AI engineers and product: Langfuse (prompts, outputs, scores, experiments, cost per feature).
| Decision | Guidance |
|---|---|
| Cloud or self-host | Cloud to start. Self-host (Docker or Kubernetes; it runs on Postgres, ClickHouse, Redis and S3-compatible storage) when data must stay in your network. |
| What content to send | Redact PII before it leaves your process; sample full content if volume is high; set retention. |
| Environments | Separate projects (or environment tags) for dev, staging and prod, so experiments never pollute production data. |
Recap
- Traces → observations (spans, generations, tools, retrievers), grouped by session and user.
- Instrument with
@observe, context managers, or the OpenAI drop-in;propagate_attributesfor user, session and tags; flush in scripts. - Scores from code, users, judges and humans turn traces into quality data.
- Prompts are versioned and linked to generations; experiments run tasks and evaluators over datasets. Develop the evaluators locally.