Interactive course · 21 lessons · ~7 hours

AI engineering: build LLM systems you can see inside.

LLM, RAG, MCP, agents, OpenTelemetry, Langfuse, Grafana: the vocabulary arrives all at once. This course takes one piece at a time. First what a model really does, then how to give it your knowledge (RAG) and your tools (MCP), then how to observe, evaluate and control the whole thing in production. All examples are in Python.

0 of 21 lessons complete0%

The one picture this course is built on

Every production LLM application, from a support bot to a coding agent, has the same four parts. Each module of this course is one box in this diagram.

User a question Your app (Python) prompt templates, the tool loop, guardrails, caching, routing between models modules 1, 3, 4 RAG · module 2 embed → vector search → rerank → top chunks your knowledge LLM · module 1 tokens in → tokens out (or a tool call) stateless, probabilistic MCP servers · module 3 tools, resources, prompts (DB, files, APIs, Git…) your actions Observability module 4 OpenTelemetry traces · metrics · logs Langfuse LLM traces, prompts, scores, datasets Grafana Prometheus · Tempo · Loki dashboards, alerts latency · tokens · cost · errors · answer quality The project (lesson 21) builds exactly this: RAG over this site's lessons, served as an MCP server, traced end to end.
The model is only one box. Most of the engineering, and most of the failures, happen in the boxes around it: what context you retrieve, which tools you expose, and whether you can see what happened.
Who this is for, and what you need

You write Python and have called an API before. The OpenAI SDK course is a good companion, since this course uses that SDK for model calls. The ideas apply to any provider. Most examples run offline, with no API key: model calls go through the real OpenAI SDK to a small fake server, the data is this website's own lessons, and telemetry goes to in-memory OpenTelemetry exporters. Running the full observability stack needs Docker, and Langfuse needs its cloud or a self-hosted server. Examples target Python 3.11+.

A fast-moving field

Libraries such as the MCP SDK, Langfuse and the OpenTelemetry GenAI conventions release often. The lessons separate concepts (stable) from API details (check the current docs), and name the versions that were current when they were written.

Pick a path

~1.5 hours

What are all these words?

Lessons 01, 05, 11, 12, 15. What an LLM, RAG, an agent, MCP and observability actually are.

~3 hours

I need to build RAG

Lessons 02, 05–10. Embeddings, vector search, chunking, retrieval and evaluation.

~2 hours

I need MCP

Lessons 11–14. Tool use, the protocol, a Python server and client, and security.

~2.5 hours

I run LLM apps in production

Lessons 15–20. OpenTelemetry, Langfuse, Grafana, guardrails and cost control.

Ships with a hands-on project

🛠 An observable docs copilot

Index every lesson on this site with hybrid search, answer questions with verified citations over a FastAPI HTTP API, and expose search_docs and ask as an MCP server that Claude Desktop or any MCP client can use. Every request is traced with OpenTelemetry into a Grafana dashboard (and optionally Langfuse), and the whole system is evaluated against a 41-question golden set. It runs offline with no API key; one docker compose up adds the Collector and the Grafana stack.

Open the project →

The curriculum