AI engineering: build LLM systems you can see inside.
LLM, RAG, MCP, agents, OpenTelemetry, Langfuse, Grafana: the vocabulary arrives all at once. This course takes one piece at a time. First what a model really does, then how to give it your knowledge (RAG) and your tools (MCP), then how to observe, evaluate and control the whole thing in production. All examples are in Python.
The one picture this course is built on
Every production LLM application, from a support bot to a coding agent, has the same four parts. Each module of this course is one box in this diagram.
You write Python and have called an API before. The OpenAI SDK course is a good companion, since this course uses that SDK for model calls. The ideas apply to any provider. Most examples run offline, with no API key: model calls go through the real OpenAI SDK to a small fake server, the data is this website's own lessons, and telemetry goes to in-memory OpenTelemetry exporters. Running the full observability stack needs Docker, and Langfuse needs its cloud or a self-hosted server. Examples target Python 3.11+.
Libraries such as the MCP SDK, Langfuse and the OpenTelemetry GenAI conventions release often. The lessons separate concepts (stable) from API details (check the current docs), and name the versions that were current when they were written.
Pick a path
What are all these words?
Lessons 01, 05, 11, 12, 15. What an LLM, RAG, an agent, MCP and observability actually are.
I need to build RAG
Lessons 02, 05–10. Embeddings, vector search, chunking, retrieval and evaluation.
I need MCP
Lessons 11–14. Tool use, the protocol, a Python server and client, and security.
I run LLM apps in production
Lessons 15–20. OpenTelemetry, Langfuse, Grafana, guardrails and cost control.
Ships with a hands-on project
🛠 An observable docs copilot
Index every lesson on this site with hybrid search, answer questions with verified
citations over a FastAPI HTTP API, and expose search_docs and ask as an
MCP server that Claude Desktop or any MCP client can use. Every request is traced with
OpenTelemetry into a Grafana dashboard (and optionally
Langfuse), and the whole system is evaluated against a 41-question golden set. It runs
offline with no API key; one docker compose up adds the Collector and the Grafana stack.