Module 3 · Agents & MCP

Build an MCP server in Python

Advanced 24 min read The official SDK v2: tools, resources, prompts, lifespan, tests

With the official MCP Python SDK, a server is a set of type-hinted Python functions with decorators. The SDK turns the type hints into JSON Schemas, validates arguments, runs the protocol over stdio or HTTP, and handles version negotiation. This lesson builds the course's docs-search server: the one the project exposes to Claude, IDEs and your own agents. It also covers how to test it, connect it to a real host, and deploy it.

What was and was not run

The SDK code in this lesson follows the SDK v2 documentation (September 2026) but was not executed here, because installing the mcp package on this machine needs your go-ahead. The schema example in section 2 was run: it uses Pydantic, which is how the SDK builds schemas. The protocol itself was shown running in lesson 12.

🧰

A well-labelled toolbox

Each tool has a clear label (name and description), a shape that only fits the right bolts (the input schema), and a note about what it produces (the output schema). The SDK is the workshop that prints the labels from your code, so they can never drift out of date.

1. Install and the smallest server

uv add "mcp[cli]"          # or: pip install "mcp[cli]"
# still on v1 code (FastMCP)? pin it until you migrate:  pip install "mcp>=1.28,<2"
from mcp.server import MCPServer

mcp = MCPServer("Demo")

@mcp.tool()
def add(a: int, b: int) -> int:
    """Add two numbers."""
    return a + b

@mcp.resource("greeting://{name}")
def greeting(name: str) -> str:
    """Greet someone by name."""
    return f"Hello, {name}!"

if __name__ == "__main__":
    mcp.run()                  # stdio by default; mcp.run(transport="streamable-http", port=3001) for HTTP
v1 → v2

In SDK v1 the class was FastMCP (from mcp.server.fastmcp import FastMCP). In v2 it is MCPServer from mcp.server, with the same decorator style. Many tutorials online still show v1, so check which major version a snippet targets.

2. Type hints become schemas

The model only sees names, descriptions and schemas, so they are the most important "code" in a server. The SDK builds them from your function signature with Pydantic. Here is the same mechanism run directly, so you can see exactly what a host receives:

import inspect
import json
from typing import Annotated

from pydantic import BaseModel, Field, TypeAdapter, ValidationError, create_model

class Hit(BaseModel):
    url: str = Field(description="Link to the lesson section")
    title: str
    score: float

def search_docs(
    query: Annotated[str, Field(description="What to look for, in plain words")],
    course: Annotated[str | None, Field(description="Limit to one course folder, e.g. 'sql'")] = None,
    k: Annotated[int, Field(ge=1, le=10, description="How many hits")] = 5,
) -> list[Hit]:
    """Search the course lessons. Use this before answering questions about the courses."""

sig = inspect.signature(search_docs)
Args = create_model("search_docsArguments", **{
    name: (p.annotation, ... if p.default is inspect.Parameter.empty else p.default)
    for name, p in sig.parameters.items()})

input_schema = Args.model_json_schema()
print("description:", inspect.getdoc(search_docs))
print("required:   ", input_schema["required"])
print("k:          ", json.dumps(input_schema["properties"]["k"]))
print("output:     ", json.dumps(TypeAdapter(list[Hit]).json_schema())[:110], "…")

try:
    Args.model_validate({"query": "joins", "k": 50})          # the SDK validates before your code runs
except ValidationError as e:
    print("rejected:   ", e.errors()[0]["loc"], e.errors()[0]["msg"])
description: Search the course lessons. Use this before answering questions about the courses. required: ['query'] k: {"default": 5, "description": "How many hits", "maximum": 10, "minimum": 1, "title": "K", "type": "integer"} output: {"$defs": {"Hit": {"properties": {"url": {"description": "Link to the lesson section", "title": "Url", "type": … rejected: ('k',) Input should be less than or equal to 10

3. The docs server, properly built

A real server loads expensive things once (the lifespan), validates inputs, returns structured output, reports tool failures with ToolError (so the model can recover), and exposes data as resources and a workflow as a prompt:

"""The learn-with-project docs search, as an MCP server."""
from collections.abc import AsyncIterator
from contextlib import asynccontextmanager
from dataclasses import dataclass
from typing import Annotated

from pydantic import BaseModel, Field
from mcp.server import MCPServer
from mcp.server.mcpserver import Context
from mcp.server.mcpserver.exceptions import ResourceNotFoundError, ToolError

from mini_rag import HybridIndex
from site_docs import COURSE_NAMES, load_sections

@dataclass
class AppContext:
    index: HybridIndex
    docs: list[dict]

@asynccontextmanager
async def lifespan(server: MCPServer) -> AsyncIterator[AppContext]:
    docs = load_sections()
    index = HybridIndex(docs)                   # built once at startup, shared by every request
    try:
        yield AppContext(index=index, docs=docs)
    finally:
        pass                                    # close pools, flush telemetry, etc.

mcp = MCPServer("docs-copilot", lifespan=lifespan,
                instructions="Search and read the lessons of the learn-with-project site. "
                             "Search first, then read the best section, then answer with its URL.")

class Hit(BaseModel):
    url: str = Field(description="Lesson section, usable with the lesson:// resource")
    title: str
    section: str
    snippet: str = Field(description="The first 200 characters of the section")

@mcp.tool()
def search_docs(
    query: Annotated[str, Field(description="What to look for, in plain words")],
    ctx: Context[AppContext],
    course: Annotated[str | None, Field(description="Limit to one course folder, e.g. 'sql'")] = None,
    k: Annotated[int, Field(ge=1, le=10)] = 5,
) -> list[Hit]:
    """Search the course lessons (keyword + vector hybrid). Returns the best matching sections."""
    if course is not None and course not in COURSE_NAMES:
        raise ToolError(f"Unknown course {course!r}. Valid values: {sorted(COURSE_NAMES)}")
    index = ctx.request_context.lifespan_context.index
    hits = index.search(query, k=k, where={"course": course} if course else None)
    return [Hit(url=h.doc["url"], title=h.doc["title"], section=h.doc["section"],
                snippet=h.doc["text"][:200]) for h in hits]

@mcp.resource("lesson://{course}/{name}")
def lesson(course: str, name: str, ctx: Context) -> str:
    """The full text of one lesson, e.g. lesson://sql/06-order-limit.html"""
    docs = ctx.request_context.lifespan_context.docs
    parts = [d["text"] for d in docs if d["course"] == course and d["url"].split("#")[0].endswith(name)]
    if not parts:
        raise ResourceNotFoundError(f"No lesson {course}/{name}")
    return "\n\n".join(parts)

@mcp.prompt()
def explain(topic: str) -> str:
    """Explain a topic from the courses for a beginner, with citations."""
    return (f"Use search_docs to find lessons about {topic!r}. Read the best section, then explain "
            f"it for a beginner in under 150 words and cite the lesson URL.")

if __name__ == "__main__":
    mcp.run()
Feature usedWhy it matters
lifespan=the index is built once, not on every call. Its finally is your shutdown hook.
ctx: Context[AppContext]injected by the SDK and hidden from the schema; gives typed access to lifespan state (plus progress and elicitation)
Return type list[Hit]becomes the tool's outputSchema, and results arrive as structuredContent
ToolErrora tool execution error (isError: true) whose message the model reads. Never return an error string: that looks like success.
instructionsreturned by discovery: guidance the host can give its model about using this server

4. Try it: the Inspector and real hosts

uv run mcp dev docs_server.py                                   # opens the MCP Inspector in a browser
uv run mcp run docs_server.py --transport streamable-http       # serve over HTTP at http://localhost:8000/mcp
claude mcp add docs -- uv run --with "mcp[cli]" mcp run /absolute/path/to/docs_server.py   # Claude Code
uv run mcp install docs_server.py                               # writes the Claude Desktop config entry
{
  "mcpServers": {
    "docs": {
      "command": "/absolute/path/to/uv",
      "args": ["run", "--frozen", "--with", "mcp[cli]", "mcp", "run", "/absolute/path/to/docs_server.py"],
      "env": {"LWP_SITE": "/absolute/path/to/learn-with-project/pyspark"}
    }
  }
}

The host starts the server in its own process with a minimal environment, which is why the paths are absolute and variables go in env. After editing, fully quit and restart the host.

5. Test it in memory

Client(mcp) connects to the server object directly: no subprocess, no port. This makes MCP servers as easy to test as any function:

import pytest
from mcp import Client

from docs_server import mcp

@pytest.fixture
def anyio_backend():
    return "asyncio"

@pytest.fixture
async def client():
    async with Client(mcp, raise_exceptions=True) as c:
        yield c

@pytest.mark.anyio
async def test_search_returns_structured_hits(client):
    result = await client.call_tool("search_docs", {"query": "keyset pagination", "course": "sql"})
    assert not result.is_error
    assert result.structured_content["result"][0]["url"].startswith("sql/")

@pytest.mark.anyio
async def test_unknown_course_is_a_tool_error_not_a_crash(client):
    result = await client.call_tool("search_docs", {"query": "x", "course": "cobol"})
    assert result.is_error
    assert "Valid values" in result.content[0].text

Per the SDK docs, scalars, lists, tuples and unions are wrapped as {"result": ...} in structured content, which is why the test reads ["result"]. Pydantic models, TypedDicts, dataclasses and dict[str, ...] are already JSON objects and come back as they are.

6. Design rules for servers people actually use

RuleReason
Few tools, shaped around user goals ("search_docs", "read_lesson"), not one per API endpointevery tool costs prompt tokens and makes the model's choice harder
Descriptions that say when to use a tool and what comes backthe model reads them as documentation
Small outputs: snippets and URLs, with a resource or a second tool for full texttool output is re-read on every later step of an agent
Stable, deterministic tool orderthe spec recommends it: hosts cache tool lists, and prompt caching needs identical prefixes
No hidden per-connection state; return explicit handles (basket_id) insteadthe 2026-07-28 protocol is stateless, and requests may arrive on any connection or replica
Read-only first; writes explicit, idempotent and authorised inside the toolhosts show confirmations, but your server is the last line of defence (lesson 14)

7. Serve it over HTTP, inside an existing app

mcp.streamable_http_app() returns an ASGI app (Starlette), so uvicorn or any ASGI host can run it, and you can mount it in a FastAPI or Starlette service. The one line everyone forgets: a mounted app's own lifespan never runs, so the parent app must start the MCP session manager.

from collections.abc import AsyncIterator
from contextlib import asynccontextmanager

from starlette.applications import Starlette
from starlette.routing import Mount

from docs_server import mcp

mcp_app = mcp.streamable_http_app()              # build first: creates mcp.session_manager

@asynccontextmanager
async def lifespan(app: Starlette) -> AsyncIterator[None]:
    async with mcp.session_manager.run():         # without this, the first /mcp request fails
        yield

app = Starlette(routes=[Mount("/", app=mcp_app)], lifespan=lifespan)   # your routes go BEFORE Mount("/")
# uvicorn app:app --host 127.0.0.1 --port 8000

Host allowlist

Out of the box the app only answers requests addressed to localhost. Behind a real hostname, configure transport_security, or every request gets a 421.

Auth

Remote servers are OAuth 2.1 resource servers: validate bearer tokens and their audience on every request (lesson 14).

Observability

The SDK ships OpenTelemetry support, and the spec reserves traceparent in _meta, so host and server spans join one trace (lesson 17).

Recap

  • MCPServer + decorators: @mcp.tool(), @mcp.resource(uri_template), @mcp.prompt(). Type hints are the schema.
  • lifespan for startup work, Context for state and progress, ToolError for recoverable failures, return types for structured output.
  • Develop with mcp dev and the Inspector; test in memory with Client(mcp); connect hosts with absolute paths.
  • Deploy as ASGI with streamable_http_app(), and start the session manager in the parent's lifespan.

Checkpoint

1 · A tool catches an exception and returns "Error: database unavailable". What is wrong?
Hosts, agents and dashboards rely on isError. A returned string hides the failure from everything except the model's reading of it.
2 · Your server works under Claude Code but Claude Desktop cannot start it. The config uses "python server.py". Likely cause?
Hosts do not start servers from your shell, so there is no current directory, virtualenv or exported variables unless you configure them.
3 · Where should an expensive index be built?
The lifespan runs once around the server's life, with a clean shutdown path. Import-time work also runs in tests, tooling, and every process that merely imports the module.