Build an MCP server in Python
With the official MCP Python SDK, a server is a set of type-hinted Python functions with decorators. The SDK turns the type hints into JSON Schemas, validates arguments, runs the protocol over stdio or HTTP, and handles version negotiation. This lesson builds the course's docs-search server: the one the project exposes to Claude, IDEs and your own agents. It also covers how to test it, connect it to a real host, and deploy it.
The SDK code in this lesson follows the SDK v2 documentation (September 2026) but was not executed
here, because installing the mcp package on this machine needs your go-ahead. The
schema example in section 2 was run: it uses Pydantic, which is how the SDK builds schemas. The
protocol itself was shown running in lesson 12.
A well-labelled toolbox
Each tool has a clear label (name and description), a shape that only fits the right bolts (the input schema), and a note about what it produces (the output schema). The SDK is the workshop that prints the labels from your code, so they can never drift out of date.
1. Install and the smallest server
uv add "mcp[cli]" # or: pip install "mcp[cli]" # still on v1 code (FastMCP)? pin it until you migrate: pip install "mcp>=1.28,<2"
from mcp.server import MCPServer
mcp = MCPServer("Demo")
@mcp.tool()
def add(a: int, b: int) -> int:
"""Add two numbers."""
return a + b
@mcp.resource("greeting://{name}")
def greeting(name: str) -> str:
"""Greet someone by name."""
return f"Hello, {name}!"
if __name__ == "__main__":
mcp.run() # stdio by default; mcp.run(transport="streamable-http", port=3001) for HTTP
In SDK v1 the class was FastMCP (from mcp.server.fastmcp import FastMCP).
In v2 it is MCPServer from mcp.server, with the same decorator style. Many
tutorials online still show v1, so check which major version a snippet targets.
2. Type hints become schemas
The model only sees names, descriptions and schemas, so they are the most important "code" in a server. The SDK builds them from your function signature with Pydantic. Here is the same mechanism run directly, so you can see exactly what a host receives:
import inspect
import json
from typing import Annotated
from pydantic import BaseModel, Field, TypeAdapter, ValidationError, create_model
class Hit(BaseModel):
url: str = Field(description="Link to the lesson section")
title: str
score: float
def search_docs(
query: Annotated[str, Field(description="What to look for, in plain words")],
course: Annotated[str | None, Field(description="Limit to one course folder, e.g. 'sql'")] = None,
k: Annotated[int, Field(ge=1, le=10, description="How many hits")] = 5,
) -> list[Hit]:
"""Search the course lessons. Use this before answering questions about the courses."""
sig = inspect.signature(search_docs)
Args = create_model("search_docsArguments", **{
name: (p.annotation, ... if p.default is inspect.Parameter.empty else p.default)
for name, p in sig.parameters.items()})
input_schema = Args.model_json_schema()
print("description:", inspect.getdoc(search_docs))
print("required: ", input_schema["required"])
print("k: ", json.dumps(input_schema["properties"]["k"]))
print("output: ", json.dumps(TypeAdapter(list[Hit]).json_schema())[:110], "…")
try:
Args.model_validate({"query": "joins", "k": 50}) # the SDK validates before your code runs
except ValidationError as e:
print("rejected: ", e.errors()[0]["loc"], e.errors()[0]["msg"])
3. The docs server, properly built
A real server loads expensive things once (the lifespan), validates inputs, returns structured
output, reports tool failures with ToolError (so the model can recover), and exposes data as
resources and a workflow as a prompt:
"""The learn-with-project docs search, as an MCP server."""
from collections.abc import AsyncIterator
from contextlib import asynccontextmanager
from dataclasses import dataclass
from typing import Annotated
from pydantic import BaseModel, Field
from mcp.server import MCPServer
from mcp.server.mcpserver import Context
from mcp.server.mcpserver.exceptions import ResourceNotFoundError, ToolError
from mini_rag import HybridIndex
from site_docs import COURSE_NAMES, load_sections
@dataclass
class AppContext:
index: HybridIndex
docs: list[dict]
@asynccontextmanager
async def lifespan(server: MCPServer) -> AsyncIterator[AppContext]:
docs = load_sections()
index = HybridIndex(docs) # built once at startup, shared by every request
try:
yield AppContext(index=index, docs=docs)
finally:
pass # close pools, flush telemetry, etc.
mcp = MCPServer("docs-copilot", lifespan=lifespan,
instructions="Search and read the lessons of the learn-with-project site. "
"Search first, then read the best section, then answer with its URL.")
class Hit(BaseModel):
url: str = Field(description="Lesson section, usable with the lesson:// resource")
title: str
section: str
snippet: str = Field(description="The first 200 characters of the section")
@mcp.tool()
def search_docs(
query: Annotated[str, Field(description="What to look for, in plain words")],
ctx: Context[AppContext],
course: Annotated[str | None, Field(description="Limit to one course folder, e.g. 'sql'")] = None,
k: Annotated[int, Field(ge=1, le=10)] = 5,
) -> list[Hit]:
"""Search the course lessons (keyword + vector hybrid). Returns the best matching sections."""
if course is not None and course not in COURSE_NAMES:
raise ToolError(f"Unknown course {course!r}. Valid values: {sorted(COURSE_NAMES)}")
index = ctx.request_context.lifespan_context.index
hits = index.search(query, k=k, where={"course": course} if course else None)
return [Hit(url=h.doc["url"], title=h.doc["title"], section=h.doc["section"],
snippet=h.doc["text"][:200]) for h in hits]
@mcp.resource("lesson://{course}/{name}")
def lesson(course: str, name: str, ctx: Context) -> str:
"""The full text of one lesson, e.g. lesson://sql/06-order-limit.html"""
docs = ctx.request_context.lifespan_context.docs
parts = [d["text"] for d in docs if d["course"] == course and d["url"].split("#")[0].endswith(name)]
if not parts:
raise ResourceNotFoundError(f"No lesson {course}/{name}")
return "\n\n".join(parts)
@mcp.prompt()
def explain(topic: str) -> str:
"""Explain a topic from the courses for a beginner, with citations."""
return (f"Use search_docs to find lessons about {topic!r}. Read the best section, then explain "
f"it for a beginner in under 150 words and cite the lesson URL.")
if __name__ == "__main__":
mcp.run()
| Feature used | Why it matters |
|---|---|
lifespan= | the index is built once, not on every call. Its finally is your shutdown hook. |
ctx: Context[AppContext] | injected by the SDK and hidden from the schema; gives typed access to lifespan state (plus progress and elicitation) |
Return type list[Hit] | becomes the tool's outputSchema, and results arrive as structuredContent |
ToolError | a tool execution error (isError: true) whose message the model reads. Never return an error string: that looks like success. |
instructions | returned by discovery: guidance the host can give its model about using this server |
4. Try it: the Inspector and real hosts
uv run mcp dev docs_server.py # opens the MCP Inspector in a browser uv run mcp run docs_server.py --transport streamable-http # serve over HTTP at http://localhost:8000/mcp claude mcp add docs -- uv run --with "mcp[cli]" mcp run /absolute/path/to/docs_server.py # Claude Code uv run mcp install docs_server.py # writes the Claude Desktop config entry
{
"mcpServers": {
"docs": {
"command": "/absolute/path/to/uv",
"args": ["run", "--frozen", "--with", "mcp[cli]", "mcp", "run", "/absolute/path/to/docs_server.py"],
"env": {"LWP_SITE": "/absolute/path/to/learn-with-project/pyspark"}
}
}
}
The host starts the server in its own process with a minimal environment, which is why the paths are
absolute and variables go in env. After editing, fully quit and restart the host.
5. Test it in memory
Client(mcp) connects to the server object directly: no subprocess, no port. This makes MCP
servers as easy to test as any function:
import pytest
from mcp import Client
from docs_server import mcp
@pytest.fixture
def anyio_backend():
return "asyncio"
@pytest.fixture
async def client():
async with Client(mcp, raise_exceptions=True) as c:
yield c
@pytest.mark.anyio
async def test_search_returns_structured_hits(client):
result = await client.call_tool("search_docs", {"query": "keyset pagination", "course": "sql"})
assert not result.is_error
assert result.structured_content["result"][0]["url"].startswith("sql/")
@pytest.mark.anyio
async def test_unknown_course_is_a_tool_error_not_a_crash(client):
result = await client.call_tool("search_docs", {"query": "x", "course": "cobol"})
assert result.is_error
assert "Valid values" in result.content[0].text
Per the SDK docs, scalars, lists, tuples and unions are wrapped as {"result": ...} in
structured content, which is why the test reads ["result"]. Pydantic models, TypedDicts,
dataclasses and dict[str, ...] are already JSON objects and come back as they are.
6. Design rules for servers people actually use
| Rule | Reason |
|---|---|
| Few tools, shaped around user goals ("search_docs", "read_lesson"), not one per API endpoint | every tool costs prompt tokens and makes the model's choice harder |
| Descriptions that say when to use a tool and what comes back | the model reads them as documentation |
| Small outputs: snippets and URLs, with a resource or a second tool for full text | tool output is re-read on every later step of an agent |
| Stable, deterministic tool order | the spec recommends it: hosts cache tool lists, and prompt caching needs identical prefixes |
No hidden per-connection state; return explicit handles (basket_id) instead | the 2026-07-28 protocol is stateless, and requests may arrive on any connection or replica |
| Read-only first; writes explicit, idempotent and authorised inside the tool | hosts show confirmations, but your server is the last line of defence (lesson 14) |
7. Serve it over HTTP, inside an existing app
mcp.streamable_http_app() returns an ASGI app (Starlette), so uvicorn or any ASGI host can
run it, and you can mount it in a FastAPI or Starlette service. The one line everyone forgets: a mounted
app's own lifespan never runs, so the parent app must start the MCP session manager.
from collections.abc import AsyncIterator
from contextlib import asynccontextmanager
from starlette.applications import Starlette
from starlette.routing import Mount
from docs_server import mcp
mcp_app = mcp.streamable_http_app() # build first: creates mcp.session_manager
@asynccontextmanager
async def lifespan(app: Starlette) -> AsyncIterator[None]:
async with mcp.session_manager.run(): # without this, the first /mcp request fails
yield
app = Starlette(routes=[Mount("/", app=mcp_app)], lifespan=lifespan) # your routes go BEFORE Mount("/")
# uvicorn app:app --host 127.0.0.1 --port 8000
Host allowlist
Out of the box the app only answers requests addressed to localhost. Behind a real hostname, configure transport_security, or every request gets a 421.
Auth
Remote servers are OAuth 2.1 resource servers: validate bearer tokens and their audience on every request (lesson 14).
Observability
The SDK ships OpenTelemetry support, and the spec reserves traceparent in _meta, so host and server spans join one trace (lesson 17).
Recap
- MCPServer + decorators:
@mcp.tool(),@mcp.resource(uri_template),@mcp.prompt(). Type hints are the schema. - lifespan for startup work, Context for state and progress, ToolError for recoverable failures, return types for structured output.
- Develop with
mcp devand the Inspector; test in memory withClient(mcp); connect hosts with absolute paths. - Deploy as ASGI with
streamable_http_app(), and start the session manager in the parent's lifespan.