Module 3 · Agents & MCP

Tool use & agents

Advanced 20 min read Letting the model decide the next step, safely

With tool calling the model can ask your code to run a function (search, query a database, call an API) and then continue with the result. An agent is that idea in a loop: the model keeps choosing tools until it decides it is done. This is powerful and easy to overuse. This lesson builds an agent loop with the guard rails it needs, and shows when a plain workflow, where your code decides the steps, is the better engineering choice.

🚕

A satnav route or a taxi driver

A workflow is a satnav route: fixed turns, predictable time, easy to debug, useless if the road is closed. An agent is a taxi driver who chooses the route as they go: flexible, able to handle surprises, but harder to predict, occasionally lost, and paid by the minute. Use the driver when the route really cannot be planned in advance.

1. From one call to autonomous agents

single callclassify, extract,summarise workflowsteps fixed in code(RAG is one) routermodel picks one ofa few known paths tool-loop agentmodel picks toolsuntil it is done multi-agentagents delegate toother agents more flexible → also more cost, latency, variance, and harder testing → start as far left as the problem allows

2. The agent loop, with guard rails

Here is a real agent loop over this site. The tools are genuine Python functions (search, read a section, count lessons). The loop sends the tool schemas, executes whatever the model asks for, returns results and errors to the model, and stops at a step limit. The model's decisions are scripted through the course's fake server so the example runs offline. With OpenAI() the model makes them itself.

import json

from fake_openai import fake_client
from mini_rag import HybridIndex
from site_docs import COURSE_NAMES, load_sections

docs = load_sections()
index = HybridIndex(docs)

# ---------------------------------------------------------------- tools: plain, typed, small outputs
def search_docs(query: str, course: str | None) -> list[dict]:
    hits = index.search(query, k=3, where={"course": course} if course else None)
    return [{"url": h.doc["url"], "section": h.doc["section"]} for h in hits]

def read_section(url: str) -> str:
    for d in docs:
        if d["url"] == url:
            return d["text"][:1200]                                    # cap what goes back into context
    raise KeyError(f"no section {url!r}; use a url returned by search_docs")

def count_lessons(course: str) -> int:
    if course not in COURSE_NAMES:
        raise ValueError(f"unknown course {course!r}; valid: {sorted(COURSE_NAMES)}")
    return len({d["url"].split("#")[0] for d in docs if d["course"] == course})

TOOLS = {f.__name__: f for f in (search_docs, read_section, count_lessons)}
SCHEMAS = [
    {"type": "function", "name": "search_docs", "strict": True,
     "description": "Search the course lessons. Returns up to 3 {url, section} hits.",
     "parameters": {"type": "object", "additionalProperties": False, "required": ["query", "course"],
                    "properties": {"query": {"type": "string"},
                                   "course": {"type": ["string", "null"], "description": "course folder or null"}}}},
    {"type": "function", "name": "read_section", "strict": True,
     "description": "Read one lesson section by the url returned from search_docs.",
     "parameters": {"type": "object", "additionalProperties": False, "required": ["url"],
                    "properties": {"url": {"type": "string"}}}},
    {"type": "function", "name": "count_lessons", "strict": True,
     "description": "Number of lessons in a course folder, e.g. 'mysql'.",
     "parameters": {"type": "object", "additionalProperties": False, "required": ["course"],
                    "properties": {"course": {"type": "string"}}}},
]

# ---------------------------------------------------------------- the model's decisions (scripted)
client = fake_client(script=[
    {"tool_call": "count_lessons", "arguments": {"course": "MySQL"}},           # wrong name: gets an error
    {"tool_call": "count_lessons", "arguments": {"course": "mysql"}},           # corrects itself
    {"tool_call": "search_docs", "arguments": {"query": "deadlock", "course": "mysql"}},
    {"tool_call": "read_section", "arguments": {"url": "mysql/09-locking.html#deadlock"}},
    "The MySQL course has 22 lessons. InnoDB detects deadlocks and rolls one transaction back "
    "with error 1213, so the application should retry it (mysql/09-locking.html#deadlock).",
])

# ---------------------------------------------------------------- the loop
def run_agent(question, max_steps=8):
    items = [{"role": "user", "content": question}]
    for step in range(1, max_steps + 1):
        r = client.responses.create(model="gpt-5.6-luna", input=items, tools=SCHEMAS,
                                    instructions="Answer questions about the course lessons using the tools.")
        calls = [o for o in r.output if o.type == "function_call"]
        if not calls:
            return r.output_text, step
        items += r.output                                   # keep the model's tool calls in the history
        for call in calls:
            args = json.loads(call.arguments)
            try:
                result = TOOLS[call.name](**args)
                shown = str(result)[:70]
            except Exception as e:                           # errors go back to the model as data
                result = {"error": f"{type(e).__name__}: {e}"}
                shown = result["error"][:70]
            print(f"step {step}: {call.name}({args}) -> {shown}")
            items.append({"type": "function_call_output", "call_id": call.call_id,
                          "output": json.dumps(result)})
    raise RuntimeError(f"agent did not finish within {max_steps} steps")

answer, steps = run_agent("How many lessons does the MySQL course have, and how does it handle deadlocks?")
print(f"\nfinal answer after {steps} model calls:\n{answer}")
step 1: count_lessons({'course': 'MySQL'}) -> ValueError: unknown course 'MySQL'; valid: ['ai-engineering', 'core-py step 2: count_lessons({'course': 'mysql'}) -> 22 step 3: search_docs({'query': 'deadlock', 'course': 'mysql'}) -> [{'url': 'mysql/09-locking.html#deadlock', 'section': '4. Deadlocks'}, step 4: read_section({'url': 'mysql/09-locking.html#deadlock'}) -> [diagram: Deadlock cycle: transaction 1 holds a lock on row A and wait final answer after 5 model calls: The MySQL course has 22 lessons. InnoDB detects deadlocks and rolls one transaction back with error 1213, so the application should retry it (mysql/09-locking.html#deadlock).

Errors are data

The bad course name came back as a message listing the valid values, and the model corrected itself. Raising the exception would have killed the whole run.

Hard limits

max_steps stops runaway loops. Production agents also cap total tokens, wall-clock time and cost per run.

Small outputs

Tools return only what the model needs (3 hits, 1,200 characters). Every returned token is re-read on every later step.

3. Designing tools a model can use well

RuleWhy
Clear names and descriptions that say when to use the toolthe description is the documentation the model reads
Strict JSON schemas with enums for closed setsfewer invalid calls; validation for free
Few, coarse tools rather than many tiny onesevery extra tool makes choosing harder and costs prompt tokens
Helpful error messages that name the valid valueslets the model recover, as in the run above
Read-only by default; writes need confirmationa wrong search is harmless, a wrong refund is not
Enforce permissions inside the tool, from the user's identitynever trust arguments the model chose to decide authorisation
Idempotent writes (idempotency keys)agents retry, and a retried payment must not charge twice

4. Often you want a workflow instead

Many "agents" in production are really workflows with one or two model decisions. They are cheaper, faster, testable step by step, and they fail in predictable ways. The common patterns:

Chain

Step A's output feeds step B: extract → validate → summarise.

Router

A cheap call classifies the request; code sends it down a fixed path.

Parallel

Run independent calls at once (several sources, several checks) and merge.

Orchestrator–workers

One call plans subtasks, workers do them, one call combines them.

Evaluator–optimiser

Generate, critique against criteria, revise. Stop after N rounds.

Agent

Only when the number and order of steps truly cannot be known in advance.

from typing import Literal

from pydantic import BaseModel
from fake_openai import fake_client

class Route(BaseModel):
    route: Literal["docs", "stats", "out_of_scope"]

client = fake_client(script=['{"route": "docs"}', '{"route": "stats"}', '{"route": "out_of_scope"}'])
ROUTER = ('Classify the question. "docs": how/why questions about the lessons. "stats": counts or '
          'lists of lessons. "out_of_scope": anything else. JSON: {"route": ...}')

def handle(question):
    r = client.responses.create(model="gpt-5.6-luna", instructions=ROUTER, input=question)
    route = Route.model_validate_json(r.output_text).route
    if route == "docs":
        return route, "→ RAG pipeline (lesson 09)"
    if route == "stats":
        return route, "→ deterministic Python over the lesson metadata, no LLM needed"
    return route, "→ polite refusal, no further model calls"

for q in ["Why is OFFSET slow?", "How many SQL lessons are there?", "Write me a poem about cats"]:
    print(f"{q:36}", *handle(q))
Why is OFFSET slow? docs → RAG pipeline (lesson 09) How many SQL lessons are there? stats → deterministic Python over the lesson metadata, no LLM needed Write me a poem about cats out_of_scope → polite refusal, no further model calls

5. Why agents fail, and how to contain it

for per_step in (0.99, 0.95, 0.90):
    print(f"per-step success {per_step:.0%}: " +
          "  ".join(f"{n} steps → {per_step ** n:.0%}" for n in (1, 5, 10, 20)))
per-step success 99%: 1 steps → 99% 5 steps → 95% 10 steps → 90% 20 steps → 82% per-step success 95%: 1 steps → 95% 5 steps → 77% 10 steps → 60% 20 steps → 36% per-step success 90%: 1 steps → 90% 5 steps → 59% 10 steps → 35% 20 steps → 12%

A 95%-reliable step sounds good until you chain ten of them. That is the core reason to keep agents short, give them few, well-designed tools, and prefer workflows.

FailureContainment
Loops (same call again and again)step, token and time budgets; detect repeated identical calls
Context bloat from tool outputstruncate and summarise outputs; drop old tool results from history
Wrong or dangerous actionread-only tools by default; human approval for writes; permission checks inside tools
Prompt injection via tool resultstreat tool output as data; never let it widen permissions (lesson 20)
Hard to debugtrace every step (tool, arguments, result, tokens), which is Module 4
Quality driftevaluate trajectories (right tools, sensible order) as well as final answers

6. Agent frameworks

The loop above is about 30 lines. Frameworks add state management, persistence (resume a run after a crash), human-in-the-loop pauses, streaming and multi-agent handoffs. The OpenAI Agents SDK is covered in OpenAI SDK lesson 17. LangGraph models an application as a graph of steps over shared state. This is its documented core API, not run here:

from typing import TypedDict

from langgraph.graph import END, START, StateGraph

class State(TypedDict):
    question: str
    sources: list
    answer: str

def retrieve(state: State) -> dict:
    return {"sources": [h.doc for h in index.search(state["question"], k=4)]}

def generate(state: State) -> dict:
    return {"answer": rag.answer(state["question"]).text}

graph = StateGraph(State)
graph.add_node("retrieve", retrieve)
graph.add_node("generate", generate)
graph.add_edge(START, "retrieve")
graph.add_edge("retrieve", "generate")
graph.add_edge("generate", END)
app = graph.compile()          # add a checkpointer to persist state and pause for human approval
print(app.invoke({"question": "What is a covering index?"})["answer"])

The next lesson, MCP, answers a different question: not how to run the loop, but how to plug tools into any agent or app through one standard protocol.

Recap

  • An agent is a tool loop where the model chooses the next step. A workflow is code choosing it.
  • Guard the loop: step, token and time limits; errors returned as data; small tool outputs.
  • Design tools carefully: clear descriptions, strict schemas, read-only by default, permissions enforced inside.
  • Prefer workflows when the steps are knowable; errors compound with every autonomous step.

Checkpoint

1 · A tool raises an exception because the model passed an invalid date. Best handling?
Models recover well from clear error messages. An empty result hides the problem; crashing wastes the work done so far.
2 · Support requests always follow: look up the order → check the refund policy → draft a reply. Agent or workflow?
Known steps belong in code: cheaper, predictable and testable. Let the model do the language work within each step.
3 · Where must "only managers can issue refunds over $500" be enforced?
Prompts can be ignored or injected. Authorisation is code, and it runs on every call.