Tool use & agents
With tool calling the model can ask your code to run a function (search, query a database, call an API) and then continue with the result. An agent is that idea in a loop: the model keeps choosing tools until it decides it is done. This is powerful and easy to overuse. This lesson builds an agent loop with the guard rails it needs, and shows when a plain workflow, where your code decides the steps, is the better engineering choice.
A satnav route or a taxi driver
A workflow is a satnav route: fixed turns, predictable time, easy to debug, useless if the road is closed. An agent is a taxi driver who chooses the route as they go: flexible, able to handle surprises, but harder to predict, occasionally lost, and paid by the minute. Use the driver when the route really cannot be planned in advance.
1. From one call to autonomous agents
2. The agent loop, with guard rails
Here is a real agent loop over this site. The tools are genuine Python functions (search, read a
section, count lessons). The loop sends the tool schemas, executes whatever the model asks for, returns
results and errors to the model, and stops at a step limit. The model's decisions are scripted
through the course's fake server so the example runs offline. With OpenAI() the model makes
them itself.
import json
from fake_openai import fake_client
from mini_rag import HybridIndex
from site_docs import COURSE_NAMES, load_sections
docs = load_sections()
index = HybridIndex(docs)
# ---------------------------------------------------------------- tools: plain, typed, small outputs
def search_docs(query: str, course: str | None) -> list[dict]:
hits = index.search(query, k=3, where={"course": course} if course else None)
return [{"url": h.doc["url"], "section": h.doc["section"]} for h in hits]
def read_section(url: str) -> str:
for d in docs:
if d["url"] == url:
return d["text"][:1200] # cap what goes back into context
raise KeyError(f"no section {url!r}; use a url returned by search_docs")
def count_lessons(course: str) -> int:
if course not in COURSE_NAMES:
raise ValueError(f"unknown course {course!r}; valid: {sorted(COURSE_NAMES)}")
return len({d["url"].split("#")[0] for d in docs if d["course"] == course})
TOOLS = {f.__name__: f for f in (search_docs, read_section, count_lessons)}
SCHEMAS = [
{"type": "function", "name": "search_docs", "strict": True,
"description": "Search the course lessons. Returns up to 3 {url, section} hits.",
"parameters": {"type": "object", "additionalProperties": False, "required": ["query", "course"],
"properties": {"query": {"type": "string"},
"course": {"type": ["string", "null"], "description": "course folder or null"}}}},
{"type": "function", "name": "read_section", "strict": True,
"description": "Read one lesson section by the url returned from search_docs.",
"parameters": {"type": "object", "additionalProperties": False, "required": ["url"],
"properties": {"url": {"type": "string"}}}},
{"type": "function", "name": "count_lessons", "strict": True,
"description": "Number of lessons in a course folder, e.g. 'mysql'.",
"parameters": {"type": "object", "additionalProperties": False, "required": ["course"],
"properties": {"course": {"type": "string"}}}},
]
# ---------------------------------------------------------------- the model's decisions (scripted)
client = fake_client(script=[
{"tool_call": "count_lessons", "arguments": {"course": "MySQL"}}, # wrong name: gets an error
{"tool_call": "count_lessons", "arguments": {"course": "mysql"}}, # corrects itself
{"tool_call": "search_docs", "arguments": {"query": "deadlock", "course": "mysql"}},
{"tool_call": "read_section", "arguments": {"url": "mysql/09-locking.html#deadlock"}},
"The MySQL course has 22 lessons. InnoDB detects deadlocks and rolls one transaction back "
"with error 1213, so the application should retry it (mysql/09-locking.html#deadlock).",
])
# ---------------------------------------------------------------- the loop
def run_agent(question, max_steps=8):
items = [{"role": "user", "content": question}]
for step in range(1, max_steps + 1):
r = client.responses.create(model="gpt-5.6-luna", input=items, tools=SCHEMAS,
instructions="Answer questions about the course lessons using the tools.")
calls = [o for o in r.output if o.type == "function_call"]
if not calls:
return r.output_text, step
items += r.output # keep the model's tool calls in the history
for call in calls:
args = json.loads(call.arguments)
try:
result = TOOLS[call.name](**args)
shown = str(result)[:70]
except Exception as e: # errors go back to the model as data
result = {"error": f"{type(e).__name__}: {e}"}
shown = result["error"][:70]
print(f"step {step}: {call.name}({args}) -> {shown}")
items.append({"type": "function_call_output", "call_id": call.call_id,
"output": json.dumps(result)})
raise RuntimeError(f"agent did not finish within {max_steps} steps")
answer, steps = run_agent("How many lessons does the MySQL course have, and how does it handle deadlocks?")
print(f"\nfinal answer after {steps} model calls:\n{answer}")
Errors are data
The bad course name came back as a message listing the valid values, and the model corrected itself. Raising the exception would have killed the whole run.
Hard limits
max_steps stops runaway loops. Production agents also cap total tokens, wall-clock time and cost per run.
Small outputs
Tools return only what the model needs (3 hits, 1,200 characters). Every returned token is re-read on every later step.
3. Designing tools a model can use well
| Rule | Why |
|---|---|
| Clear names and descriptions that say when to use the tool | the description is the documentation the model reads |
| Strict JSON schemas with enums for closed sets | fewer invalid calls; validation for free |
| Few, coarse tools rather than many tiny ones | every extra tool makes choosing harder and costs prompt tokens |
| Helpful error messages that name the valid values | lets the model recover, as in the run above |
| Read-only by default; writes need confirmation | a wrong search is harmless, a wrong refund is not |
| Enforce permissions inside the tool, from the user's identity | never trust arguments the model chose to decide authorisation |
| Idempotent writes (idempotency keys) | agents retry, and a retried payment must not charge twice |
4. Often you want a workflow instead
Many "agents" in production are really workflows with one or two model decisions. They are cheaper, faster, testable step by step, and they fail in predictable ways. The common patterns:
Chain
Step A's output feeds step B: extract → validate → summarise.
Router
A cheap call classifies the request; code sends it down a fixed path.
Parallel
Run independent calls at once (several sources, several checks) and merge.
Orchestrator–workers
One call plans subtasks, workers do them, one call combines them.
Evaluator–optimiser
Generate, critique against criteria, revise. Stop after N rounds.
Agent
Only when the number and order of steps truly cannot be known in advance.
from typing import Literal
from pydantic import BaseModel
from fake_openai import fake_client
class Route(BaseModel):
route: Literal["docs", "stats", "out_of_scope"]
client = fake_client(script=['{"route": "docs"}', '{"route": "stats"}', '{"route": "out_of_scope"}'])
ROUTER = ('Classify the question. "docs": how/why questions about the lessons. "stats": counts or '
'lists of lessons. "out_of_scope": anything else. JSON: {"route": ...}')
def handle(question):
r = client.responses.create(model="gpt-5.6-luna", instructions=ROUTER, input=question)
route = Route.model_validate_json(r.output_text).route
if route == "docs":
return route, "→ RAG pipeline (lesson 09)"
if route == "stats":
return route, "→ deterministic Python over the lesson metadata, no LLM needed"
return route, "→ polite refusal, no further model calls"
for q in ["Why is OFFSET slow?", "How many SQL lessons are there?", "Write me a poem about cats"]:
print(f"{q:36}", *handle(q))
5. Why agents fail, and how to contain it
for per_step in (0.99, 0.95, 0.90):
print(f"per-step success {per_step:.0%}: " +
" ".join(f"{n} steps → {per_step ** n:.0%}" for n in (1, 5, 10, 20)))
A 95%-reliable step sounds good until you chain ten of them. That is the core reason to keep agents short, give them few, well-designed tools, and prefer workflows.
| Failure | Containment |
|---|---|
| Loops (same call again and again) | step, token and time budgets; detect repeated identical calls |
| Context bloat from tool outputs | truncate and summarise outputs; drop old tool results from history |
| Wrong or dangerous action | read-only tools by default; human approval for writes; permission checks inside tools |
| Prompt injection via tool results | treat tool output as data; never let it widen permissions (lesson 20) |
| Hard to debug | trace every step (tool, arguments, result, tokens), which is Module 4 |
| Quality drift | evaluate trajectories (right tools, sensible order) as well as final answers |
6. Agent frameworks
The loop above is about 30 lines. Frameworks add state management, persistence (resume a run after a crash), human-in-the-loop pauses, streaming and multi-agent handoffs. The OpenAI Agents SDK is covered in OpenAI SDK lesson 17. LangGraph models an application as a graph of steps over shared state. This is its documented core API, not run here:
from typing import TypedDict
from langgraph.graph import END, START, StateGraph
class State(TypedDict):
question: str
sources: list
answer: str
def retrieve(state: State) -> dict:
return {"sources": [h.doc for h in index.search(state["question"], k=4)]}
def generate(state: State) -> dict:
return {"answer": rag.answer(state["question"]).text}
graph = StateGraph(State)
graph.add_node("retrieve", retrieve)
graph.add_node("generate", generate)
graph.add_edge(START, "retrieve")
graph.add_edge("retrieve", "generate")
graph.add_edge("generate", END)
app = graph.compile() # add a checkpointer to persist state and pause for human approval
print(app.invoke({"question": "What is a covering index?"})["answer"])
The next lesson, MCP, answers a different question: not how to run the loop, but how to plug tools into any agent or app through one standard protocol.
Recap
- An agent is a tool loop where the model chooses the next step. A workflow is code choosing it.
- Guard the loop: step, token and time limits; errors returned as data; small tool outputs.
- Design tools carefully: clear descriptions, strict schemas, read-only by default, permissions enforced inside.
- Prefer workflows when the steps are knowable; errors compound with every autonomous step.