Responses API vs Chat Completions
OpenAI has two APIs for generating text. Chat Completions
(client.chat.completions.create) is the older, widely copied format: a list of
messages in, one message out. The Responses API
(client.responses.create) is the newer, recommended one. It takes input items and
returns a list of typed output items, which lets it represent tool calls, reasoning and built-in
tools cleanly, and it can keep conversation state server-side. You will meet both, so learn to read both.
1. Roles: who is speaking
| Role | Written by | Use for |
|---|---|---|
developer (older: system) | you, the application | rules, persona, format, constraints. Highest priority after OpenAI's own policies |
user | the end user | the request itself. Treat it as untrusted input |
assistant | the model (earlier turns) | replaying conversation history |
In the Responses API, the instructions parameter is a shortcut for a developer message, and it
is not carried over automatically when you chain responses with previous_response_id.
Send it on every call.
2. The same request in both APIs
import json
from fake_openai import fake_client
client = fake_client(script=["Kochi has 214 customers.", "Kochi has 214 customers."])
history = [
{"role": "user", "content": "Which city has the most customers?"},
{"role": "assistant", "content": "Kochi."},
{"role": "user", "content": "How many?"},
]
# Responses API: instructions + input items
r = client.responses.create(model="gpt-5.6-luna",
instructions="You are a retail analyst. Be brief.",
input=history)
print("responses:", r.output_text)
# Chat Completions: everything is a message, including the system/developer prompt
c = client.chat.completions.create(model="gpt-5.6-luna",
messages=[{"role": "developer", "content": "You are a retail analyst. Be brief."},
*history])
print("chat: ", c.choices[0].message.content)
for req in client.fake.requests:
print(req["path"], "->", sorted(req["body"]))
3. Output items: why the Responses API is easier to build on
from fake_openai import fake_client
client = fake_client(script=[{"tool_calls": [("get_weather", {"city": "Kochi"}),
("get_weather", {"city": "Pune"})]}])
r = client.responses.create(model="gpt-5.6-luna", input="Weather in Kochi and Pune?",
tools=[{"type": "function", "name": "get_weather",
"parameters": {"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]}}])
for item in r.output:
if item.type == "message":
print("message:", item.content[0].text)
elif item.type == "function_call":
print("function_call:", item.name, item.arguments, item.call_id)
print("output_text is empty when the model only asked for tools:", repr(r.output_text))
4. Which one should you use?
Responses API
New projects; agents and tool use; built-in tools (web search, file search, code interpreter, MCP); reasoning models; server-side conversation state. OpenAI's recommended default.
Chat Completions
Existing code; libraries and other providers that implement "OpenAI-compatible" chat endpoints (many local model servers do). It is still supported, but new capabilities arrive in Responses first.
chat.completions.create(messages=[...]) → responses.create(input=[...], instructions="...")
system / developer message → instructions="..." (or a developer input item)
choices[0].message.content → response.output_text
tool_calls[i].function.name / .arguments → output item type "function_call": .name / .arguments
{"role": "tool", "tool_call_id": ..., ...} → {"type": "function_call_output", "call_id": ..., "output": ...}
response_format={"type": "json_schema", ...} → text={"format": {...}} (or responses.parse, lesson 06)
max_tokens / max_completion_tokens → max_output_tokens
Recap
- Roles: developer (your rules), user (untrusted request), assistant (earlier replies).
- Responses API: input items in, typed output items out;
instructionsmust be resent each call. output_textis a convenience; loop overoutputto handle tool calls and other items.- Chat Completions remains common and "OpenAI-compatible"; new features land in Responses.
Checkpoint
response.output_text is empty. What is the most likely reason?
response.output, run the tools, and send the results back (lesson 07).