How the OpenAI API works
Under every chatbot, agent and "AI feature" is the same exchange: your code sends an HTTP request containing a model name and some input, and gets back output plus a count of tokens used. That is all the model does. It has no memory between calls, it cannot run your code, and it charges per token. Understanding this exchange precisely makes every later topic (tools, memory, streaming, cost) obvious.
A letter to a brilliant stranger
Each call is a letter to an expert who has never met you and forgets you the moment they reply. If you want them to remember the conversation, you must include it in the letter. If you want them to use your database, they can only ask you to look something up, and you post the answer back. You pay per word, for the words you send and the words they write.
1. One request, end to end
2. See the actual request and response
Every runnable example in this course uses fake_openai.py (lesson 02), which runs the
real openai SDK against a local fake server. Nothing leaves your machine, no key is
needed, and you can inspect the exact JSON the SDK sends.
import json
from fake_openai import fake_client # real code: from openai import OpenAI; client = OpenAI()
client = fake_client(script=["Delivered on 12 Oct; refund of INR 1,500 approved."])
response = client.responses.create(
model="gpt-5.6-luna",
instructions="Answer in one sentence.",
input="What happened with order 1042?",
)
sent = client.fake.requests[0]
print("POST", sent["path"])
print(json.dumps(sent["body"], indent=2))
print("---")
print(type(response).__name__, response.id)
print("output items:", [item.type for item in response.output])
print("output_text: ", response.output_text)
print("usage: ", response.usage.input_tokens, "in /", response.usage.output_tokens, "out")
output_text is a convenience property that joins the text of all message items.
output is the real structure: a list of items such as messages, function calls,
reasoning and web search results. Later lessons use those other item types.
3. Tokens: the unit of everything
What a token is
A chunk of text, often part of a word. English averages about 4 characters or ¾ of a word per token. Other languages and code tokenise differently, often into more tokens per character.
Context window
The maximum tokens a model can consider in one call: input (instructions, history, documents, tool results) plus output. Exceed it and the request fails or must be truncated.
Billing
You pay per input token and per output token. Output usually costs several times more. Cached input can be cheaper (lesson 13).
4. The model is stateless
from fake_openai import fake_client
client = fake_client(script=["Nice to meet you, Asha!", "I don't know your name; you haven't told me."])
first = client.responses.create(model="gpt-5.6-luna", input="Hi, my name is Asha.")
second = client.responses.create(model="gpt-5.6-luna", input="What is my name?")
print(first.output_text)
print(second.output_text)
print("second request contained:", client.fake.requests[1]["body"]["input"])
The second request contained only the second question, so a real model could not know the name either. "Chat memory" is always your code (or OpenAI's stored conversation state) resending earlier turns. Lesson 15 covers both approaches and what each costs.
5. The constraints you design around
| Constraint | What it means | Where it is handled |
|---|---|---|
| Latency | Seconds per call; longer outputs take longer. Time to first token matters for chat UIs | streaming (05), smaller models (04) |
| Rate limits | Requests per minute and tokens per minute, per project and model | async with limits (11), retries (12), Batch API (13) |
| Cost | Tokens × price, per model | model choice (04), caching and budgets (13) |
| Non-determinism | The same input can produce different outputs | structured outputs (06), evals (18) |
| Knowledge cut-off | The model does not know recent or private data | tools (07–08), retrieval (10, 14) |
Recap
- A call is an HTTPS POST: model + input items → output items + token usage.
- The SDK handles auth headers, JSON, retries, timeouts and typed results.
- Tokens measure context size and cost; output tokens usually cost more.
- Stateless: memory, tools and knowledge are things your code provides.