Module 1 · Foundations

How the OpenAI API works

Beginner 16 min read What really happens on every call

Under every chatbot, agent and "AI feature" is the same exchange: your code sends an HTTP request containing a model name and some input, and gets back output plus a count of tokens used. That is all the model does. It has no memory between calls, it cannot run your code, and it charges per token. Understanding this exchange precisely makes every later topic (tools, memory, streaming, cost) obvious.

✉️

A letter to a brilliant stranger

Each call is a letter to an expert who has never met you and forgets you the moment they reply. If you want them to remember the conversation, you must include it in the letter. If you want them to use your database, they can only ask you to look something up, and you post the answer back. You pay per word, for the words you send and the words they write.

1. One request, end to end

Your Python code client.responses.create( model="gpt-5.6-luna", input="Summarise …") the SDK adds: API key header, retries, timeouts, JSON encoding, typed results POST /v1/responses OpenAI's servers ① check key, rate limits ② tokenise the input ③ run the model: predict output tokens one at a time ④ assemble output items ⑤ count tokens → your bill 200 OK + JSON Response id output: [items] output_text usage: input_tokens output_tokens Nothing else is shared between calls unless you send it again, or ask OpenAI to store it (lesson 15).

2. See the actual request and response

Every runnable example in this course uses fake_openai.py (lesson 02), which runs the real openai SDK against a local fake server. Nothing leaves your machine, no key is needed, and you can inspect the exact JSON the SDK sends.

import json
from fake_openai import fake_client          # real code: from openai import OpenAI; client = OpenAI()

client = fake_client(script=["Delivered on 12 Oct; refund of INR 1,500 approved."])
response = client.responses.create(
    model="gpt-5.6-luna",
    instructions="Answer in one sentence.",
    input="What happened with order 1042?",
)

sent = client.fake.requests[0]
print("POST", sent["path"])
print(json.dumps(sent["body"], indent=2))
print("---")
print(type(response).__name__, response.id)
print("output items:", [item.type for item in response.output])
print("output_text: ", response.output_text)
print("usage:       ", response.usage.input_tokens, "in /", response.usage.output_tokens, "out")
POST /v1/responses { "input": "What happened with order 1042?", "instructions": "Answer in one sentence.", "model": "gpt-5.6-luna" } --- Response resp_2 output items: ['message'] output_text: Delivered on 12 Oct; refund of INR 1,500 approved. usage: 13 in / 12 out

output_text is a convenience property that joins the text of all message items. output is the real structure: a list of items such as messages, function calls, reasoning and web search results. Later lessons use those other item types.

3. Tokens: the unit of everything

What a token is

A chunk of text, often part of a word. English averages about 4 characters or ¾ of a word per token. Other languages and code tokenise differently, often into more tokens per character.

Context window

The maximum tokens a model can consider in one call: input (instructions, history, documents, tool results) plus output. Exceed it and the request fails or must be truncated.

Billing

You pay per input token and per output token. Output usually costs several times more. Cached input can be cheaper (lesson 13).

4. The model is stateless

from fake_openai import fake_client

client = fake_client(script=["Nice to meet you, Asha!", "I don't know your name; you haven't told me."])
first = client.responses.create(model="gpt-5.6-luna", input="Hi, my name is Asha.")
second = client.responses.create(model="gpt-5.6-luna", input="What is my name?")
print(first.output_text)
print(second.output_text)
print("second request contained:", client.fake.requests[1]["body"]["input"])
Nice to meet you, Asha! I don't know your name; you haven't told me. second request contained: What is my name?

The second request contained only the second question, so a real model could not know the name either. "Chat memory" is always your code (or OpenAI's stored conversation state) resending earlier turns. Lesson 15 covers both approaches and what each costs.

5. The constraints you design around

ConstraintWhat it meansWhere it is handled
LatencySeconds per call; longer outputs take longer. Time to first token matters for chat UIsstreaming (05), smaller models (04)
Rate limitsRequests per minute and tokens per minute, per project and modelasync with limits (11), retries (12), Batch API (13)
CostTokens × price, per modelmodel choice (04), caching and budgets (13)
Non-determinismThe same input can produce different outputsstructured outputs (06), evals (18)
Knowledge cut-offThe model does not know recent or private datatools (07–08), retrieval (10, 14)

Recap

  • A call is an HTTPS POST: model + input items → output items + token usage.
  • The SDK handles auth headers, JSON, retries, timeouts and typed results.
  • Tokens measure context size and cost; output tokens usually cost more.
  • Stateless: memory, tools and knowledge are things your code provides.

Checkpoint

1 · A user says "as I mentioned earlier…" but your app sends only the latest message. What does the model see?
Each request is independent. Continuity comes from resending previous turns (or referencing stored responses), and those tokens are billed again.
2 · Which usually costs more per token?
Generating tokens one at a time costs more compute than reading input. Prices list them separately, with output typically several times the input price.