How an AI agent works — a program that sends the whole conversation to a model, runs the tools the model asks for, and loops until the model answers — with the JSON of every request and reply in the Anthropic Messages or OpenAI Chat Completions format, the context window filling up, and a skill loaded as a tool.
An AI agent is a program that runs a loop around a model.
- The user types a prompt.
- The model makes every decision: call a tool, or answer.
- The agent carries out the decisions, and keeps the conversation. A program around a model is also called its
harness.
- The top bar picks the API: Anthropic Messages or OpenAI Chat Completions. The JSON, the roles and the
words follow it.
Reading the drawing
- The user, left: the chat, with the prompts and the answers.
- The agent, middle: the loop as a graph, our instructions (the system prompt),
messages[], and the tools.- A dashed message has not been sent to the model yet.
- Under the tools: what stands behind three of them, outside the agent.
- The model, top right: what it is doing, and its context window.
- A request fills the empty window: a copy of everything the agent sends flies there as one chunk.
- A reply is written at the end of the window; once it leaves, the window is wiped.
- The wire, bottom middle: the JSON of the step.
- In a request, the messages sent before are dimmed and the new ones lit.
- The tool definitions are cut to one line each: they are the same in every request.
The system prompt: our instructions
The system prompt is ours: we write it into the agent, and the agent sends it with every request.
- The model does not write it. It reads it first, before the conversation.
- It sets the model's job; each prompt sets a task.
- The agent fills in today's date when it starts, since the model has no clock.
- The user never sees it in the chat.
The agent loop
One pass around the loop is one request to the model.
- Call the model with the whole conversation.
- Read the reply. It is one of two things:
- a tool call: run the tool, append its result, call the model again;
- an answer: hand it to the user. The turn ends.
- The model decides which way the loop goes; the agent's code only follows.
The first prompt takes three requests: a forecast, a reminder, then the answer.
Roles: who wrote each message
Every message in messages[] carries a role, and the two APIs place the same conversation differently.
| In the conversation | Anthropic Messages | OpenAI Chat Completions |
|---|
| The instructions | the request's system field | a first message, role developer |
| The prompt | role user | role user |
| The model's reply | role assistant | role assistant |
| A tool call | a tool_use block in the reply's content | an entry in the reply's tool_calls |
| The call's arguments | input, a JSON object | arguments, a JSON object written as a string |
| A tool's result | a tool_result block, in a user message | a message of its own, role tool |
| What links a result to its call | tool_use_id | tool_call_id |
| The reply is a tool call | stop_reason: tool_use | finish_reason: tool_calls |
| The reply is the answer | stop_reason: end_turn | finish_reason: stop |
A tool call is JSON in the model's reply, and the model runs nothing.
- The agent declares each tool in every request: a name, a description and a JSON Schema of its parameters.
- The model chooses by the descriptions, and writes a call: the tool's name, an id and the arguments.
- The agent runs the code behind the name. The model never sees that code.
- The result goes back as a message carrying the call's id, so the model can match the two.
The context window
The context window is everything one request puts in front of the model.
- The model keeps nothing between requests.
messages[] is its only memory, so the agent sends all of it
every time. - It fills with:
- the tool definitions, in every request, used or not;
- our instructions, the system prompt;
- every message so far, tool results included;
- the reply, written one token at a time at its end.
- A token is a piece of a word. The page counts about four characters of JSON to a token; the exact count, and
the layout inside the window, are the provider's.
- The window has a fixed size per model, hundreds of thousands of tokens today. A long session fills it, and the
agent then has to drop or summarise old messages.
A skill reaches the model as a tool like any other: a name, a description and parameters.
- A function: code inside the agent.
get_forecast, add_reminder. - A skill: a folder with a
SKILL.md, instructions for one kind of task. load_skill.- Only its name and description sit in the context window beforehand.
- Its instructions enter the window when the model loads it.