How retrieval-augmented generation works inside an AI agent — the notes cut into passages and indexed by their vectors in advance, then on every prompt the agent's own code embeds the prompt, finds the nearest passages and puts them in front of it before the model sees it, with the JSON of every request in the Anthropic Messages or OpenAI Chat Completions format.
RAG, retrieval-augmented generation, is text our code adds to a prompt before the model sees it.
- The agent searches the user's notes on every prompt, and puts the closest passages in front of it.
- The model answers from what the request holds. It still makes every decision: call a tool, or answer.
- The agent, the loop and the context window are those of the
Simple Loop AI Agent
; this page adds the search.
Reading the drawing
The user, the agent, the model and the wire are the Simple Loop AI Agent's; two things are new.
- The loop starts with search the notes. While it runs, a dashed box over
messages[] holds the prompt. - Retrieval, under the agent: the notes, the embedding model and the vector index.
- A click on a file in
notes/ opens it. - The map in the index draws each passage as a dot, coloured by its file; a click on a dot shows its text.
- The prompt's search is a diamond, with a line to each passage it found and its score.
The notes index
The index is filled by a pipeline apart from the loop, which never stops: notes are added and edited all the time.
- Cut: each note becomes passages. Here one per line; real indexers cut by size or by heading.
- Embed: an embedding model turns each passage into a vector, a list of numbers (1,024 here).
- Texts with similar meanings get vectors that lie close together.
- An embedding model is much smaller than a chat model, and writes no text.
- Store: the vector index keeps each passage with its vector, and finds the nearest vectors fast.
The search on every prompt
The search is the loop's first step: it runs once per prompt, before the first request, the same way every time.
- Embed the prompt with the same embedding model, so its vector compares with the passages'.
- Find the nearest passages: the index returns a fixed number (three here), each with a score.
- The score is the similarity of two vectors: 1 is the same meaning, near 0 unrelated.
- Similar meanings score high even with no word in common.
- Keep what scores high enough (0.6 here) and put it in front of the prompt, inside
<notes>. - Append the two to
messages[] as one user message. Only then is the model called.
The number of passages and the lowest score kept are the agent's settings, not the model's.
What the model sees
To the model, the passages are part of the user message; nothing marks them as a search.
- Our instructions say where the passages will be (
<notes>) and to use them. - The chat never shows them: the user sees only what they typed.
- Our agent keeps them in
messages[], so every later request carries them too.- That is how the second prompt, "Will it rain there?", works: its own search finds nothing, but the first
turn's passage names the place.
Retrieval and the model's decisions
The search decides nothing: it runs on every prompt, and the threshold only filters what it adds.
- The model decides everything after it: answer from the passages, or call a tool.
- A tool call does not search again: the loop goes back to "call the model", past the search.
- A search tool the model calls when it chooses is also called RAG (agentic RAG); to the model it is one
more tool call.