Introduction
An agent that has finished a task hands back two things: a result, and a record of the steps it took to get there. The record is usually long, orderly, and reassuring.
This walkthrough follows an agent through one small research task, one loop at a time. It labels each added capability where it appears, and it shows a single misreading travel from one loop into the final report. The agent, the sources, and all the figures are invented. No real product or study is described.
The Starting Point
A human-resources manager gives an AI agent this instruction:
Find three recent studies on four-day work weeks and summarize what they found.
The agent is a language model set up with four additions. It writes out reasoning before it acts. It has a search tool and a tool for opening web pages. It can save notes to a memory store outside its context window, the text it has in front of it as it writes. And it runs in a loop, deciding its own next step until it judges the task finished.
The loop it follows is the one described in a 2022 paper by researchers at Princeton University and Google: the model writes a thought, takes an action, observes the result, and repeats (Yao et al. 2022, sec. 2).
Walking Through the Loops
Step 1: The agent writes a plan
Before using any tool, the agent writes:
I need three studies from the past few years. I'll search for trials of a four-day week, choose three with reported results, open each one, note its main findings, and then write a summary.
The capability at work here is reasoning. The plan is generated text, and it becomes part of the context that every later step is predicted from. It's a sensible plan. It contains no step for checking a note against its source.
Step 2: It searches, reads the results, and picks sources
The agent writes a tool call, a request in a set format, to its search tool. The system runs the search and returns ten results as text: titles, short excerpts, and links.
The agent reads them and chooses three: a research institute's report on a trial at 41 companies, a government evaluation of a pilot in public offices, and a survey of employees at firms that had made the change.
Two capabilities are at work: tool use and retrieval. Retrieval means placing relevant outside text in front of the model so that it can work from that text instead of from its parameters, the numbers set by training (Lewis et al. 2020). The choices rest on titles and excerpts alone. Nothing so far has gone wrong.
Step 3: It opens one source and misreads a figure
The agent opens the first report. The page it retrieves includes this sentence:
Of the 41 participating companies, 62 percent reported that productivity was maintained or improved.
The agent writes its thought:
The trial found that productivity improved by 62 percent.
The source gives a share of companies, and the companies in that share reported productivity that held steady or rose. The agent has turned it into the size of a productivity gain. The two claims are far apart. Both sentences contain "productivity," "improved," and "62 percent," and the agent's version is a common kind of sentence in summaries of research.
Retrieval put the right text in front of the model. The model still generated a likely sentence, and this time the likely sentence was wrong.
Step 4: It saves a note, continues, and builds on the misreading
The agent saves a note to memory:
Study 1 (41 companies): productivity improved by 62 percent.
The capability here is memory. From this point the agent works from its note. The original page drops out of the context window to make room for the next source, and the sentence that would have exposed the error goes with it.
The agent opens the second source, which reports no measurable change in output, and the third, in which employees rated their own productivity slightly higher. It notes both accurately. Then it reasons across its three notes:
Results vary widely, from no change to a 62 percent improvement. The largest gain came from the private-sector trial, which suggests that company-led schemes are more effective than public-sector ones.
This is a compounding error. The comparison is correctly reasoned from the notes. One of the notes is wrong, so the conclusion about private and public schemes is built on a figure that doesn't exist.
Step 5: It reports success, and a reviewer finds the error
The agent writes its final message:
Task complete. I found three recent studies and summarized their findings. Productivity effects ranged from no change to a 62 percent improvement, with the strongest results in company-led trials.
The summary is tidy, and it says the task is complete. From the agent's side it is. Every step in its plan was carried out.
The manager doesn't use the summary straight away. She opens the saved record of steps, finds the note for the first study, and follows its link to the source. She reads the sentence about 41 companies and sees that the note doesn't match it. The check takes her about two minutes. She corrects the note and discards the conclusion about company-led schemes, since nothing now supports it.
Key Considerations
Each of the four additions did what it's for. Reasoning produced a workable plan. Search and retrieval brought in real sources. Memory let the agent handle more material than its context window holds. The loop carried the task through without help. The error came from the model's reading of one sentence, and each addition then passed it along.
The common mistake is to assume that a long, tidy record of work means the work is right. A record shows what the agent did. It can't show whether each step was done correctly, because the record is written by the same model that made the mistake. The 2026 International AI Safety Report notes that as tasks grow longer, agents "often lose track of their progress and cannot reliably deal with unexpected inputs" (Bengio and others 2026, sec. 1.2). A longer record means more places for an error to sit.
The record is still what made the error findable. An agent that returned only its summary would have left the manager nothing to check.
Summary
One misreading in the third loop passed through memory into the agent's reasoning and its final report, and a person comparing one note with its source caught it.
| Loop | What the agent did | Addition that made it possible | Where a check would have caught the error |
|---|
| 1 | Wrote a plan | Reasoning | A plan with a step for checking notes against sources |
| 2 | Searched and chose three sources | Tool use and retrieval | No error yet |
| 3 | Opened a source and misread a figure | Retrieval | Comparing the agent's thought with the source sentence |
| 4 | Saved the note and reasoned from it | Memory and reasoning | Rereading the source before drawing a comparison |
| 5 | Reported the task complete | The agent loop | A person opening the saved steps and one source |
- The error entered at loop 3 and was never re-examined, because later loops read the note and not the source.
- Every row from 3 onward offered a chance to catch it, and each needed a comparison with something outside the agent's own text.
- The check that worked was the cheapest one: a person opened one source.
References
- Bengio, Yoshua, and others. 2026. International AI Safety Report 2026. DSIT 2026/001. Published February 3, 2026.
- Lewis, Patrick, Ethan Perez, Aleksandra Piktus, and 9 others. 2020. "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks." arXiv:2005.11401. NeurIPS 2020.
- Yao, Shunyu, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022. "ReAct: Synergizing Reasoning and Acting in Language Models." arXiv:2210.03629. ICLR 2023.