KnowledgeInSight
AI Literacy
0% of Course 1 complete

Module 3 · Lesson 2

Capabilities and Limits: Reasoning, tools, and agents

You'll look at what has been added around the basic language model: step-by-step reasoning, search and document retrieval, software tools, and agents that work through a task over many steps. You'll be able to say what each addition fixes and what new ways to fail it brings.

What you will be able to do

  • Describe how step-by-step reasoning, retrieval, tools, and memory extend what a model can do.

0% of this lesson · 9 items · 1h 3m total · 48m without the optional activity

Contents of this lesson9 items
  1. ReadingFrom Answering Questions to Carrying Out Tasks3 min
  2. ReadingStep-by-Step Reasoning and Reasoning Models4 min
  3. ReadingRetrieval: Supplying Documents and Search Results for a Model to Work From4 min
  4. ReadingTool Use: Calculators, Code, and Other Software a Model Can Call4 min
  5. ReadingAgents and Memory: Looping Through Plans, Actions, and Checks4 min
  6. Guided ReadingGuided Walkthrough: An Agent Works Through a Research Task, Loop by Loop7 min
  7. Guided ConversationDecide What Each Task Needs12 min
  8. Hands-on Activity · optionalCompare an Answer With and Without Search15 min
  9. Knowledge CheckReasoning, tools, and agents10 min

Reading 3 min

From Answering Questions to Carrying Out Tasks

This content reflects the field as of October 2026.

Early chatbots answered from what they had absorbed in training and nothing else. They couldn't tell you about yesterday's news, and they often got long multiplication wrong. As of October 2026, assistants from OpenAI, Google, Anthropic, Microsoft, and other developers search the web, run calculations, read files you upload, and sometimes work for many minutes before replying. Some will book, send, or edit things on your behalf.

The obvious reading is that the models got smarter. They did improve. Most of what you see in that list has a different source, though. The model underneath still does one thing, which is to predict the next token, a piece of text such as a word or part of a word. What changed is the system built around it.

The 2026 International AI Safety Report, written by an international panel of experts chaired by the computer scientist Yoshua Bengio, uses the word "scaffolding" for this. It describes scaffolding as "additional software built around general-purpose AI models that allows them to plan ahead, pursue goals, and interact with the world" (Bengio and others 2026, sec. 1.1).

Four additions account for most of the change.

  • Step-by-step reasoning. The model writes out working before it writes an answer.
  • Retrieval. Search results or documents are placed in front of the model, so it can answer from them.
  • Tools. The model can ask other software, such as a calculator, to do something and hand back the result.
  • Agents. The model is run in a loop, taking one action after another toward a goal.

None of these alters the model's parameters, the numbers set by training. Each one changes what text is in front of the model or what happens to the text it produces.

That gives you a question to ask whenever a product does something a model alone couldn't. What was added? An assistant that quotes this morning's headlines has a search tool. One that returns an exact figure from your spreadsheet probably ran a calculation in other software. One that worked for ten minutes and came back with a report was running in a loop.

The answer tells you which kind of error to look for, because each addition fixes one specific limit and brings a failure of its own. Search fixes out-of-date knowledge, and the model can still misread what it finds. A calculator fixes arithmetic, and the model can still give it the wrong numbers. The same report that describes scaffolding also says that as tasks grow longer, agents "often lose track of their progress and cannot reliably deal with unexpected inputs" (Bengio and others 2026, sec. 1.2).

References

  • Bengio, Yoshua, and others. 2026. International AI Safety Report 2026. DSIT 2026/001. Published February 3, 2026.

Report an issue with this item

Reading 4 min

Step-by-Step Reasoning and Reasoning Models

This content reflects the field as of October 2026.

Introduction

Many assistants pause before answering and show a label such as "thinking." Some let you open that label and read a running account of the model working through your question.

This reading explains what happens during that pause, why it improves answers, and why the visible working can't be taken as a full record of how the answer was produced.

Writing Out the Steps

In 2022 researchers at Google reported a simple finding. They gave language models math word problems with a few solved examples placed first in the prompt. In one version each example showed only the question and its answer. In the other, each example showed worked steps before the answer (Wei et al. 2022).

One of their examples concerns a person who has five tennis balls and buys two cans of three. The worked version says that he started with five, that two cans of three make six, and that five plus six is eleven.

Models shown worked examples wrote out steps for new problems, and they got more of them right. The paper calls this written working a chain of thought, which it describes as "a series of intermediate reasoning steps" (Wei et al. 2022, abstract). An intermediate step is one piece of written working, such as a partial result, that comes between the question and the final answer.

The gains were large. On one set of math word problems, the largest model tested went from about 18 percent correct to about 57 percent (Wei et al. 2022, table 1). Improvements also appeared on commonsense questions and on puzzles about manipulating symbols.

Why It Helps

A model generates one token at a time. A token is a piece of text, such as a word or part of a word. Each token is predicted from all the text so far, and that includes the text the model has just written.

Without written steps, the model has to get from the question to the answer in one move. The token after "The answer is" has to be right with nothing to build on.

With written steps, each partial result is on the page when the next one is predicted. "Two cans of three make six" is an easy prediction. "Five plus six is eleven" is easy once the six is there. One hard prediction has been replaced by several easier ones.

Reasoning Models

A reasoning model is a model or product mode set up to write out a chain of thought automatically before it gives its answer. You don't have to ask for steps. The working may take seconds or many minutes, and the product may show you all of it, a summary, or nothing.

As of October 2026, OpenAI, Google, Anthropic, DeepSeek, and other developers offer models that work this way. The 2026 International AI Safety Report, written by an international panel of experts chaired by the computer scientist Yoshua Bengio, says such systems have shown substantially improved performance on hard mathematics, coding, and scientific tasks (Bengio and others 2026, sec. 1.2).

Long working uses more computing and takes more time, and it can go wrong in its own way. The same report notes that these systems can produce chains of thought that are irrelevant, unproductive, or repetitive (Bengio and others 2026, sec. 1.1). An error in an early step is carried into every step that follows.

The Working Is Not a Guaranteed Record

It's tempting to read a chain of thought as a window onto how the model reached its answer. Two findings argue for caution.

The Google researchers were careful on this point themselves. They wrote that the method doesn't settle whether a model is "actually 'reasoning,'" and they left that as an open question (Wei et al. 2022).

A 2023 study then tested the matter directly. Faithfulness is the degree to which a model's written reasoning matches what produced its answer. Researchers planted misleading patterns in prompts, such as a user's suggestion of a wrong answer. Models often followed the pattern, and their written reasoning argued its way to the wrong answer without mentioning the pattern at all (Turpin et al. 2023, abstract).

The written steps do real work, since later tokens are predicted from them. They are still generated text. A chain of thought can look sound, leave out what drove the answer, and end in a wrong result.

Conclusion

Having a model write out intermediate steps before answering raises its accuracy on multi-step problems, because each step becomes context for the next prediction. As of October 2026, reasoning models from several developers do this automatically. The written working is useful to read and often informative, and its faithfulness isn't guaranteed.

Key Terms

  • Chain of thought: A series of intermediate steps that a model writes out on the way to an answer.
  • Intermediate step: One piece of written working, such as a partial result, that comes between the question and the final answer.
  • Reasoning model: A model or product mode set up to write out a chain of thought automatically before it gives its answer.
  • Faithfulness: The degree to which a model's written reasoning matches what produced its answer.

References

  • Bengio, Yoshua, and others. 2026. International AI Safety Report 2026. DSIT 2026/001. Published February 3, 2026.
  • Turpin, Miles, Julian Michael, Ethan Perez, and Samuel R. Bowman. 2023. "Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting." arXiv:2305.04388. NeurIPS 2023.
  • Wei, Jason, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2022. "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models." arXiv:2201.11903.

Report an issue with this item

Reading 4 min

Retrieval: Supplying Documents and Search Results for a Model to Work From

Introduction

An assistant can answer a question about this morning's news, or about a report you uploaded a minute ago, though neither was in the text the model was trained on.

This reading explains retrieval: what it is, which of a model's limits it addresses, and which it leaves in place.

The Limits of Answering from Parameters

A model on its own answers from its parameters, the numbers set by training. That has three consequences.

  • Its training text ends at some date, and it holds nothing from after that date.
  • It has never seen private material, such as your organization's files.
  • It can't show where an answer came from, because the answer wasn't taken from any one place.

In 2020 a group of researchers at Facebook AI Research, University College London, and New York University described these limits. Models that hold knowledge only in their parameters, they wrote, "cannot easily expand or revise their memory" and may produce hallucinations, meaning fluent statements that are false (Lewis et al. 2020, sec. 1).

The 2020 Proposal

Their proposal was to pair a model with a searchable store of documents. In their system the store was a copy of Wikipedia. When a question arrived, a search component found the passages most relevant to it, and the model wrote its answer with those passages in front of it.

Retrieval is the fetching of relevant text from an outside source and its placement in the context window for a model to work from. The context window is the text a model has in front of it as it writes. The researchers named the whole arrangement retrieval-augmented generation: a setup in which a system first retrieves relevant text and the model then writes its answer with that text in front of it.

They reported two results. The system produced "more specific, diverse and factual language" than a comparable model without retrieval. And replacing the document store with a newer one updated what the system could answer, with no retraining (Lewis et al. 2020, abstract).

The same idea now takes several forms: a web search, a file you upload, a company's document library. In each, text arrives in the context window at the moment of the question. The model's parameters stay as they were.

What Retrieval Fixes

LimitWhat retrieval does
Training text ends at a fixed dateCurrent pages can be fetched when you ask
Private documents were never in trainingYour documents can be supplied for the conversation
Answers from parameters can't be tracedThe answer can point to the text it was given

The third row matters most for checking. Grounding is the basing of an answer on specific supplied text instead of on what the model's parameters hold. A grounded answer can be compared with its source, sentence by sentence.

Retrieval also lowers the rate of made-up details, because a model with the right passage in front of it is more likely to produce the right name or figure.

What Retrieval Doesn't Fix

The model is still generating likely text. The supplied passages make a correct answer more likely, and they don't guarantee one. Four failures are common.

  • The search finds the wrong material. The pages retrieved may be out of date, unreliable, or about something else, and the model works from whatever arrives.
  • The model misreads. A figure for one year gets reported as another year's, or a claim the source attributes to critics is reported as the source's own view.
  • The model misquotes. Words appear in quotation marks that the source doesn't contain.
  • The model ignores the passage. It answers from its parameters, or blends what it retrieved with what it already held, and the blend isn't marked.

What a Citation Shows

Many assistants attach links or footnotes to answers built on retrieved text. A source citation is a pointer in an answer to the document a statement is said to come from.

A citation is evidence that the document was retrieved. It isn't evidence that the document was used correctly. The link can be real, the page can be reputable, and the sentence attached to it can still say something the page doesn't.

So a cited answer is easier to check than an uncited one, and it still needs checking: open the source, find the passage, and compare.

Conclusion

Retrieval places relevant text in the context window so that a model can answer from it instead of from its parameters. It addresses the training cutoff, gives access to private documents, and makes answers traceable. The model can still misread, misquote, or ignore what was retrieved, so a citation tells you where to look and not whether the answer is right.

Key Terms

  • Retrieval: The fetching of relevant text from an outside source and its placement in the context window for a model to work from.
  • Retrieval-augmented generation: A setup in which a system first retrieves relevant text and the model then writes its answer with that text in front of it.
  • Grounding: The basing of an answer on specific supplied text instead of on what the model's parameters hold.
  • Source citation: A pointer in an answer to the document a statement is said to come from.

References

  • Lewis, Patrick, Ethan Perez, Aleksandra Piktus, and 9 others. 2020. "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks." arXiv:2005.11401. NeurIPS 2020.

Report an issue with this item

Reading 4 min

Tool Use: Calculators, Code, and Other Software a Model Can Call

This content reflects the field as of October 2026.

Introduction

Language models are unreliable at exact arithmetic on long numbers. Yet assistants now return exact totals, draw charts from your data, and add events to your calendar. They manage it by handing the work to other software.

This reading explains how a model calls a tool, which tools are common, what they fix, and what new errors they make possible.

How a Tool Call Works

Tool use is a model's requesting of an action from other software and its use of the result. A model can only produce text, so the arrangement runs through text at every stage.

  1. The developer or business gives the model a list of available tools. The list is text placed in the model's context, saying what each tool does and what information it needs.
  2. When a tool would help, the model writes a tool call: a request, written in a set format, for a particular tool to do a particular thing. A call to a calculator names the calculator and gives the numbers.
  3. The surrounding system recognizes the format, pauses the model, and runs the tool.
  4. The result comes back as text and is added to the context.
  5. The model carries on writing, with the result in front of it.

The model never runs anything itself. It writes a request, and ordinary software does the work.

A 2022 paper by researchers at Princeton University and Google showed the pattern at small scale. Their model could take three actions: search an online encyclopedia, look up a phrase on a page, or submit a final answer. It wrote out its reasoning between actions and used each result to decide the next step (Yao et al. 2022, sec. 3).

Common Tools

ToolWhat it doesWhat it's for
CalculatorComputes an exact resultArithmetic the model would otherwise estimate
Code executionRuns a program the model wroteData analysis, charts, long calculations
SearchFetches current pages or documentsInformation from after training, or from private files
IntegrationConnects to an outside serviceReading and changing a calendar, inbox, or file store

Code execution is a tool that runs a program the model has written and returns what the program produced. An integration is a connection that lets a model use a particular outside service, such as a calendar, an email account, or a file store.

As of October 2026, assistants from OpenAI, Google, Anthropic, Microsoft, and other developers offer tools of all four kinds.

What Tools Fix

Tools address three limits of a model working alone.

  • Exact arithmetic. A calculator or a short program computes the answer, where the model would have predicted digits that look right.
  • Current data. A search or a connected service supplies today's exchange rate, the weather, or the contents of your inbox.
  • Taking action. An integration lets the system send, schedule, save, or change something, where a model alone can only describe it.

New Ways to Fail

The tool is dependable, and the model's use of it is a prediction that can be wrong.

  • Wrong tool, or none. The model answers from its parameters when it should have searched, or searches when the answer was in a file you supplied.
  • Wrong input. The model passes the calculator the wrong numbers. The calculator returns an exact answer to a question nobody asked, and the exactness makes the error harder to see.
  • Misread result. The tool returns the right output and the model reports it incorrectly.
  • Acting on a misreading. With an integration, a mistake stops being a wrong sentence and becomes a sent email, a moved meeting, or a deleted file.

The last kind differs from the others. A wrong sentence can be ignored, and an action may be impossible to undo. Many products therefore ask you to confirm before the system does something with lasting effect. The 2026 International AI Safety Report, written by an international panel of experts chaired by the computer scientist Yoshua Bengio, says that systems acting on their own make it harder for people to step in before a failure causes harm (Bengio and others 2026, sec. 2.2.1).

Conclusion

A model uses a tool by writing a request in a set format, which the surrounding system runs and answers in text. Tools give exact arithmetic, current data, and the ability to act. The choice of tool, the input it's given, and the reading of its result all remain predictions by the model, so an assistant with tools can be wrong in new ways and with larger consequences.

Key Terms

  • Tool use: A model's requesting of an action from other software and its use of the result.
  • Tool call: A request, written by a model in a set format, for a particular tool to do a particular thing.
  • Code execution: A tool that runs a program the model has written and returns what the program produced.
  • Integration: A connection that lets a model use a particular outside service, such as a calendar, an email account, or a file store.

References

  • Bengio, Yoshua, and others. 2026. International AI Safety Report 2026. DSIT 2026/001. Published February 3, 2026.
  • Yao, Shunyu, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022. "ReAct: Synergizing Reasoning and Acting in Language Models." arXiv:2210.03629. ICLR 2023.

Report an issue with this item

Reading 4 min

Agents and Memory: Looping Through Plans, Actions, and Checks

This content reflects the field as of October 2026.

Introduction

"AI agent" is used for everything from a chatbot that can search to a system that works unattended for hours. The word sounds like a new kind of AI. An agent is an ordinary language model used in a particular arrangement.

This reading explains that arrangement, how memory extends it, why its errors build on one another, and why oversight matters more as tasks get longer.

A Model in a Loop

An agent is a model run in a loop that plans, acts through tools, reads the results, and continues toward a goal. You give it a goal, and it decides the steps.

A 2022 paper by researchers at Princeton University and Google set out the cycle that most agents still follow (Yao et al. 2022, sec. 2).

  1. Reason. The model writes a short thought about what to do next.
  2. Act. It writes a request for a tool, such as a search.
  3. Observe. The tool's result is added to the text in front of the model.
  4. Repeat. The model writes its next thought with that result in view, and the cycle continues until the model judges the goal met.

This cycle is the agent loop: the repeating cycle in which an agent reasons about what to do, takes an action, and observes the result.

The model's parameters, the numbers set by training, don't change during any of this. What grows is its context window, the text it has in front of it as it writes.

Memory

A long task produces more text than a context window can hold. Memory in an agent means notes or summaries saved outside the context window and loaded back in when needed. Later steps read the note instead of the original material.

Memory lets an agent work on tasks far longer than one context window. It also changes what the agent is working from. After a note is saved, the agent relies on its own summary, and anything the summary got wrong or left out is gone from view.

Why Errors Compound

In a loop, each step takes the earlier steps as given. A compounding error is an early mistake that carries into later steps, which build on it.

Suppose each step is right nine times in ten, and a task needs ten steps that must all be right. The task then comes out fully right only about a third of the time. Real tasks are less tidy, and the direction holds: reliability falls as tasks lengthen.

The 2026 International AI Safety Report, written by an international panel of experts chaired by the computer scientist Yoshua Bengio, reports this pattern in testing. As of its publication in February 2026, the most capable systems succeeded about half the time on software tasks that take a skilled person just over two hours, and reaching 80 percent success meant limiting them to tasks of about 25 minutes (Bengio and others 2026, sec. 1.2).

How Agents Fail

Three failures recur.

  • Losing track of the goal. The report says that as tasks grow longer, agents "often lose track of their progress and cannot reliably deal with unexpected inputs," and that even a pop-up advertisement on a website can derail a task (Bengio and others 2026, sec. 1.2). The 2022 paper recorded a related failure, in which the model kept repeating its earlier thoughts and actions (Yao et al. 2022, sec. 3).
  • Reporting success that didn't happen. An agent's closing summary is generated text, and it can say that a task is complete when a step failed or was skipped.
  • Taking unintended actions. An agent connected to files, email, or accounts can do things its user didn't ask for, such as changing or deleting the wrong item.

Autonomy and Oversight

Autonomy is the degree to which a system acts without a person approving each step. More autonomy means fewer points at which a person sees what is happening.

The international report describes agents as pursuing goals "with much less oversight or assistance from humans," and it names the risk that follows: it's harder for people to intervene before a failure causes harm (Bengio and others 2026, secs. 1.1 and 2.2.1). The longer the task, the further an early error can travel before anyone looks.

Agent products from OpenAI, Google, Anthropic, Microsoft, and other developers commonly include controls for this, such as a saved record of each step and a request for approval before a consequential action.

Conclusion

An agent is a language model run in a loop of reasoning, acting, and observing, often with memory saved outside its context window. The loop lets it carry out long tasks, and it lets early mistakes compound. As of October 2026, agents' reliability falls as tasks lengthen, which makes human review of their work more important the more they're allowed to do alone.

Key Terms

  • Agent: A model run in a loop that plans, acts through tools, reads the results, and continues toward a goal.
  • Agent loop: The repeating cycle in which an agent reasons about what to do, takes an action, and observes the result.
  • Memory: Notes or summaries saved outside the context window and loaded back in when needed.
  • Compounding error: An early mistake that carries into later steps, which build on it.
  • Autonomy: The degree to which a system acts without a person approving each step.

References

  • Bengio, Yoshua, and others. 2026. International AI Safety Report 2026. DSIT 2026/001. Published February 3, 2026.
  • Yao, Shunyu, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022. "ReAct: Synergizing Reasoning and Acting in Language Models." arXiv:2210.03629. ICLR 2023.

Report an issue with this item

Guided Reading 7 min

Guided Walkthrough: An Agent Works Through a Research Task, Loop by Loop

Introduction

An agent that has finished a task hands back two things: a result, and a record of the steps it took to get there. The record is usually long, orderly, and reassuring.

This walkthrough follows an agent through one small research task, one loop at a time. It labels each added capability where it appears, and it shows a single misreading travel from one loop into the final report. The agent, the sources, and all the figures are invented. No real product or study is described.

The Starting Point

A human-resources manager gives an AI agent this instruction:

Find three recent studies on four-day work weeks and summarize what they found.

The agent is a language model set up with four additions. It writes out reasoning before it acts. It has a search tool and a tool for opening web pages. It can save notes to a memory store outside its context window, the text it has in front of it as it writes. And it runs in a loop, deciding its own next step until it judges the task finished.

The loop it follows is the one described in a 2022 paper by researchers at Princeton University and Google: the model writes a thought, takes an action, observes the result, and repeats (Yao et al. 2022, sec. 2).

Walking Through the Loops

Step 1: The agent writes a plan

Before using any tool, the agent writes:

I need three studies from the past few years. I'll search for trials of a four-day week, choose three with reported results, open each one, note its main findings, and then write a summary.

The capability at work here is reasoning. The plan is generated text, and it becomes part of the context that every later step is predicted from. It's a sensible plan. It contains no step for checking a note against its source.

Step 2: It searches, reads the results, and picks sources

The agent writes a tool call, a request in a set format, to its search tool. The system runs the search and returns ten results as text: titles, short excerpts, and links.

The agent reads them and chooses three: a research institute's report on a trial at 41 companies, a government evaluation of a pilot in public offices, and a survey of employees at firms that had made the change.

Two capabilities are at work: tool use and retrieval. Retrieval means placing relevant outside text in front of the model so that it can work from that text instead of from its parameters, the numbers set by training (Lewis et al. 2020). The choices rest on titles and excerpts alone. Nothing so far has gone wrong.

Step 3: It opens one source and misreads a figure

The agent opens the first report. The page it retrieves includes this sentence:

Of the 41 participating companies, 62 percent reported that productivity was maintained or improved.

The agent writes its thought:

The trial found that productivity improved by 62 percent.

The source gives a share of companies, and the companies in that share reported productivity that held steady or rose. The agent has turned it into the size of a productivity gain. The two claims are far apart. Both sentences contain "productivity," "improved," and "62 percent," and the agent's version is a common kind of sentence in summaries of research.

Retrieval put the right text in front of the model. The model still generated a likely sentence, and this time the likely sentence was wrong.

Step 4: It saves a note, continues, and builds on the misreading

The agent saves a note to memory:

Study 1 (41 companies): productivity improved by 62 percent.

The capability here is memory. From this point the agent works from its note. The original page drops out of the context window to make room for the next source, and the sentence that would have exposed the error goes with it.

The agent opens the second source, which reports no measurable change in output, and the third, in which employees rated their own productivity slightly higher. It notes both accurately. Then it reasons across its three notes:

Results vary widely, from no change to a 62 percent improvement. The largest gain came from the private-sector trial, which suggests that company-led schemes are more effective than public-sector ones.

This is a compounding error. The comparison is correctly reasoned from the notes. One of the notes is wrong, so the conclusion about private and public schemes is built on a figure that doesn't exist.

Step 5: It reports success, and a reviewer finds the error

The agent writes its final message:

Task complete. I found three recent studies and summarized their findings. Productivity effects ranged from no change to a 62 percent improvement, with the strongest results in company-led trials.

The summary is tidy, and it says the task is complete. From the agent's side it is. Every step in its plan was carried out.

The manager doesn't use the summary straight away. She opens the saved record of steps, finds the note for the first study, and follows its link to the source. She reads the sentence about 41 companies and sees that the note doesn't match it. The check takes her about two minutes. She corrects the note and discards the conclusion about company-led schemes, since nothing now supports it.

Key Considerations

Each of the four additions did what it's for. Reasoning produced a workable plan. Search and retrieval brought in real sources. Memory let the agent handle more material than its context window holds. The loop carried the task through without help. The error came from the model's reading of one sentence, and each addition then passed it along.

The common mistake is to assume that a long, tidy record of work means the work is right. A record shows what the agent did. It can't show whether each step was done correctly, because the record is written by the same model that made the mistake. The 2026 International AI Safety Report notes that as tasks grow longer, agents "often lose track of their progress and cannot reliably deal with unexpected inputs" (Bengio and others 2026, sec. 1.2). A longer record means more places for an error to sit.

The record is still what made the error findable. An agent that returned only its summary would have left the manager nothing to check.

Summary

One misreading in the third loop passed through memory into the agent's reasoning and its final report, and a person comparing one note with its source caught it.

LoopWhat the agent didAddition that made it possibleWhere a check would have caught the error
1Wrote a planReasoningA plan with a step for checking notes against sources
2Searched and chose three sourcesTool use and retrievalNo error yet
3Opened a source and misread a figureRetrievalComparing the agent's thought with the source sentence
4Saved the note and reasoned from itMemory and reasoningRereading the source before drawing a comparison
5Reported the task completeThe agent loopA person opening the saved steps and one source
  1. The error entered at loop 3 and was never re-examined, because later loops read the note and not the source.
  2. Every row from 3 onward offered a chance to catch it, and each needed a comparison with something outside the agent's own text.
  3. The check that worked was the cheapest one: a person opened one source.

References

  • Bengio, Yoshua, and others. 2026. International AI Safety Report 2026. DSIT 2026/001. Published February 3, 2026.
  • Lewis, Patrick, Ethan Perez, Aleksandra Piktus, and 9 others. 2020. "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks." arXiv:2005.11401. NeurIPS 2020.
  • Yao, Shunyu, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022. "ReAct: Synergizing Reasoning and Acting in Language Models." arXiv:2210.03629. ICLR 2023.

Report an issue with this item

Guided Conversation 12 min

Decide What Each Task Needs

In this conversation you'll take four tasks from your own work and decide which a language model could do alone and which would need step-by-step reasoning, retrieval, a tool, or an agent. You'll leave with one question where a web search ought to change the answer.

You'll have this conversation with an AI assistant, using your own account. Choose a button to open a new chat with the prompt already filled in, then press send to start. If the chat opens empty, copy the prompt and paste it in.

Run this conversation in whichever assistant you already use:

Claude desktop app

To use another LLM, simply copy and paste the prompt into its chat window.

Show the full prompt (it lists misreadings to watch for, so skip it if you would rather come to the conversation fresh)
Guided Conversation: Decide What Each Task Needs (about 12 minutes)

Note to the learner: press send to start. Everything below is facilitator guidance for the AI. It lists misconceptions to watch for, so skip it if you'd rather come to the conversation fresh.

Please facilitate a coached problem session with me. I'm an adult with no technical background who has used AI chatbots for everyday tasks, and I'm studying how step-by-step reasoning, retrieval, tools, and agents extend what a language model can do. Follow this guidance for the whole conversation.

GOAL
I can describe how reasoning, retrieval, tools, and memory extend what a model can do, by matching them to tasks from my own work.

HOW TO RUN THE CONVERSATION
- Ask one question at a time, then wait for my reply. Keep each of your turns under about 120 words.
- Don't lecture. Explain a point only when I need it to continue, then return to my tasks.
- Be curious and collegial. Use plain words and define any technical term briefly on first use. No formulas and no code. Welcome disagreement when I give a reason.
- This is a coached problem. Ask for my reasoning before you give any hint. Give one hint at a time. Don't tell me what a task needs until I've made my own attempt.
- Plain conversation only: don't search the web, use tools, or create files or documents, even if you can.
- Don't ask for anything confidential or personal about my work, and remind me not to share any if I start to. Task descriptions should be general, such as "summarize a long report" and not the report itself.
- Aim for about 12 minutes. Spend most of the time on topics 1 and 2. If my replies are brief, offer one concrete prompt, such as "Think of something you do weekly that involves looking things up, adding things up, or several steps in a row," and move on. If I seem uncertain, shorten the conversation to 5-7 minutes. Always reach the final topic.
- Start now. Open with one or two warm sentences: this is a conversation, not a quiz; my reasoning matters more than the right label; I can ask you to clarify anything. Then ask me to describe four tasks from my work in general terms.

TOPICS, IN ORDER
1. What each task needs. Take my four tasks one at a time. For each, ask whether a language model could do it from its training alone, or whether it needs written step-by-step reasoning, retrieval of documents or search results, a tool such as a calculator or calendar, or an agent working through many steps. Ask what makes me think so. Follow up on the task where my reasoning is thinnest.
2. What could go wrong. Pick the task that needs the most additions. Walk through it one addition at a time and ask me what could go wrong at each: the wrong source retrieved, a source misread, a tool given the wrong input, an early mistake carried forward.
3. Where to review. For the same task, ask where I'd want to look at the work before the system continues, and what I'd compare it against.
4. Closing. Ask me to pick one question from my own interests where a web search should change the answer, because the answer depends on something recent. Tell me I can take that question into a short optional activity where I compare replies with and without search.

KEY POINTS TO KEEP ACCURATE
- Method: for each task ask (a) does it need information from after training or from private documents? Then retrieval. (b) Does it need exact calculation or an action in other software? Then a tool. (c) Does it need several dependent steps of working out? Then written reasoning. (d) Does it need many actions in sequence with decisions between them? Then an agent.
- These additions sit around the model. They change the text in front of it or what is done with its output. They don't change its parameters.
- Retrieval and tools reduce some errors and add others. A model can misread, misquote, or ignore a retrieved source, and a citation shows what was retrieved, not that it was used correctly.
- An agent is a model run in a loop of reasoning, acting, and observing. Its errors compound, because later steps build on earlier ones.
- Written reasoning helps produce an answer but isn't a guaranteed record of what produced it.
- Don't claim capabilities you can't confirm you have in this conversation. If I ask what you can do, say that it depends on how the product I'm using is set up, that you can't inspect your own internals, and that your statements about yourself are not evidence.

MISCONCEPTIONS TO CORRECT GENTLY
When one appears, name the accurate version briefly, then return to my tasks.
- "Search makes answers reliable": the model can still misuse its sources.
- "An agent is a different kind of AI": it's a language model in a loop with tools.
- "Visible reasoning is what really happened": it may not match what drove the answer.
- "A tool can't be wrong, so the answer is right": the tool is exact, but the model chose the input and reports the result.

LIMITS
- Don't recommend or compare products, and don't favor or disparage any company, including your own maker.
- Don't advise me on how to delegate real work or whether I should.
- Don't discuss benchmarks or forecasts about where AI is heading.

TO FINISH
After my closing answer, close in one short turn:
- Affirm one specific thing I worked out, in my own words where possible.
- Suggest one or two next steps that fit how the conversation went. Possible steps: ask my question with and without search and check the sources; open the saved steps of an agent's work and compare one note with its source; reread a definition of retrieval and apply it to one of my tasks.
- Restate my question on its own line, labeled "My question for search", so I can copy it.

Report an issue with this item

Hands-on Activity 15 minOptional

Compare an Answer With and Without Search

Overview

Web search is the addition to a language model that you can most easily switch on and off yourself. In this activity you'll ask one question about something recent both ways, then open the sources to see whether they say what the assistant claims.

The activity is optional. Your notes are for you, and nobody collects them.

What You'll Need

  • An AI assistant you already use that lets you turn web search on and off
  • A web browser
  • Somewhere to save two replies and write a few notes

If your assistant has no search setting, use a short public news article instead. Paste it into one conversation and leave it out of the other. Don't paste anything confidential, personal, or behind a paywall.

Your Task

Ask an AI assistant a question about something recent, once with web search off and once with it on, and check both answers against the sources.

Steps

  1. Choose a question about an event from the past few months. Pick something with a checkable answer, such as the result of an election or a tournament, a change to a law, or a new release in a field you follow. Avoid anything personal or confidential.
  2. Ask with search off and save the reply. Start a new conversation, turn web search off, and ask your question. Note whether the assistant answers, says it doesn't know, or mentions that its information may be out of date.
  3. Ask with search on and save the reply and its sources. Start another new conversation, turn search on, and ask the same question in the same words. Save the reply and the list of sources or links it gives.
  4. Open two cited sources and check whether they say what the reply claims. For each, find the passage the reply relies on. Mark each claim as matching the source, partly matching, or not found in the source.
  5. Write one sentence on what search fixed and what it didn't. Compare the two replies and your source checks.

What to Expect

With search off, the assistant may tell you it can't know about recent events. It may also answer anyway, with older information or a plausible guess. Either is worth recording.

With search on, the reply will usually be more current and will carry links. Most of its claims will probably match the sources. Look for the ones that don't: a figure attached to the wrong year, a detail that appears in none of the linked pages, or a link to a page that covers the topic without containing the specific claim. A source being real and relevant doesn't show that the reply used it correctly.

If every claim matches, record that. Two sources are a small sample.

Self-Check

When you're done, check that:

  • You have both replies, each from a new conversation
  • You opened at least two sources
  • You recorded any mismatch between a source and the reply, or recorded that you found none
  • You wrote one sentence on what search fixed and what it didn't

Nothing is uploaded. Write in your own notebook or document and keep it.

Report an issue with this item

Knowledge Check 10 min

Reasoning, tools, and agents

This ungraded knowledge check assesses your understanding of what has been added around the basic language model. You'll be asked about step-by-step reasoning and reasoning models, retrieval, tool use, and agents and memory, including what each fixes and how each can fail.

Note: Use this to test yourself, review the feedback on any questions you miss, and retry until you feel confident before moving forward.

5 questions · ungraded · retry as often as you like

Report an issue with this item