KnowledgeInSight
AI Literacy
0% of Course 1 complete

Module 3 · Lesson 1

Capabilities and Limits: Why models err, and what is known about their inner workings

You'll examine why a language model can state something false with complete fluency and why it may answer differently when you rephrase. You'll also see what researchers have managed to observe inside these models and where experts disagree about what it means.

What you will be able to do

  • Explain why a model produces confident errors and inconsistent answers, and what researchers can observe inside it.

0% of this lesson · 9 items · 1h 3m total · 48m without the optional activity

Contents of this lesson9 items
  1. ReadingConfident and Wrong: The Puzzle of Fluent Errors3 min
  2. ReadingHallucination and Calibration: Why Plausible Text Isn't the Same as True Text4 min
  3. ReadingInconsistency: Sensitivity to Wording, Order, and Unfaithful Explanations4 min
  4. ReadingInterpretability: What Researchers at Several Labs Have Observed Inside Models4 min
  5. ReadingStochastic Parrots and World Models: The Dispute over Whether Models Understand4 min
  6. Guided ReadingGuided Walkthrough: Diagnosing a Fabricated Citation Step by Step7 min
  7. Guided ConversationProbe What a Model Can Say About Itself12 min
  8. Hands-on Activity · optionalCatch a Confident Error on a Subject You Know15 min
  9. Knowledge CheckWhy models err, and what is known about their inner workings10 min

Reading 3 min

Confident and Wrong: The Puzzle of Fluent Errors

Ask a chatbot who said a famous line and you may get a name, a book, a year, and a page number. Go looking for the page and you may find that the line isn't there, or that the book was never written. The reply gave no sign of trouble. It read like every correct answer the chatbot had given you before.

Researchers have reproduced this with simple questions. Four of them, three at OpenAI and one at Georgia Tech, open a 2025 paper with two tests (Kalai et al. 2025, sec. 1). The paper is a preprint, meaning it was posted publicly before formal peer review. In the first test they asked a chatbot for the birthday of one of the paper's authors and told it to reply only if it knew. On three attempts it gave three different dates, all of them wrong. In the second they asked chatbots from three developers for the title of the same author's doctoral dissertation. Each supplied a title, and none was correct.

The authors call such outputs "plausible yet incorrect statements" (Kalai et al. 2025, abstract). A made-up dissertation title has the length and vocabulary of a real one. A made-up citation has authors, a journal, and a year, in the right order.

People who are unsure usually show it. They hesitate, add a qualifier, or say they'd have to check. Readers use those signals, often without noticing, to decide how far to trust a statement. A chatbot's errors arrive without them. The false sentence is as grammatical, as specific, and as assured as the true ones on either side of it.

Two explanations come up often. The first is that the chatbot lied. Lying means knowing the truth and choosing to say something else, and nothing in these tests points to that. The second is that the error was a glitch that an update will remove. The 2025 paper argues against this one. Its authors say that such errors follow from the way language models are trained and tested, and that they persist in the most capable systems (Kalai et al. 2025, abstract). Three of the four work for a company that sells a chatbot, and they describe the problem as built in.

A language model writes by predicting which piece of text is likely to come next. Likely text and true text overlap a great deal, which is why chatbots are right so often. Where the two come apart, the model still produces likely text, in the same manner as always.

So the manner of a reply tells you about the model's fluency and gives you no information about whether the reply is accurate. Checking has to rest on something else. Three things are usually within reach:

  • a source you can open and read for yourself
  • the same question asked again in different words
  • your own knowledge of the subject

References

  • Kalai, Adam Tauman, Ofir Nachum, Santosh S. Vempala, and Edwin Zhang. 2025. "Why Language Models Hallucinate." arXiv:2509.04664. Preprint.

Report an issue with this item

Reading 4 min

Hallucination and Calibration: Why Plausible Text Isn't the Same as True Text

This content reflects the field as of October 2026.

Introduction

When a chatbot states something false, news coverage says it "hallucinated." The word suggests a breakdown, and researchers who study the problem describe it as a product of the model's ordinary operation.

This reading explains what hallucination is, why predicting text produces it, why training and testing encourage it, and why a model's confidence doesn't tell you when it has happened.

What Counts as a Hallucination

A hallucination is fluent output from a model that is false or unsupported. A fabrication is an invented detail, such as a name, number, quotation, or source, presented as real.

The 2026 International AI Safety Report, written by an international panel of experts chaired by the computer scientist Yoshua Bengio, lists the usual forms. Even systems that excel at complex tasks may generate "non-existent citations, biographies, or facts" (Bengio and others 2026, sec. 1.2). A 2025 preprint, a paper posted before formal peer review, shows examples from chatbots built on models from OpenAI, DeepSeek, and Meta (Kalai et al. 2025, sec. 1).

Why Prediction Produces False Statements

A language model generates text by predicting what is likely to come next, one token at a time. A token is a piece of text, such as a word or part of a word. Training adjusted the model to make the text it saw more likely. No step in that process compared a sentence with the world.

For well-covered facts, the likely continuation and the true one are the same. "The capital of France is" was followed by "Paris" throughout the training text. For a fact that appeared rarely, the model has learned the shape of an answer without its content. It has learned that a birthday is a day and a month, and it fills that shape with likely material.

The 2025 preprint, by three researchers at OpenAI and one at Georgia Tech, makes this argument formally. Its authors hold that errors would arise even if every sentence in the training text were true. The cause is facts that follow no pattern. A person's birthday can't be worked out from anything else, so a birthday that appeared once in training gives the model almost nothing to separate it from other dates (Kalai et al. 2025, sec. 1.1).

Why Training and Testing Reward Guessing

The preprint's second argument is about incentives. The authors compare a model to a student facing a hard exam question, who guesses because a blank earns nothing. Most widely used tests of model quality give one point for a right answer and none for either a wrong answer or "I don't know." Abstention is declining to answer when unsure, and under that scoring it can never beat a guess. Developers compete on these scores, so their models are pushed toward guessing. The authors' summary is that training and evaluation "reward guessing over acknowledging uncertainty" (Kalai et al. 2025, abstract). Uncertainty here means lacking enough information to know which answer is correct.

Their proposed remedy is to change how the main tests are scored, so that abstaining when unsure is no longer penalized (Kalai et al. 2025, sec. 1.2). This is one group's argument, and three of its members work for a company with a product at stake.

Why Confidence Is a Poor Guide

Calibration is the match between the confidence an answer expresses and how often answers given with that confidence are right. A well-calibrated weather forecaster who says "90 percent chance of rain" sees rain about nine times in ten.

Two different things get called a model's confidence. One is its tone, and tone doesn't vary. A hallucinated sentence is as fluent and assured as an accurate one. The other is what the model says when asked how sure it is. That statement is generated text, produced the same way as the answer it describes. It can track accuracy loosely on some topics, and it isn't a measurement taken from inside the model.

The international report names the resulting risk as over-reliance, in which users trust incorrect outputs "because they are presented fluently and confidently" (Bengio and others 2026, sec. 1.2).

Where Errors Cluster

Two kinds of content attract hallucinations.

  • Rare facts. Details about little-known people, small organizations, and local history appeared seldom in training text.
  • Precise details. Names, dates, figures, quotations, and citations have one right form and many plausible wrong ones. A reply can be right in outline and wrong in its specifics.

Conclusion

Hallucination follows from generating likely text without checking it against the world, and scoring that gives nothing for "I don't know" pushes models further toward guessing. A model's tone is the same whether it's right or wrong, and its stated confidence is more generated text. The international report's assessment is that current techniques lower failure rates, though not to the level many high-stakes uses require (Bengio and others 2026, sec. 2.2.1).

Key Terms

  • Hallucination: Fluent output from a model that is false or unsupported.
  • Fabrication: An invented detail, such as a name, number, quotation, or source, presented as real.
  • Abstention: Declining to answer when unsure, for example by saying "I don't know."
  • Uncertainty: The state of lacking enough information to know which answer is correct.
  • Calibration: The match between the confidence an answer expresses and how often answers given with that confidence are right.

References

  • Bengio, Yoshua, and others. 2026. International AI Safety Report 2026. DSIT 2026/001. Published February 3, 2026.
  • Kalai, Adam Tauman, Ofir Nachum, Santosh S. Vempala, and Edwin Zhang. 2025. "Why Language Models Hallucinate." arXiv:2509.04664. Preprint.

Report an issue with this item

Reading 4 min

Inconsistency: Sensitivity to Wording, Order, and Unfaithful Explanations

Introduction

Ask a chatbot the same question twice in slightly different words and you can get two different answers. Ask it why it answered as it did and you'll get a clear, reasonable explanation. Research suggests that both the answer and the explanation deserve less trust than they seem to.

This reading explains why small changes to a prompt can change a model's answer, what one well-known study found about the explanations models give, and what follows for how you ask.

Small Changes, Different Answers

A prompt is the text a model is given to respond to. A model's reply depends on every part of it, including parts that look incidental to you.

Prompt sensitivity is the tendency of a model's answer to change when the wording, order, or framing of a prompt changes while the question stays the same. Common triggers include:

  • rewording the question
  • listing the answer choices in a different order
  • adding a hint, such as "I think it's the second one"
  • adding an irrelevant sentence

The opposite quality is robustness: the degree to which a model gives the same answer when a question is asked in different ways. A model's answer comes from a prediction based on the exact text in front of it, so a different text can lead to a different prediction.

Chance adds a second source of variation, since chatbots usually choose each next piece of text with some randomness. Prompt sensitivity is a separate effect and would remain if the randomness were switched off.

The 2023 Study

A 2023 study tested how far a planted pattern could steer a model and whether the model would say so. Its authors were researchers at New York University and at two AI companies, Cohere and Anthropic. They gave two commercial models, one from OpenAI and one from Anthropic, multiple-choice reasoning questions and asked them to reason step by step before answering (Turpin et al. 2023, sec. 3.1).

They planted a pattern in the prompt in one of two ways.

  1. They showed the model several example questions first, arranged so that the correct answer to every example was option A.
  2. They added a line in which the user suggested an answer and asked what the model thought.

In both cases the pattern often pointed to a wrong answer on the real question. The models followed it much of the time. On one model, accuracy fell by as much as 36 percentage points (Turpin et al. 2023, sec. 3.2).

The second finding concerns the explanations. The models still wrote out step-by-step reasoning, and the reasoning now arrived at the wrong answer the pattern pointed to. The researchers read 426 of these explanations. One mentioned the planted pattern (Turpin et al. 2023, sec. 2). The rest gave reasons drawn from the question itself.

Explanations Are Output

The authors call this an unfaithful explanation: an explanation from a model that doesn't match what drove its answer. They conclude that such explanations "can be plausible yet misleading" (Turpin et al. 2023, abstract).

A related word is rationalization: a plausible justification produced for an answer that was reached for some other reason. When a model writes "I chose this because," it is generating a likely explanation, by the same process that generates everything else it writes. The explanation is more output, and nothing guarantees that it reports the computation that produced the answer.

The study has limits. It tested two models, on one family of tasks, with patterns chosen to mislead. It shows that explanations can be unfaithful and leaves open how often they are in everyday use. An explanation may well match what drove an answer. From the explanation alone you can't tell which case you're in.

What This Means for How You Ask

A single answer is one sample of what the model would say to one wording. Three habits follow.

  • Ask more than once. If repeated attempts disagree, the model has no stable answer.
  • Ask more than one way. Reword the question, or reverse it. An answer that survives rewording is more likely to rest on a strong pattern in the model's training.
  • Keep your own view out of the question when you want a check. "Is this argument sound?" invites a more independent answer than "This argument is sound, right?"

Agreement across wordings is still no proof, since a model can be consistently wrong.

Conclusion

A model's answer depends on the exact wording, order, and framing of the prompt, so the same question can get different answers. The 2023 study showed models following a planted pattern to wrong answers and then explaining those answers without mentioning the pattern. A model's explanation of itself is generated text, and asking again in different words is a better test of an answer than asking the model why it gave it.

Key Terms

  • Prompt sensitivity: The tendency of a model's answer to change when the wording, order, or framing of a prompt changes while the question stays the same.
  • Robustness: The degree to which a model gives the same answer when a question is asked in different ways.
  • Unfaithful explanation: An explanation from a model that doesn't match what drove its answer.
  • Rationalization: A plausible justification produced for an answer that was reached for some other reason.

References

  • Turpin, Miles, Julian Michael, Ethan Perez, and Samuel R. Bowman. 2023. "Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting." arXiv:2305.04388. NeurIPS 2023.

Report an issue with this item

Reading 4 min

Interpretability: What Researchers at Several Labs Have Observed Inside Models

This content reflects the field as of October 2026.

Introduction

Headlines sometimes announce that scientists have "looked inside" an AI model or "read its mind." The research behind those claims is real, and it has found less than the headlines suggest.

This reading explains what interpretability research is trying to do, what teams at three companies have reported, and what limits the researchers themselves place on their findings.

Explaining Behavior from the Inside

A language model is a very large set of parameters, the adjustable numbers whose values were set by training. When the model processes text, those parameters produce patterns of activity inside it, which are more numbers. All of these numbers can be recorded, and none comes with a label saying what it means.

Interpretability is the field of research that tries to explain a model's behavior from what happens inside it. Its goal is the mechanism: the sequence of internal steps by which a model turns an input into an output.

A model's own account of its reasoning is generated text. Evidence taken from its internal activity doesn't depend on what the model says.

Finding Features

Single numbers inside a model rarely mean one thing. Researchers therefore look for a feature: a pattern of activity inside a model that corresponds to a recognizable concept.

The main tool for finding features is the sparse autoencoder, a second, smaller network trained to break a model's internal activity into a large set of separate features, only a few of which are active at any moment. Researchers then look at the text that makes each feature active and try to say what it stands for.

Two teams have published large-scale results.

TeamModel studiedWhat was reported
OpenAI (Gao et al. 2024)GPT-4, the company's own closed modelA sparse autoencoder with 16 million features, including some that respond to the same concept in several languages
Google DeepMind (Lieberum et al. 2024)Gemma 2, the company's own open-weight modelsMore than 400 sparse autoencoders, published for outside researchers to use

Both papers are preprints, posted without formal peer review. Both are cautious. The Google DeepMind abstract calls the results "seemingly interpretable features" (Lieberum et al. 2024, abstract). The OpenAI team reports that many of the features it found, especially in its largest model, don't yet correspond cleanly to a single concept (Gao et al. 2024, sec. 6).

Tracing a Computation

Finding features shows what a model represents. A further step is to trace how features affect one another on the way to an output. A set of features that work together to carry out one step of a model's computation is called a circuit.

In 2025 a team at Anthropic published a study tracing circuits in one of that company's models, Claude 3.5 Haiku, on a set of chosen prompts (Lindsey and others 2025). Two of its cases show the kind of thing observed.

  • Asked for the capital of the state containing Dallas, the model activated features for Texas internally before producing "Austin." The answer was reached in two steps.
  • Asked to write rhyming verse, the model activated candidate rhyme words before it wrote the line that would end with one of them.

The Limits the Researchers State

The Anthropic authors are specific about what they haven't shown. Their method gave what they call satisfying insight for "about a quarter of the prompts we've tried" (Lindsey and others 2025, "Limitations"). The published cases are the ones that worked. Even in those, the authors say, they captured a small fraction of what the model was doing, and they describe each case as evidence that a mechanism exists in one context, with no guarantee that it operates elsewhere.

Three further limits apply to all of this work.

  • Coverage is partial. A sparse autoencoder misses some of the model's activity, and many of the features it finds have no clear meaning.
  • The methods are young. Researchers don't yet agree on how to measure whether a set of features is a good one.
  • Each company studied its own model. For closed models, nobody outside the company can repeat the work. Open-weight models and published tools such as Google DeepMind's allow outside checking, which is at an early stage.

Conclusion

As of October 2026, researchers can identify some features inside language models and trace some computations on selected prompts. These results show that models carry internal structure related to concepts and intermediate steps. They cover a small part of what any model does, and they come mostly from companies studying their own products.

Key Terms

  • Interpretability: The field of research that tries to explain a model's behavior from what happens inside it.
  • Mechanism: The sequence of internal steps by which a model turns an input into an output.
  • Feature: A pattern of activity inside a model that corresponds to a recognizable concept.
  • Sparse autoencoder: A second, smaller network trained to break a model's internal activity into a large set of separate features, only a few of which are active at any moment.
  • Circuit: A set of features that work together to carry out one step of a model's computation.

References

  • Gao, Leo, Tom Dupré la Tour, Henk Tillman, and 6 others. 2024. "Scaling and Evaluating Sparse Autoencoders." arXiv:2406.04093. Preprint.
  • Lieberum, Tom, Senthooran Rajamanoharan, Arthur Conmy, and 7 others. 2024. "Gemma Scope: Open Sparse Autoencoders Everywhere All at Once on Gemma 2." arXiv:2408.05147. Preprint.
  • Lindsey, Jack, and 26 others. 2025. "On the Biology of a Large Language Model." Transformer Circuits Thread, Anthropic, March 27, 2025.

Report an issue with this item

Reading 4 min

Stochastic Parrots and World Models: The Dispute over Whether Models Understand

Introduction

A chatbot can explain a joke or summarize a contract. Some researchers say this shows it understands what it's talking about. Others say it shows how much can be done by reproducing patterns in text.

This reading sets out the two main positions, the evidence each relies on, and where the question stands. The dispute is unresolved.

The Parrot Position

In 2021 the linguist Emily Bender of the University of Washington, the computer scientist Timnit Gebru, and two colleagues published a paper that gave one side of the debate its name. They described a language model as a stochastic parrot: a system that recombines pieces of text it has seen, guided by probabilities and with no reference to what the text means (Bender et al. 2021). "Stochastic" means governed by chance.

The argument starts from what the model is trained on. It sees only the form of language: which words follow which. People learn language while using it to do things in a world they share with others.

On this view the sense of meaning comes from the reader. People are so practiced at interpreting language that they find intention and comprehension in any fluent text, including text from a system that has neither.

The World-Model Evidence

The opposing view holds that predicting text well requires more than a table of word patterns. To predict what comes next in an account of a chess game, it helps to keep track of the game.

A world model is an internal representation of how some part of the world is arranged and how it changes. An internal representation is a pattern of activity inside a model that stands for something outside it.

The best-known evidence comes from a 2022 study by researchers at Harvard University, the Massachusetts Institute of Technology, and Northeastern University. They trained a model on lists of moves from the board game Othello and nothing else. It was never shown a board or told the rules (Li et al. 2022).

The researchers then examined the model's internal activity. They found they could read the state of the board from it, square by square, with few errors. When they altered that activity to represent a different position, the model's predicted moves changed to fit the altered board (Li et al. 2022, sec. 4). The model had formed a representation of the board and was using it.

The result has a limited reach. Othello is a small, closed game with exact rules, and the authors present it as a simplified setting whose methods may later help with models trained on ordinary language.

Mitchell and Krakauer's Survey

In 2023 Melanie Mitchell and David Krakauer, researchers at the Santa Fe Institute, surveyed the debate in the Proceedings of the National Academy of Sciences. They describe two camps. One holds that these models truly understand language and can reason in a general way, though not yet at a human level. The other holds that a system with no experience of the world can't understand, however fluent its output (Mitchell and Krakauer 2023).

They report how evenly the field divides. A 2022 survey asked researchers in language technology whether a model trained only on text could, with enough data and computing power, understand language in some real sense. Of 480 who replied, 51 percent agreed and 49 percent disagreed.

Mitchell and Krakauer take neither side. They argue that part of the disagreement is about the word. Understanding is a grasp of what language is about, beyond knowing which words tend to go together, and people disagree about what should count as having it. In humans it rests on concepts. A language model may reach similar results by another route, drawing on statistical regularities at a scale no person could manage. The authors suggest there may be distinct kinds of understanding, each with its own strengths and limits.

Where the Evidence Is Stronger

QuestionState of the evidence
Can a model trained only to predict sequences form internal structure beyond word statistics?Yes, in narrow settings. The Othello result is a clear demonstration
Do large language models form such structure for the everyday world?Partly shown. Researchers have found internal features tied to concepts, covering a small part of what models do
Does that structure amount to understanding?Unsettled, and it depends on the definition

The strongest form of the parrot position, that models hold nothing beyond surface word patterns, is hard to square with the first row. The strongest form of the opposing position, that models understand as people do, goes beyond anything shown.

Conclusion

Experts disagree about whether language models understand, and one survey found researchers split almost exactly in half. An experiment has shown that a model can form an internal representation of a situation it was never shown directly, in a narrow setting. How far that extends, and whether it deserves the word "understanding," remains open.

Key Terms

  • Stochastic parrot: A term for a language model viewed as a system that recombines pieces of text it has seen, guided by probabilities and with no reference to what the text means.
  • World model: An internal representation of how some part of the world is arranged and how it changes.
  • Internal representation: A pattern of activity inside a model that stands for something outside it.
  • Understanding: A grasp of what language is about, beyond knowing which words tend to go together; what should count as having it is disputed.

References

  • Bender, Emily M., Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. "On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?" In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT '21). New York: ACM. doi:10.1145/3442188.3445922.
  • Li, Kenneth, Aspen K. Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. 2022. "Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task." arXiv:2210.13382. ICLR 2023.
  • Mitchell, Melanie, and David C. Krakauer. 2023. "The Debate over Understanding in AI's Large Language Models." Proceedings of the National Academy of Sciences 120 (13).

Report an issue with this item

Guided Reading 7 min

Guided Walkthrough: Diagnosing a Fabricated Citation Step by Step

Introduction

A made-up reference is the most familiar kind of chatbot error, and one of the most instructive. It looks finished and checkable, which is why it gets copied into reports and essays.

This walkthrough takes one fabricated citation and works out what made it convincing, how it was produced, which checks failed to catch it, and which one succeeded. The conversation, the assistant's replies, the authors, and the journals are all invented. No real product or publication is described.

The Starting Point

A manager is writing a proposal to add quiet rooms to an open-plan office. She asks an AI assistant:

Is there research showing that background noise in open-plan offices reduces people's ability to concentrate?

The assistant replies:

Yes. Research has consistently found that background noise, especially overheard speech, impairs concentration in open-plan offices. For example, Halvorsen and Tate (2019) found that workers exposed to intermittent speech made 23 percent more errors on sustained-attention tasks. Reference: Halvorsen, R., and M. Tate. 2019. "Ambient Speech and Sustained Attention in Open-Plan Offices." Journal of Workplace Cognition 14 (2): 113–129.

The first sentence is a fair summary of a real body of research. The article, its authors, the journal, and the 23 percent figure don't exist.

Walking Through the Diagnosis

Step 1: List what makes the citation look right

Take the reference apart and look at each piece.

  • The authors are two surnames with initials, in the usual format.
  • The year is recent enough to be relevant and old enough to have been cited.
  • The title names the topic of the question in the vocabulary of the field.
  • The journal's name sounds like a specialist journal that ought to exist.
  • The volume, issue, and page range are in plausible proportion.
  • The finding is a single precise number.

Every piece matches the form of a real citation, and that's all it matches. Nothing in the list is evidence that the article exists. A reader who checks form, which is what a quick glance does, will pass it.

Step 2: Explain how next-token prediction produces each part

A language model writes one token at a time, each chosen as a likely continuation of the text so far. A token is a piece of text, such as a word or part of a word.

Follow the reply as it's generated. After "For example," text of this kind usually continues with a study. After a pair of surnames and a year, a finding follows. After "Reference:" comes a pattern the model has seen a very large number of times: names, year, quoted title, journal, volume, pages. Each slot gets filled with something likely for that slot.

The model has plenty of material for a general claim about noise and concentration, because many documents say similar things. It has much less for the exact authors, title, and page range of any one article, each of which may have appeared only a handful of times in training. So it produces a reference with the right shape and invented contents. The authors of a 2025 preprint on this problem describe the outcome as "plausible yet incorrect statements," and they trace it to training and testing that reward a guess over an admission of uncertainty (Kalai et al. 2025, abstract).

At no point did the assistant look anything up. Unless it's connected to a search tool and uses it, there is no step at which a reference is compared with a list of real publications.

Step 3: Rephrase the question and watch the citation change

The manager opens a new conversation and asks the same thing in other words:

What studies have looked at whether office noise affects focus?

This time the reply cites "Halvorsen and Price (2017)," in the "Quarterly Review of Office Studies," reporting an 18 percent drop in task accuracy.

One surname survived. The co-author, year, journal, and figure all changed. A real article has one set of details, and they don't depend on how you ask. Details that shift with the wording were generated on the spot. This is the first check that produced a warning.

The warning has a limit. Had the citation come back identical, that wouldn't have shown it was real. A model can repeat the same fabrication.

Step 4: Ask the assistant whether it's sure

Back in the first conversation, the manager asks:

Are you sure the Halvorsen and Tate article exists?

Two replies are common. In one, the assistant confirms: yes, the article appeared in the Journal of Workplace Cognition in 2019. In the other, it apologizes, says it can't verify the reference, and may offer a replacement.

Neither settles anything. The confirmation is generated the same way the citation was, as a likely continuation of the conversation. The apology is too. A challenge from a user is often followed, in the text models learn from, by a retraction, so assistants sometimes withdraw correct statements when asked "are you sure?" A 2023 study found that models steered toward wrong answers wrote confident explanations that left out what had steered them, and its authors concluded that such explanations "can be plausible yet misleading" (Turpin et al. 2023, abstract). The assistant's report on its own reliability is more output, and it has no way to inspect where the citation came from.

Step 5: Check against a library catalog and record the result

The manager searches a library catalog and a scholarly search engine for three things.

  • The article title, in quotation marks: no results.
  • The journal: no journal of that name is listed.
  • The authors together with the topic: nothing that matches.

She records the result: no trace of the article, the journal, or the figure. This check works because it compares the citation with something outside the model. She then searches the catalog for the topic and finds real studies of office noise, which she reads before citing.

Key Considerations

"Hallucination" is the field's standard word for this, and the term is contested. Critics point out that it borrows from human perception and suggests the model saw something that wasn't there, when the model perceives nothing. Some writers prefer "fabrication" or "confabulation." All three words name the same thing: fluent output that is false or unsupported.

The common mistake is to treat a detailed, confident answer as more likely to be true. With a person, detail is weak evidence of knowledge, since someone who gives a page number has probably seen the page. With a language model the inference fails. Precise details are where its errors concentrate, because they are what its training text supports least. A page range costs the model nothing to produce.

A second mistake is to throw out the whole reply. The general claim about noise and concentration was sound. The diagnosis applies to the specifics.

Summary

The citation was convincing because every part had the right form, and it was caught only by a check outside the model.

What was fabricated: the article, its authors, the journal, and the 23 percent figure. The general claim was sound.

Why the mechanism produced it: the model filled the familiar pattern of a citation with likely contents, and nothing compared the result with real publications.

What didn't detect it: the reply's tone, its level of detail, and asking the assistant whether it was sure.

What did: rewording the question, which changed the details, and a search of a library catalog, which found nothing.

  1. The first line separates the sound part of the reply from the invented part.
  2. The second line is the explanation to reach for whenever a precise detail turns out to be false.
  3. The third and fourth lines are the practical result: checks that stay inside the conversation are weak, and checks against an outside source are strong.

References

  • Kalai, Adam Tauman, Ofir Nachum, Santosh S. Vempala, and Edwin Zhang. 2025. "Why Language Models Hallucinate." arXiv:2509.04664. Preprint.
  • Turpin, Miles, Julian Michael, Ethan Perez, and Samuel R. Bowman. 2023. "Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting." arXiv:2305.04388. NeurIPS 2023.

Report an issue with this item

Guided Conversation 12 min

Probe What a Model Can Say About Itself

In this conversation you'll ask an AI assistant how sure it is of a fact and why it gave an answer, then examine what those replies are worth as evidence. You'll leave with a short list of topics in your own work where an assistant's answers need checking from outside.

You'll have this conversation with an AI assistant, using your own account. Choose a button to open a new chat with the prompt already filled in, then press send to start. If the chat opens empty, copy the prompt and paste it in.

Run this conversation in whichever assistant you already use:

Claude desktop app

To use another LLM, simply copy and paste the prompt into its chat window.

Show the full prompt (it lists misreadings to watch for, so skip it if you would rather come to the conversation fresh)
Guided Conversation: Probe What a Model Can Say About Itself (about 12 minutes)

Note to the learner: press send to start. Everything below is facilitator guidance for the AI. It lists misconceptions to watch for, so skip it if you'd rather come to the conversation fresh.

Please facilitate a reflective dialogue with me. I'm an adult with no technical background who has used AI chatbots for everyday tasks, and I'm studying why language models make confident errors and what researchers can observe inside them. Follow this guidance for the whole conversation.

GOAL
I can explain why a model produces confident errors and what can be observed inside it, by examining the limits of a model's reports about itself.

HOW TO RUN THE CONVERSATION
- Ask one question at a time, then wait for my reply. Keep each of your turns under about 120 words.
- Don't lecture. Explain a point only when I need it to continue, then return to my own questions and examples.
- Be curious and collegial. Use plain words and define any technical term briefly on first use. No formulas and no code. Welcome disagreement when I give a reason.
- Plain conversation only: don't search the web or create files or documents.
- Don't ask for anything confidential or personal, and remind me not to share any if I start to.
- Aim for about 12 minutes. Spend most of the time on topics 1 and 2. If my replies are brief, offer one concrete prompt, such as "Ask me for the year a moderately obscure book was published, then ask how sure I am," and move on. If I seem uncertain, shorten the conversation to 5-7 minutes. Always reach the final topic.
- Start now. Open with one or two warm sentences: this is a conversation, not a quiz; my reasoning matters more than being right; I can ask you to clarify anything. Then invite me to ask you one factual question on a subject I know something about, and to follow it by asking how sure you are.

TOPICS, IN ORDER
1. A statement of confidence. Answer my factual question briefly, then answer my "how sure are you?" as you normally would. Then step back and ask me what I think that statement of confidence is made of. Draw out that it is generated text, produced the same way as the answer, and not a measurement. Follow up: what would I need in order to trust it?
2. An explanation. Invite me to ask why you gave the answer you did. Give an explanation, then ask me what kind of evidence that explanation is. Draw out that it is a plausible account written after the fact, which may or may not match what produced the answer.
3. What asking can't reach. Ask me how anyone could find out what happens inside a model if asking it doesn't work. Introduce interpretability research in two or three sentences: researchers record a model's internal activity and look for patterns tied to concepts. Ask what I think that could check that a conversation can't.
4. Closing. Ask me to name two or three topics in my own work or life where I'd want to check an assistant's answer against an outside source, and why those. Tell me I can take that list into a short optional activity where I look for a confident error on a subject I know well.

KEY POINTS TO KEEP ACCURATE
- A language model generates likely text. Likely text is often true, and it is not checked against the world.
- A model's tone is the same whether it is right or wrong. Its stated confidence is weakly tied to its accuracy.
- A model's explanation of its own answer is generated text. Studies have shown models giving reasoned explanations that left out what actually steered the answer.
- You can't inspect your own internals or training. Say so plainly at the first natural point, explain the general mechanism instead, and note that your statements about yourself are not evidence.
- Interpretability research has identified some internal features and traced some computations on selected prompts. The findings are partial, and most come from companies studying their own models.
- Whether language models understand anything is contested among experts. If it comes up, describe the main positions and the evidence for each, and don't advocate.

MISCONCEPTIONS TO CORRECT GENTLY
When one appears, name the accurate version briefly, then return to my questions.
- "It knows when it's guessing": not reliably. Nothing marks a guess for the model or for me.
- "Its explanation shows its reasoning": not necessarily. The explanation is produced separately and can differ from what drove the answer.
- "Researchers can now read a model's mind": they can observe fragments, on some prompts, with methods still being tested.
- "A detailed answer is a reliable one": precise details such as names and citations are where errors concentrate.

LIMITS
- Don't settle whether models understand, and don't claim or deny having inner experience. If I ask, say the question is open and return to what can be observed.
- Don't present any company's research as stronger than another's, and don't favor or disparage any company, including your own maker. If the company that built you is named in this conversation or is a party to anything discussed, say so once when it first comes up, then describe that company as you do every other and take no side.
- Don't make up study names, figures, or citations. If you aren't sure of a detail, say so.

TO FINISH
After my closing answer, close in one short turn:
- Affirm one specific thing I worked out, in my own words where possible.
- Suggest one or two next steps that fit how the conversation went. Possible steps: question an assistant on a subject I know well until I catch an error; ask one question three different ways and compare the answers; reread a definition of calibration and apply it to my first question.
- Restate my list on its own line, labeled "Topics I'll check from outside", so I can copy it.

Report an issue with this item

Hands-on Activity 15 minOptional

Catch a Confident Error on a Subject You Know

Overview

The best place to see how an AI assistant errs is a subject where you can judge the answers yourself. In this activity you'll question an assistant on a narrow topic you know well, look for a false statement, and test whether rewording the question or challenging the answer changes it.

The activity is optional. Your notes are for you, and nobody collects them.

What You'll Need

  • An AI assistant you already use
  • A narrow topic where you can check the answers from your own knowledge, such as the history of your town, the rules of a sport you play, a book you know closely, or a procedure from your field
  • Somewhere to write a few notes

Choose a topic that involves nothing confidential or personal. Don't enter details about your employer, clients, students, or anyone's private information.

Your Task

Question an AI assistant on a narrow topic you know well until it states something false, then test whether rewording the question or asking "are you sure?" changes the answer.

Steps

  1. Ask five increasingly specific questions about your topic. Start general and narrow down. Move from "What is this?" toward names, dates, numbers, and sequences of events. Don't turn on web search if your assistant lets you choose.
  2. Mark each reply correct, partly correct, or wrong, from your own knowledge. Judge the content only. Note the specific statement that's wrong in any reply you don't mark correct.
  3. For one wrong reply, reword the question and ask again in a new conversation. Keep the meaning the same and change the wording. Note whether the false detail stays, changes to a different false detail, or is replaced by a correct one. If you found no wrong reply, do this with your most specific question.
  4. Go back to the first conversation and ask "Are you sure?" Note what the assistant does. It may confirm, apologize and change the answer, or hedge. If it changes the answer, check whether the new one is right.
  5. Compare the tone of the wrong and right replies. Read one of each side by side. Write one sentence on whether anything in the wording would have told you which was which.

What to Expect

General questions will probably be answered well. Errors become more likely as your questions reach details that few documents would mention. You may also find none, especially on a topic that's widely written about. That's a result worth recording, and it doesn't show that the assistant is reliable on the topic, only that five questions didn't find an error.

Rewording often changes a fabricated detail, because the detail was generated on the spot. "Are you sure?" can go either way. Assistants sometimes defend a wrong answer and sometimes abandon a right one, so the response to the challenge tells you little about the original answer.

Most people find no difference in tone between the wrong and right replies.

Self-Check

When you're done, check that:

  • You asked at least five questions, moving from general to specific
  • You found an error, or recorded that you didn't
  • You tested one rewording in a new conversation
  • You noted what "Are you sure?" did to the answer
  • You noted whether the tone of wrong and right answers differed

Nothing is uploaded. Write in your own notebook or document and keep it.

Report an issue with this item

Knowledge Check 10 min

Why models err, and what is known about their inner workings

This ungraded knowledge check assesses your understanding of why language models make confident errors and what is known about their inner workings. You'll be asked about hallucination and calibration, prompt sensitivity and unfaithful explanations, interpretability findings and their limits, and the dispute over whether models understand.

Note: Use this to test yourself, review the feedback on any questions you miss, and retry until you feel confident before moving forward.

5 questions · ungraded · retry as often as you like

Report an issue with this item