KnowledgeInSight
AI Literacy
0% of Course 2 complete

Module 1 · Lesson 2

Evaluating AI: Verifying claims, sources, and citations

You'll learn how AI systems produce citations and facts that look right and aren't, and how professional fact checkers verify a claim quickly. You'll be able to apply a short routine to any AI answer and decide how much checking it needs.

What you will be able to do

  • Apply a verification routine to the claims, sources, and citations in an AI output.

0% of this lesson · 9 items · 1h 3m total · 48m without the optional activity

Contents of this lesson9 items
  1. ReadingSix Court Cases That Didn't Exist: The Sanctions Order in Mata v. Avianca3 min
  2. ReadingFabricated and Erroneous Citations: How Often They Occur and What They Look Like4 min
  3. ReadingLateral Reading: Checking a Claim by Leaving the Page4 min
  4. ReadingA Verification Routine for Claims, Sources, Quotations, and Numbers4 min
  5. ReadingDeciding How Much to Check: Stakes, Specificity, and Your Own Expertise4 min
  6. Guided ReadingGuided Walkthrough: Verifying an AI-Written Paragraph with Three Citations7 min
  7. Guided ConversationPlan the Checks for an Answer You'd Rely On12 min
  8. Hands-on Activity · optionalVerify Five Claims from One AI Answer15 min
  9. Knowledge CheckVerifying claims, sources, and citations10 min

Reading 3 min

Six Court Cases That Didn't Exist: The Sanctions Order in Mata v. Avianca

In June 2023 a federal judge in New York fined two lawyers and their law firm $5,000. They had filed a legal brief that cited court decisions which no court had ever issued. A chatbot had made them up. The story was widely reported, and it's often told as a warning against using AI at all. The judge's order supports a narrower reading.

The case began as an ordinary injury claim. A passenger, Roberto Mata, sued the airline Avianca, saying a metal serving cart had struck his knee during a flight. When the airline asked the court to dismiss the claim, Mata's lawyers filed a response that cited earlier decisions in support of their argument (Mata v. Avianca 2023).

The airline's lawyers couldn't find several of those decisions. Neither could the judge, P. Kevin Castel of the US District Court for the Southern District of New York. He ordered Mata's lawyers to supply copies. They filed documents that looked like excerpts of the opinions. Six of the cases turned out not to exist. They had names, courts, dates, and quoted passages, and all of it had been generated by ChatGPT.

The penalty was for what the lawyers did after the citations were questioned. In the judge's words, they "continued to stand by the fake opinions after judicial orders called their existence into question." The order treats the use of AI as a separate matter: "There is nothing inherently improper about using a reliable artificial intelligence tool for assistance" (Mata v. Avianca 2023).

One detail shows how the mistake survived. The lawyer who had done the research did try to check. He asked the chatbot whether one of the cases was real, and it said yes. He asked whether the others were fake, and it said they were real too and could be found in standard legal databases (Mata v. Avianca 2023). He treated that reply as confirmation. Nobody looked the cases up in one of those databases until the court insisted.

This reading describes one US court's sanctions order. It isn't legal advice, and other courts and professions set their own rules.

The same kind of error can come from any chatbot built on a language model, whether it's made by OpenAI, Google, Anthropic, or another developer. The lesson of the order applies well beyond law. A citation in an AI answer looks like evidence. It has an author, a title, and a date, and it arrives in the format you'd expect from a careful source. The order shows that all of those features can be present when the source doesn't exist.

A citation from an AI system is therefore better treated as a claim. It asserts that a source exists and that the source says a particular thing. Both parts can be checked, and neither can be checked by asking the system that produced them.

References

  • Mata v. Avianca, Inc. 2023. Opinion and Order on Sanctions, No. 22-cv-1461 (PKC), US District Court, Southern District of New York, June 22, 2023.

Report an issue with this item

Reading 4 min

Fabricated and Erroneous Citations: How Often They Occur and What They Look Like

This content reflects the field as of October 2026.

Introduction

Stories about chatbots inventing sources have been in the news since 2023. They leave two practical questions open: how often this happens, and whether a made-up reference looks any different from a real one.

This reading gives the standard term for the problem, reports what one careful study measured, separates the kinds of failure, and explains why none of them is visible on the page.

The Term: Confabulation

The National Institute of Standards and Technology (NIST), the US government's standards agency, published a profile of the risks of generative AI in 2024. One of the twelve risks it names is confabulation: the production of confidently stated but false content by an AI system. The profile's own wording is "confidently stated but erroneous or false content," and it notes that the same thing is commonly called "hallucination" or "fabrication" (NIST 2024, "Confabulation").

The word "confidently" carries the weight. A confabulated statement comes in the same fluent, assured prose as a correct one.

What One Study Measured

In 2023, William Walters of Manhattan College and Esther Wilder of the City University of New York tested how reliable chatbot citations were. They had ChatGPT write short papers with references on 42 topics, using two versions of the underlying model, GPT-3.5 and the newer GPT-4. They then checked all 636 citations by hand (Walters and Wilder 2023).

Older version (GPT-3.5)Newer version (GPT-4)
Citations that were fabricated55 percent18 percent
Real citations with substantive errors43 percent24 percent

With the older version, more than half the references pointed to works that didn't exist. The newer version did much better, and nearly one reference in five was still invented. The authors' summary is that the newer version was a major improvement and that "problems remain" (Walters and Wilder 2023).

Three Kinds of Failure

The study separates two problems, and a third belongs beside them.

A fabricated citation is a reference to a source that doesn't exist. No such article, book, or court decision was ever published.

A citation error is a mistake in the details of a real reference, such as a wrong year, volume, or page number. Walters and Wilder found that wrong numbers of this kind were the most common errors in the real citations (Walters and Wilder 2023).

Misattribution is crediting a real source with a claim it doesn't make. The article exists and the details are right, and it doesn't say what the AI answer says it does. The study didn't measure this, since it checked the references and not how each was used. NIST's profile points to the same risk when it warns of confabulated "citations that purport to justify" an answer (NIST 2024, "Confabulation").

Why Fabricated Citations Look Real

A language model produces text by predicting what plausibly comes next. A plausible reference has a recognizable shape: authors, a year, a title that fits the topic, a journal, page numbers. The model can produce that shape without having any particular source behind it.

The parts are often real. Walters and Wilder report that most of the fabricated citations included the names of real journals, publishers, and organizations (Walters and Wilder 2023). A fabricated reference can pair a real journal with a believable title and a likely year. Nothing on its surface separates it from a genuine one. In the study, each one had to be looked up.

Have Things Improved?

The figures above describe two model versions from 2023, and they shouldn't be read as rates for the assistants you use now. The drop from 55 to 18 percent between versions shows the direction of change within the study itself.

Since then, assistants from OpenAI, Anthropic, Google, and other developers have been connected to web search. An assistant that searches can link to pages it retrieved, which makes a wholly invented reference less likely. It doesn't remove the other failures. An assistant can link to a real page and still misstate what the page says.

The reasonable position is that the problem is reduced and not gone. The rate for your subject and your kind of question is something only checking will show.

Conclusion

AI systems can produce references that don't exist, real references with wrong details, and real references credited with things they don't say. A 2023 study found the first two at high rates in one chatbot, with clear improvement in the newer version. Rates have fallen since and remain above zero, and a wrong citation looks the same on the page as a right one.

Key Terms

  • Confabulation: The production of confidently stated but false content by an AI system.
  • Fabricated citation: A reference to a source that doesn't exist.
  • Citation error: A mistake in the details of a real reference, such as a wrong year, volume, or page number.
  • Misattribution: Crediting a real source with a claim it doesn't make.

References

  • National Institute of Standards and Technology. 2024. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1. Gaithersburg, MD: NIST, July 2024.
  • Walters, William H., and Esther Isabelle Wilder. 2023. "Fabrication and Errors in the Bibliographic Citations Generated by ChatGPT." Scientific Reports 13: 14045.

Report an issue with this item

Reading 4 min

Lateral Reading: Checking a Claim by Leaving the Page

Introduction

Most of us were taught to judge a source by examining it closely: read it carefully, look at who wrote it, see whether it seems professional. A study of how experts evaluate websites found that the people who are best at this do something else.

This reading describes that study, the two ways of reading it identified, and why the better one suits AI output.

The Stanford Study

Sam Wineburg and Sarah McGrew, education researchers then at Stanford University, watched three groups evaluate unfamiliar websites on social and political topics. The groups were 10 historians with doctorates, 10 professional fact checkers, and 25 Stanford undergraduates. Each person thought aloud while working, and the researchers recorded what they did (Wineburg and McGrew 2019).

The groups behaved differently from the first minute. The historians and the students mostly stayed on the page they'd been given. They read its text, looked at its design, and checked its "About" section. The fact checkers scanned the page briefly and then left it. They opened new browser tabs and searched for what other sources said about the organization behind the site.

The fact checkers reached better conclusions, and they reached them sooner. In one task, people compared two websites about children's health. One belonged to a large, long-established professional association of pediatricians. The other belonged to a much smaller advocacy group with a similar name. Every fact checker judged the large association's site the more reliable. The historians were divided, and most of the students favored the advocacy group's site or couldn't choose.

In another task, the fact checkers needed on average less than a minute to find out who funded a website about the minimum wage. The other two groups took several times as long, and many of them never found out (Wineburg and McGrew 2019).

Vertical and Lateral Reading

The researchers gave the two approaches names.

Vertical reading is evaluating a source by staying on it and examining its own content and appearance. The reader moves up and down the page.

Lateral reading is evaluating a source by leaving it and checking what other sources say about it. The reader moves sideways, across tabs.

Vertical readingLateral reading
Where you lookAt the source itselfAt other sources
What you rely onThe source's own account and appearanceWhat independent sources report
Main weaknessA source controls how it presents itselfTakes a search, and the other sources must be judged too

Vertical reading fails for a simple reason. Everything on a page was put there by the people behind it. A professional layout, a list of references, and a confident "About" statement are all easy to produce, whoever is producing them. Careful reading of the page tells you how the source wants to be seen.

Why This Fits AI Output

An AI answer is a page of the same kind. Its tone, its structure, and its citations were all produced by the same system, and they look the same whether the content is right or wrong. Studying the answer more closely doesn't help, and asking the system about its own answer amounts to reading the "About" section.

Lateral reading supplies what the answer can't. Source evaluation, the work of judging whether a source can be trusted on a particular point, has to draw on something outside the source. For an AI answer, that means leaving the conversation and looking the claim up.

The aim is corroboration: confirmation of a claim by a source that is independent of the one that made it. Two AI assistants agreeing with each other is weak corroboration, since both may have learned from the same mistaken text. A primary document, a reference work, or an expert's published account is stronger.

Three Questions

The Digital Inquiry Group, a nonprofit that grew out of the Stanford research group behind the study, turned these findings into a free curriculum called Civic Online Reasoning. It's built around three questions (Digital Inquiry Group n.d.):

  1. Who's behind the information?
  2. What's the evidence?
  3. What do other sources say?

For an AI answer, the first question has an odd reply. The author is a model that predicts plausible text, so the question shifts to the sources the answer names. The second and third questions apply as written. You look for the evidence behind a claim, and you look for it somewhere other than the answer.

Conclusion

In the Stanford study, professional fact checkers judged websites faster and more accurately than historians and students by leaving each page and checking what others said about it. That approach is lateral reading, and its opposite, vertical reading, fails because a source controls its own presentation. An AI answer gives no reliable sign of its own accuracy, so checking it means going outside the conversation.

Key Terms

  • Vertical reading: Evaluating a source by staying on it and examining its own content and appearance.
  • Lateral reading: Evaluating a source by leaving it and checking what other sources say about it.
  • Source evaluation: The work of judging whether a source can be trusted on a particular point.
  • Corroboration: Confirmation of a claim by a source that is independent of the one that made it.

References

  • Digital Inquiry Group. n.d. "Teaching Lateral Reading." Civic Online Reasoning. Accessed October 3, 2026.
  • Wineburg, Sam, and Sarah McGrew. 2019. "Lateral Reading and the Nature of Expertise: Reading Less and Learning More When Evaluating Digital Information." Teachers College Record 121 (11): 1–40.

Report an issue with this item

Reading 4 min

A Verification Routine for Claims, Sources, Quotations, and Numbers

Introduction

"Always check AI output" is common advice, and it's hard to follow because it doesn't say what to check or how. People who verify information for a living work from a fixed order of steps, which keeps the job short.

This reading sets out a routine in that spirit: one step to decide what needs checking, then specific checks for sources, quotations, and numbers, and a rule for what to do when a check fails.

Step One: List the Checkable Claims

A checkable claim is a statement that could be shown true or false by looking something up. "Remote work rose after 2020" is checkable. "Remote work is a mixed blessing" is an opinion, and no lookup will settle it.

Read the answer once and list its checkable claims. Then mark the ones the answer depends on. Those are the claims that would change your conclusion or your next action if they were wrong. A wrong date in a passing remark may not matter. A wrong figure that the whole recommendation rests on does.

Sources: Two Questions

Every source named in an AI answer raises two questions, and they have to be asked in order.

  1. Does it exist? Search for the title and the author outside the conversation, in a search engine, a library catalog, or the publisher's own site. A 2023 study checked 636 citations produced by one chatbot and found that 55 percent from an older model version and 18 percent from a newer one were fabricated (Walters and Wilder 2023). Those figures describe 2023 models, and the rates have fallen since. They haven't reached zero.
  2. Does it say this? Open the source and find the passage that supports the claim. A real article can be cited for something it never said.

Where you can, go to the primary source: the original document in which a claim, finding, or statement first appeared. A news report about a study is a step removed from the study, and errors enter at each step.

Quotations and Numbers

A quotation is a claim that a particular person wrote or said particular words. Search for a distinctive phrase from the quotation, in quotation marks, and look for it in the original.

AI systems often produce a quotation that captures the gist of what someone said in words the person never used. If you can find the idea in the source and not the words, the honest fix is to drop the quotation marks and paraphrase.

A number can be wrong in more ways than a sentence can. For each figure that matters, find four things in the source.

What to findWhy it matters
The figure itselfIt may be misremembered or invented
Its unitPercent and percentage points, or thousands and millions, are easy to swap
Its dateA correct figure from ten years ago may be presented as current
What it counts"82 percent of managers surveyed" is a different claim from "82 percent of companies"

Why "Are You Sure?" Isn't a Check

It's tempting to verify an answer by asking the assistant to confirm it. The reply tells you very little.

A 2023 study by researchers at Anthropic tested five AI assistants from Anthropic, OpenAI, and Meta. One test gave an assistant a factual question, waited for its answer, and then challenged the answer. The assistants often changed a correct answer to an incorrect one when challenged, and one of them wrongly admitted to mistakes on nearly every question (Sharma et al. 2023). Developers have worked on this behavior since 2023, and it's reduced in current assistants without having disappeared.

A challenge can produce an apology where the first answer was right. It can also produce reassurance where the first answer was wrong, because the same process that generated the error generates the confirmation.

When a Claim Can't Be Traced

Traceability is the ability to follow a claim back to a source you can open and read. Some claims fail this test. The source can't be found, or it exists and doesn't contain the claim.

An unverified claim is one you haven't been able to confirm from a source outside the conversation. It may be true. Until you know, leave it out, or keep it and say plainly that you couldn't confirm it.

Conclusion

The routine has a fixed order. List the checkable claims and mark the ones that matter, confirm that each source exists and says what's claimed, find quotations word for word, and pin down each number's value, unit, date, and meaning. Asking the assistant to confirm itself isn't one of the checks. Anything that can't be traced is labeled as unverified.

Key Terms

  • Checkable claim: A statement that could be shown true or false by looking something up.
  • Primary source: The original document in which a claim, finding, or statement first appeared.
  • Traceability: The ability to follow a claim back to a source you can open and read.
  • Unverified claim: A claim you haven't been able to confirm from a source outside the conversation.

References

  • Sharma, Mrinank, Meg Tong, Tomasz Korbak, and 16 others. 2023. "Towards Understanding Sycophancy in Language Models." arXiv:2310.13548. ICLR 2024.
  • Walters, William H., and Esther Isabelle Wilder. 2023. "Fabrication and Errors in the Bibliographic Citations Generated by ChatGPT." Scientific Reports 13: 14045.

Report an issue with this item

Reading 4 min

Deciding How Much to Check: Stakes, Specificity, and Your Own Expertise

Introduction

Nobody can verify every sentence an AI assistant produces, and nobody should try. Checking everything would cost more time than the assistant saved. Checking nothing leaves you with errors you don't know about.

This reading covers three things that should set how much you check, and one that shouldn't: how confident the answer sounds.

Stakes

Stakes are what would be lost, and by whom, if a claim turned out to be wrong. Two questions bring them out: what happens if this is wrong, and who does it happen to.

A wrong suggestion in a list of ideas for a team outing costs almost nothing. A wrong dosage, a wrong deadline in a contract, or a wrong statistic in a published report can cost a good deal, and the cost may fall on someone other than you.

Stakes also depend on whether the error can be undone. A mistake in a draft that a colleague will review can be caught. A mistake in an email already sent to two hundred people can't.

Specificity

Specificity is how exact and particular a claim is. "Many companies adopted remote work" is general. "A 2021 study in a named journal found a 9 percent rise in output" is specific.

Errors concentrate in specific claims. Names, dates, figures, quotations, and citations each have exactly one right form and many wrong ones. A language model produces the form that seems most plausible, and for precise details the plausible form and the correct form often differ.

Specific claims are also the easiest to look up. A date or a citation takes a minute.

Your Own Expertise

What you already know changes both the risk and the cost of checking.

In a field you know well, you'll notice many errors on a first read, and you'll know where to look to confirm the rest. Checking is quick.

In a field you don't know, the same answer reads as equally convincing, and you have fewer ways to catch a mistake. This is where people lean on AI most, since they're asking because they don't know. You're therefore most exposed where you know least.

One piece of evidence shows that expertise doesn't remove the risk. In a field experiment with 758 consultants, a test with real workers randomly assigned to work with or without AI, trained professionals using a model were less likely to reach the correct answer on a task that was beyond the model's ability than colleagues working alone (Dell'Acqua et al. 2026). Professional training didn't protect them.

Over-Reliance

The US government's standards agency, the National Institute of Standards and Technology (NIST), describes this pattern in its 2024 profile of generative AI risks. Under a risk it calls "Human-AI Configuration," the profile warns that people may "over-rely" on these systems, or judge their content to be of higher quality than it is (NIST 2024, "Human-AI Configuration").

Over-reliance is trusting a system's output more than its accuracy warrants. It grows with fluency. An answer that sounds certain invites less checking, and how certain an answer sounds has no dependable link to whether it's correct. A model will state a fabricated citation and a real one in the same tone.

A Minimum Check for Each Kind of Claim

Claim typeTypical riskMinimum check
Citation or named sourceThe source doesn't exist, or doesn't say thisFind it by title, then find the passage
QuotationThe words were never said or writtenSearch for the exact phrase in the original
Number or dateWrong value, unit, or yearFind the figure in a primary source
General factual statementOut of date or oversimplifiedConfirm in one independent reference
Summary of a document you suppliedSomething left out or addedCompare against the document on the key points
Idea, outline, or wording suggestionLow: you judge it by reading itNone beyond your own judgment

For long outputs with low or moderate stakes, a spot check is often enough: verifying a small sample of claims to judge how far the rest can be trusted. If three specific claims chosen from different parts of an answer all hold, the rest is more likely to be sound. If one fails, the whole answer needs a closer look.

A spot check doesn't suit high-stakes work. There, every claim the conclusion depends on gets checked.

Conclusion

How much to check depends on the stakes, on how specific the claim is, and on how well you know the field. It doesn't depend on how confident the answer sounds. Citations, quotations, and numbers always get a look, low-stakes suggestions rarely need one, and a spot check covers the ground in between.

Key Terms

  • Stakes: What would be lost, and by whom, if a claim turned out to be wrong.
  • Specificity: How exact and particular a claim is.
  • Over-reliance: Trusting a system's output more than its accuracy warrants.
  • Spot check: Verifying a small sample of claims to judge how far the rest can be trusted.

References

  • Dell'Acqua, Fabrizio, Edward McFowland III, Ethan Mollick, Hila Lifshitz, Katherine C. Kellogg, Saran Rajendran, Lisa Krayer, François Candelon, and Karim R. Lakhani. 2026. "Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality." Organization Science 37 (2): 403–423.
  • National Institute of Standards and Technology. 2024. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1. Gaithersburg, MD: NIST, July 2024.

Report an issue with this item

Guided Reading 7 min

Guided Walkthrough: Verifying an AI-Written Paragraph with Three Citations

Introduction

A paragraph with three citations looks well supported. Checking it takes a few minutes when the checks are done in a fixed order, and the result is often a shorter and more accurate paragraph.

This walkthrough follows one such check from start to finish. The paragraph, its three sources, the authors' names, and every figure in it were invented for this exercise. None of them refers to a real study, and nothing here is evidence about remote work.

The Starting Point

Daniel manages a department and is writing a short proposal to let his team work from home two days a week. He asks an AI assistant for a paragraph on remote work and productivity, with sources. This is what he gets.

Remote work raises productivity. A two-year study of 800 call-center employees found that those working from home completed 9 percent more calls than their office-based colleagues (Halvorsen and Adeyemi 2021). A national survey found that 82 percent of managers saw output rise after their teams went remote (Brightwater Institute 2022). Remote workers are also far less likely to quit, with turnover falling by half (Marchetti 2020).

The paragraph is fluent, specific, and exactly what he hoped to read. The proposal will go to his director, who may act on it.

Walking Through the Check

Step 1: List the checkable claims and mark the ones the paragraph depends on

Daniel goes through the paragraph sentence by sentence and lists what could be looked up.

  1. Remote work raises productivity.
  2. A two-year study of 800 call-center employees found 9 percent more calls completed at home (Halvorsen and Adeyemi 2021).
  3. A national survey found 82 percent of managers saw output rise (Brightwater Institute 2022).
  4. Turnover fell by half among remote workers (Marchetti 2020).

Claim 1 is the conclusion, and it rests on the other three. Claims 2, 3, and 4 each carry a citation and a number, so all three need checking. He marks claims 3 and 4 as the ones his proposal leans on most, since his director cares about managers' experience and about staff leaving.

Step 2: Search for each source by title and author outside the conversation

He doesn't ask the assistant whether the sources are real. He opens a browser and searches for each one by its authors and by the title the assistant gave.

  • Halvorsen and Adeyemi 2021. A search finds the article on a journal's website, with the same authors, year, and title.
  • Brightwater Institute 2022. A search finds the survey report on the institute's own site.
  • Marchetti 2020. A search for the title returns nothing. He tries the journal the assistant named and looks through its contents for 2020. No such article appears. A researcher called Marchetti does publish in a nearby field, on a different subject, which may be why the name seemed believable. He searches a library catalog as well and finds nothing.

Two sources exist. One can't be found.

Step 3: For the sources that exist, find the passage that supports the claim

Finding a source is half the check. Daniel now opens each one and looks for the specific claim.

In the Halvorsen and Adeyemi article, the summary on the first page reports a two-year study of 800 call-center employees at one company, with home-based staff completing 9 percent more calls. The figure, the unit, and what it counts all match. He also notes a detail the paragraph left out: the study covered a single company and a single kind of work.

In the Brightwater report, he searches for "82." The figure is there, and the sentence around it reads differently from the paragraph. The report says 82 percent of managers reported that output "stayed the same or rose." A table on the next page breaks this down. Only 31 percent said output rose. The other 51 percent said it stayed the same.

Step 4: Record a verdict for each

He writes one line per source.

SourceVerdictNote
Halvorsen and Adeyemi 2021ConfirmedFigure matches. One company, call-center work only
Brightwater Institute 2022Misdescribed82 percent said "stayed the same or rose." 31 percent said it rose
Marchetti 2020Not foundNo trace by title, in the journal's 2020 contents, or in a library catalog

The claim about turnover is now unverified. It might be true, and a real study might support it. Daniel has no source for it, so he can't present it as established.

Step 5: Rewrite the paragraph using only what was confirmed

He rewrites each sentence to say what its source supports and no more. The opening claim has to change too, because the evidence beneath it is now thinner than it looked.

Key Considerations

The verdict words are Daniel's working labels. "Confirmed" means he found the source and the passage. "Misdescribed" means the source is real and says something different. "Not found" means he couldn't trace the source, which is weaker than proving it doesn't exist.

The common mistake is to stop once a source is found to exist. If Daniel had stopped after Step 2, the Brightwater survey would have passed, and his proposal would have told his director that 82 percent of managers saw output rise. A study of 636 chatbot-generated citations found that many of the references that did exist still contained errors, alongside the ones that were invented outright (Walters and Wilder 2023). Existence and accuracy are separate checks.

Daniel also did all of his checking outside the conversation. Research on how professional fact checkers work found that they judge a source by leaving it and seeing what other sources say, and that this is faster and more accurate than studying the source itself (Wineburg and McGrew 2019). The assistant's paragraph can't vouch for its own citations.

Summary

Daniel listed the claims, searched for each source, read the passages, recorded three verdicts, and rewrote the paragraph. The original and rewritten versions are set side by side below.

Original sentenceRewritten sentenceVerdict behind the change
Remote work raises productivity.The evidence I could confirm on remote work and productivity is limited and mixed.Rests on the three claims below, of which one was confirmed
A two-year study of 800 call-center employees found that those working from home completed 9 percent more calls than their office-based colleagues (Halvorsen and Adeyemi 2021).A two-year study of 800 call-center employees at one company found that those working from home completed 9 percent more calls than their office-based colleagues (Halvorsen and Adeyemi 2021).Confirmed. "At one company" added from the source
A national survey found that 82 percent of managers saw output rise after their teams went remote (Brightwater Institute 2022).In a national survey, 31 percent of managers said output rose after their teams went remote, and a further 51 percent said it stayed the same (Brightwater Institute 2022).Misdescribed. Figures corrected from the report's table
Remote workers are also far less likely to quit, with turnover falling by half (Marchetti 2020).(Removed.)Not found. The claim is unverified

The rewritten paragraph: The evidence I could confirm on remote work and productivity is limited and mixed. A two-year study of 800 call-center employees at one company found that those working from home completed 9 percent more calls than their office-based colleagues (Halvorsen and Adeyemi 2021). In a national survey, 31 percent of managers said output rose after their teams went remote, and a further 51 percent said it stayed the same (Brightwater Institute 2022).

  1. The paragraph is shorter and makes a weaker case. That is the accurate result of the check, and Daniel's director can now rely on every sentence in it.
  2. The turnover claim is gone from the paragraph. If Daniel wants it back, he needs to find a real source for it.

References

  • Walters, William H., and Esther Isabelle Wilder. 2023. "Fabrication and Errors in the Bibliographic Citations Generated by ChatGPT." Scientific Reports 13: 14045.
  • Wineburg, Sam, and Sarah McGrew. 2019. "Lateral Reading and the Nature of Expertise: Reading Less and Learning More When Evaluating Digital Information." Teachers College Record 121 (11): 1–40.

Report an issue with this item

Guided Conversation 12 min

Plan the Checks for an Answer You'd Rely On

In this conversation you'll take a situation where you'd act on an AI answer and plan how you'd check it: which claims matter, and where you'd go to confirm each one. You'll leave with your own three-step minimum routine.

You'll have this conversation with an AI assistant, using your own account. Choose a button to open a new chat with the prompt already filled in, then press send to start. If the chat opens empty, copy the prompt and paste it in.

Run this conversation in whichever assistant you already use:

Claude desktop app

To use another LLM, simply copy and paste the prompt into its chat window.

Show the full prompt (it lists misreadings to watch for, so skip it if you would rather come to the conversation fresh)
Guided Conversation: Plan the Checks for an Answer You'd Rely On (about 12 minutes)

Note to the learner: press send to start. Everything below is facilitator guidance for the AI. It lists misconceptions to watch for, so skip it if you'd rather come to the conversation fresh.

Please facilitate a coached problem session with me. I'm an adult with no technical background who has used AI chatbots for everyday tasks, and I'm studying how to verify the claims, sources, and citations in an AI answer. Follow this guidance for the whole conversation.

GOAL
I can apply a verification routine to an AI output I might use, by planning which claims to check and where to check them.

HOW TO RUN THE CONVERSATION
- Ask one question at a time, then wait for my reply. Keep each of your turns under about 120 words.
- Don't lecture. Explain a point only when I need it to continue, then return to my situation.
- Be curious and collegial. Use plain words and define any technical term briefly on first use. Welcome disagreement when I give a reason.
- This is a coached problem. The problem is: plan the checks for one AI answer I would rely on. Ask for my own attempt at each part before you give any hint. Give one hint at a time. Don't write my routine for me.
- Plain conversation only: don't search the web or create files or documents.
- Don't ask for confidential, personal, or student information. I should describe my situation in general terms. If I start to share private details, remind me to leave them out.
- Aim for about 12 minutes. Spend most of the time on topics 2 and 3. If my replies are brief, offer one concrete prompt, such as "Think of the last time you copied a fact or figure from an AI answer into something you sent to someone," and move on. If I seem uncertain, shorten the conversation to 5-7 minutes. Always reach the final topic.
- Start now. Open with one or two warm sentences: this is a conversation, not a quiz; a plan I'd really follow matters more than a thorough one; I can ask you to clarify anything. Then ask me to describe a recent or likely situation where I would act on an answer from an AI assistant.

TOPICS, IN ORDER
1. The situation and the stakes. Ask what I'd do with the answer, what would happen if it were wrong, and who it would happen to. Follow up if I haven't said whether the mistake could be undone.
2. The checkable claims. Ask what kinds of claims an answer like that would contain: sources, quotations, numbers, dates, general statements. Ask which ones my decision would depend on. Draw out that specific details are where errors concentrate.
3. Where I'd check. For each claim that matters, ask where I would go outside this conversation to confirm it, and what I'd look for there. Draw out the two questions for any source: does it exist, and does it say this. Follow up on one claim where my plan is vague.
4. Closing. Ask me to state my own minimum routine in three steps. Tell me I can take it into a short optional activity where I check the claims in a real AI answer.

KEY POINTS TO KEEP ACCURATE
- Method: list the checkable claims and mark the ones the answer depends on; for each source, confirm that it exists and then that it says what's claimed; find quotations word for word; for numbers, find the figure, its unit, its date, and what it counts; label anything that can't be traced as unverified.
- Fluency and a confident tone are not evidence. A wrong answer reads the same as a right one.
- Checking happens outside the conversation, in a search engine, a library catalog, a publisher's site, or the original document.
- How much to check depends on the stakes, on how specific the claim is, and on how well I know the field.
- Your own reassurance is not verification. Say this about yourself plainly when it's relevant: if I ask you whether something is true or whether a source is real, your answer comes from the same process that could have produced an error.
- If I ask how you work, explain the general mechanism in one or two sentences and say plainly that you can't inspect your own internals, so your statements about yourself are not evidence.

MISCONCEPTIONS TO CORRECT GENTLY
When one appears, name the accurate version briefly, then return to my situation.
- "A citation means it's sourced": AI systems can produce citations to works that don't exist, so a citation is a claim to check.
- "Search-enabled assistants don't make things up": an assistant that searches can still link to a real page and misstate what it says.
- "Asking it to double-check is enough": when challenged, an assistant may simply agree with itself or reverse a correct answer, and neither tells me about the claim.

LIMITS
- Don't verify anything for me in this conversation, and don't offer to. Don't tell me whether any specific claim, source, or figure I mention is accurate. If I ask, say that checking it outside this conversation is the point.
- Don't search the web.
- Don't recommend or compare any AI product.
- Don't give legal, medical, or financial advice about my situation.
- Don't ask for or accept confidential details.

TO FINISH
After my closing answer, close in one short turn:
- Affirm one specific thing I worked out, in my own words where possible.
- Suggest one or two next steps that fit how the conversation went. Possible steps: apply my routine to one real AI answer and record a verdict for each claim; time how long the checks take; ask a colleague where they would check the same claims; review the difference between a source existing and a source saying what's claimed.
- Restate my routine on its own lines, labeled "My minimum routine", so I can copy it.

Report an issue with this item

Hands-on Activity 15 minOptional

Verify Five Claims from One AI Answer

Overview

Checking an AI answer is quicker than most people expect, and doing it once on a real answer shows you where the errors in your own field tend to be. In this activity you'll ask one question, pick five claims from the answer, and trace each one.

The activity is optional. Your notes are for you, and nobody collects them.

What You'll Need

  • An AI assistant you already use
  • A web browser or a library catalog
  • Somewhere to write a few notes, or any checking routine of your own that you've already drafted

Choose a question about your field in general. Don't enter confidential, personal, or student information.

Your Task

Ask an AI assistant a factual question in your field that calls for sources, then check five of its claims outside the conversation.

Steps

  1. Ask a question that needs specific facts and sources, and save the answer. Pick a topic you know reasonably well, and ask for sources by name. Copy the full answer into your notes before you do anything else.
  2. Pick five checkable claims, including every citation. Choose statements that could be shown true or false by looking them up: a named source, a quotation, a figure, a date. If the answer has more than five citations, check the ones the answer leans on most.
  3. Check each outside the conversation and record confirmed, misdescribed, or not found. Look up every citation by its title. If the source exists, open it and find the passage. Write the verdict and where you checked. Don't ask the assistant whether its own claims are right.
  4. Note how long the checking took and which claim was hardest to trace. A sentence on each is enough.

What to Expect

You may find that all five claims hold. You may find a source that exists and says something a little different from what the answer reported, which is the most common problem and the easiest to miss. Less often, you may find a source that you can't trace at all.

"Not found" means you couldn't trace it in the time you had. It's a reason to treat the claim as unverified, and it doesn't prove the source is invented.

Self-Check

When you're done, check that:

  • You checked five claims
  • Every citation was looked up by title
  • Each claim has a verdict and a note on where you checked
  • You recorded at least one thing you couldn't confirm, or noted that all five held

Nothing is uploaded. Write in your own notebook or document and keep it.

Report an issue with this item

Knowledge Check 10 min

Verifying claims, sources, and citations

This ungraded knowledge check assesses your understanding of how to verify what an AI answer tells you. You'll be asked about confabulation and fabricated citations, lateral reading, the steps of a verification routine, and how stakes and specificity set the amount of checking.

Note: Use this to test yourself, review the feedback on any questions you miss, and retry until you feel confident before moving forward.

5 questions · ungraded · retry as often as you like

Report an issue with this item