KnowledgeInSight
AI Literacy
0% of Course 1 complete

Course Project

Explain a Model's Answer

You'll take one real reply from an AI assistant and write a plain-language account of how it came to be, from training through to the words on your screen. It's worth doing because explaining one concrete case shows you what you understand and where your account still has gaps.

What you will be able to do

  • Explain how a specific AI reply was produced, from pretraining to generation
  • Separate what is established about that process from what is uncertain and from what the model can't report about itself

0% of this lesson · 3 items · 1h 10m total

Contents of this lesson3 items
  1. ReadingExplaining One Answer: Why a Single Case Tests What You Know3 min
  2. ReadingProject Requirements and Instructions7 min
  3. Writing AssignmentExplain a Model's Answer60 min

Reading 3 min

Explaining One Answer: Why a Single Case Tests What You Know

"Chatbots predict the next word." "They're trained on huge amounts of text." "They sometimes make things up." You've probably heard all three statements, and each is accurate as far as it goes. They're also easy to repeat without being able to apply them to anything.

A single reply is harder. Suppose a colleague forwards you an answer an AI assistant wrote and asks, "How does it come up with this, and can I trust it?" The general statements don't settle either question. Your colleague is asking about this reply: where its facts came from, why it's laid out as a tidy list, why it opens with a friendly sentence, and whether the date in its third paragraph is right.

Answering means assigning parts of the reply to causes. The broad knowledge in it most likely traces to pretraining, the first and largest stage of training, in which a model learns to predict the next piece of text across a very large collection. The Royal Society's report on machine learning describes the general approach: such systems learn from examples and data, where conventional programs follow rules written out step by step (Royal Society 2017, chap. 1). The helpful manner and the format more likely trace to later training and to instructions placed ahead of the conversation. The words themselves were produced one piece at a time. If the product searched the web, some of the detail came from pages fetched while you waited.

Each of those assignments is a choice, and a careful reader can press on it with three doubts.

  • How do you know? Some parts of your account rest on published research about how these systems are built and how they behave.
  • What are you inferring? Other parts are your best reading of the evidence in front of you. You can see the reply, and you can't see the training that led to it.
  • What can't anyone outside the developer know? The text a model was trained on and the instructions a product gives it are mostly unpublished.

The 2026 International AI Safety Report, written by an international panel of experts chaired by the computer scientist Yoshua Bengio, says that "the inner workings of models remain poorly understood" and that developers have "incentives to keep important information proprietary" (Bengio and others 2026, sec. 3.1). An account of one reply will have gaps, and a good one marks where they are.

Asking the assistant how it produced the reply doesn't close those gaps. Its answer is more generated text, produced the same way as the reply you're trying to explain, and the model can't look inside itself and report what happened there.

An explainer can fail your colleague in two directions. If it overclaims, they'll trust the tool too much. If it's vague, they'll learn nothing. In this project you'll choose one real reply from an AI assistant, check one detail in it, and write a short explainer of how it came to be, with your claims sorted by how sure anyone can be of them.

References

  • Bengio, Yoshua, and others. 2026. International AI Safety Report 2026. DSIT 2026/001. Published February 3, 2026.
  • Royal Society. 2017. Machine Learning: The Power and Promise of Computers That Learn by Example. London: The Royal Society.

Report an issue with this item

Reading 7 min

Project Requirements and Instructions

Deliverable Overview

You'll write an explainer of 600 to 900 words about one real reply from an AI assistant. It gives a plain-language account of how that reply was produced, reports one detail you checked, and sorts your claims by how sure anyone can be of them.

Write for a colleague with no technical background who has seen the reply and asked, "How does it come up with this, and can I trust it?" Nothing is submitted, and nobody grades it. You check your own work against the success criteria below. Plan on about an hour.

Detailed Requirements

The explainer has four components. The 600 to 900 words cover your own writing. The prompt and reply you quote don't count toward the range.

The Prompt and the Reply

Open with the prompt you typed and the reply you got. Quote both in full if they're short. If the reply runs past about 150 words, summarize it and quote the two or three sentences your account discusses. Say whether the product showed that it searched the web or used another tool, for example by listing sources or links.

Use a prompt on a subject you can check, with nothing confidential or personal in it. Don't enter details about your employer, clients, students, or anyone's private information.

The Four-Stage Account

Explain the reply in four stages, one short paragraph each.

StageWhat to explain
What pretraining suppliedPretraining is the first and largest stage of training, in which a model learns to predict the next piece of text across a very large collection. Say which parts of the reply most likely rest on it, such as general knowledge, vocabulary, and the standard form of this kind of text
What fine-tuning and instructions shapedFine-tuning is further training of an already-trained model on a smaller, chosen set of examples, in order to shift its behavior. System instructions are text placed ahead of a conversation that tells the model how to behave in that setting. Say which features of manner and format most likely come from these, such as answering the request, the tone, the layout, or a caution added at the end
How it was generatedThe model produced the reply one token at a time. A token is a piece of text, often a word or part of a word. At each point the model rated the possible next tokens, and one was chosen with an element of chance. Say what that means for this reply, including that the same prompt could give a different reply
What any search or tools addedIf the product searched the web or used another tool, say which details most likely came from the fetched text. If it didn't, or you can't tell, say so in one or two sentences
The Check

Pick one specific detail in the reply, such as a name, date, figure, quotation, or source. Check it against an outside source you trust: a reference work, an official page, or the original document. Report what you checked, where, and what you found, whether the detail was right, wrong, or impossible to confirm.

The Three Lists

End with three labeled lists of at least two items each.

  • Established. Claims supported by published research or by your own check. Example: the reply was generated one token at a time.
  • Uncertain. Claims that are your inference, or that depend on information the developer hasn't published. Example: whether the bullet-point layout comes from training or from the product's instructions.
  • Not knowable from the model's own account. Things you could ask the assistant about but couldn't confirm from its answer. Example: why it gave one figure instead of another.

Success Criteria

Your explainer is complete when:

  • The prompt and reply are included or summarized
  • All four stages are covered, each in plain language
  • Every technical term is defined on first use
  • One detail of the reply was checked against an outside source, and the result is reported
  • The three lists are present, with at least two items each
  • Nothing rests on the assistant's own account of how it works
  • A reader with no technical background could follow it

Step-by-Step Guidance

  1. Choose a prompt on a subject you can check, and save the reply. Ask a question whose answer runs to a paragraph or two and contains specific details. Good subjects include the history of a place you know, the rules of a game, or a standard procedure from your field described in general terms. Use nothing confidential. Copy the prompt and the reply into a document, and note whether the product showed any sources. Allow about 10 minutes.
  2. Mark up the reply. Go through it and label each part as one of three kinds. General knowledge is what many texts on the subject would say. Manner and format covers tone, layout, and any offer or caution. Specific details are the names, dates, figures, and sources. This takes about 5 minutes and gives you the raw material for the four stages.
  3. Draft the four-stage account, one short paragraph per stage. Tie each paragraph to something you marked. "The reply lists the steps in the usual order, which is the kind of pattern pretraining supplies" is stronger than a general sentence about pretraining. Use "most likely" where you're inferring. Allow about 20 minutes.
  4. Check one specific detail against an outside source and record what you found. Choose a detail that would matter if it were wrong. Write down the source you used and the result. Don't ask the assistant to confirm its own detail, since a confirmation is generated the same way the detail was. Allow about 10 minutes.
  5. Write the three lists. Go through your draft claim by claim and put each under established, uncertain, or not knowable from the model's own account. If a list has fewer than two items, look again at your draft. An account with nothing uncertain in it usually claims too much. Allow about 5 minutes.
  6. Revise for a reader with no technical background. Define each term the first time you use it, in a clause or a short sentence. Cut any claim you can't support with published research, your check, or a clearly marked inference. Then count your words. Allow about 10 minutes.

Resources

You need an AI assistant you already use, somewhere to write, and one outside source for the check. Draw on what you know about how language models work. Notes of your own are welcome if you have them.

The two works in the References below are free to read and optional. If you cite a source in your explainer, give its author or publisher, its title, and its date, with a link if it's online. That's enough for this purpose.

Common Pitfalls

Quoting the assistant's explanation of itself as evidence. If you ask an assistant how it produced a reply, the answer may describe the general mechanism accurately. It still isn't a report from inside the model. A 2023 study by researchers at New York University and two AI companies, Cohere and Anthropic, found that models' step-by-step explanations of their answers "can be plausible yet misleading" (Turpin et al. 2023, abstract). Build your account from what you know and what you checked.

Describing a specific product's hidden instructions as fact. Most products don't publish their system instructions. You can say that a feature of the reply "is consistent with an instruction to keep answers brief." You can't say what the instructions contain. Put any such claim in your uncertain list.

Saying the model "looked up" or "knew" something. A model doesn't consult its training text when it replies. What training left behind is a set of parameters, the adjustable numbers whose values training sets, and the reply was generated from those. "Looked up" is accurate only for text a search or other tool fetched during the conversation. For the rest, write "the model produced" or "the reply states."

Leaving out the check. The check is the only part of the explainer that tests the reply against the world. It also answers the second half of your colleague's question. Capable systems may still generate "non-existent citations, biographies, or facts" (Bengio and others 2026, sec. 1.2), and the tone of a reply is no guide to which details are sound. One checked detail doesn't show that the whole reply is reliable or unreliable. Report it as one result.

References

  • Bengio, Yoshua, and others. 2026. International AI Safety Report 2026. DSIT 2026/001. Published February 3, 2026.
  • Turpin, Miles, Julian Michael, Ethan Perez, and Samuel R. Bowman. 2023. "Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting." arXiv:2305.04388. NeurIPS 2023.

Report an issue with this item

Writing Assignment 60 min

Explain a Model's Answer

Overview

You'll write a short explainer of how an AI assistant produced one specific reply, for a colleague who asked how the tool comes up with its answers and whether to trust them. The work stays with you. Nobody collects or grades it, and you judge it yourself against the Self-Check below.

Your Task

Write a 600 to 900 word plain-language explainer of how an AI assistant produced one specific reply. Cover training, shaping, generation, and any tools, and state what is established, what is uncertain, and what the model can't tell you about itself.

Use a prompt on a subject you can check. Don't enter anything confidential or personal, and no details about your employer, clients, or students.

Tips

  • Pick a reply short enough to explain in full. Two paragraphs with a few specific details give you plenty to work with.
  • Keep pointing at the reply. Each stage of your account should name a sentence or feature of it.
  • Mark inferences as inferences. "Most likely" and "is consistent with" are accurate, and your reader needs to see which claims are firm.
  • Report the check whatever it showed. A detail that turned out right is as useful to your reader as one that turned out wrong.
  • Draw on notes or practice work of your own if you have any.
  • Sources are welcome. Give the author or publisher, the title, and the date, with a link if it's online.
  • Read the finished explainer as your colleague would. Wherever you'd have to stop and ask what a word means, add a definition.

Self-Check

When you're done, check that:

  • The prompt and reply are included or summarized
  • All four stages are covered, each in plain language
  • Every technical term is defined on first use
  • One detail of the reply was checked against an outside source, and the result is reported
  • The three lists are present, with at least two items each
  • Nothing rests on the assistant's own account of how it works
  • A reader with no technical background could follow it

Nothing is uploaded. Write in your own notebook or document and keep it.

Report an issue with this item