Deliverable Overview
You'll write an explainer of 600 to 900 words about one real reply from an AI assistant. It gives a plain-language account of how that reply was produced, reports one detail you checked, and sorts your claims by how sure anyone can be of them.
Write for a colleague with no technical background who has seen the reply and asked, "How does it come up with this, and can I trust it?" Nothing is submitted, and nobody grades it. You check your own work against the success criteria below. Plan on about an hour.
Detailed Requirements
The explainer has four components. The 600 to 900 words cover your own writing. The prompt and reply you quote don't count toward the range.
The Prompt and the Reply
Open with the prompt you typed and the reply you got. Quote both in full if they're short. If the reply runs past about 150 words, summarize it and quote the two or three sentences your account discusses. Say whether the product showed that it searched the web or used another tool, for example by listing sources or links.
Use a prompt on a subject you can check, with nothing confidential or personal in it. Don't enter details about your employer, clients, students, or anyone's private information.
The Four-Stage Account
Explain the reply in four stages, one short paragraph each.
| Stage | What to explain |
|---|
| What pretraining supplied | Pretraining is the first and largest stage of training, in which a model learns to predict the next piece of text across a very large collection. Say which parts of the reply most likely rest on it, such as general knowledge, vocabulary, and the standard form of this kind of text |
| What fine-tuning and instructions shaped | Fine-tuning is further training of an already-trained model on a smaller, chosen set of examples, in order to shift its behavior. System instructions are text placed ahead of a conversation that tells the model how to behave in that setting. Say which features of manner and format most likely come from these, such as answering the request, the tone, the layout, or a caution added at the end |
| How it was generated | The model produced the reply one token at a time. A token is a piece of text, often a word or part of a word. At each point the model rated the possible next tokens, and one was chosen with an element of chance. Say what that means for this reply, including that the same prompt could give a different reply |
| What any search or tools added | If the product searched the web or used another tool, say which details most likely came from the fetched text. If it didn't, or you can't tell, say so in one or two sentences |
The Check
Pick one specific detail in the reply, such as a name, date, figure, quotation, or source. Check it against an outside source you trust: a reference work, an official page, or the original document. Report what you checked, where, and what you found, whether the detail was right, wrong, or impossible to confirm.
The Three Lists
End with three labeled lists of at least two items each.
- Established. Claims supported by published research or by your own check. Example: the reply was generated one token at a time.
- Uncertain. Claims that are your inference, or that depend on information the developer hasn't published. Example: whether the bullet-point layout comes from training or from the product's instructions.
- Not knowable from the model's own account. Things you could ask the assistant about but couldn't confirm from its answer. Example: why it gave one figure instead of another.
Success Criteria
Your explainer is complete when:
- The prompt and reply are included or summarized
- All four stages are covered, each in plain language
- Every technical term is defined on first use
- One detail of the reply was checked against an outside source, and the result is reported
- The three lists are present, with at least two items each
- Nothing rests on the assistant's own account of how it works
- A reader with no technical background could follow it
Step-by-Step Guidance
- Choose a prompt on a subject you can check, and save the reply. Ask a question whose answer runs to a paragraph or two and contains specific details. Good subjects include the history of a place you know, the rules of a game, or a standard procedure from your field described in general terms. Use nothing confidential. Copy the prompt and the reply into a document, and note whether the product showed any sources. Allow about 10 minutes.
- Mark up the reply. Go through it and label each part as one of three kinds. General knowledge is what many texts on the subject would say. Manner and format covers tone, layout, and any offer or caution. Specific details are the names, dates, figures, and sources. This takes about 5 minutes and gives you the raw material for the four stages.
- Draft the four-stage account, one short paragraph per stage. Tie each paragraph to something you marked. "The reply lists the steps in the usual order, which is the kind of pattern pretraining supplies" is stronger than a general sentence about pretraining. Use "most likely" where you're inferring. Allow about 20 minutes.
- Check one specific detail against an outside source and record what you found. Choose a detail that would matter if it were wrong. Write down the source you used and the result. Don't ask the assistant to confirm its own detail, since a confirmation is generated the same way the detail was. Allow about 10 minutes.
- Write the three lists. Go through your draft claim by claim and put each under established, uncertain, or not knowable from the model's own account. If a list has fewer than two items, look again at your draft. An account with nothing uncertain in it usually claims too much. Allow about 5 minutes.
- Revise for a reader with no technical background. Define each term the first time you use it, in a clause or a short sentence. Cut any claim you can't support with published research, your check, or a clearly marked inference. Then count your words. Allow about 10 minutes.
Resources
You need an AI assistant you already use, somewhere to write, and one outside source for the check. Draw on what you know about how language models work. Notes of your own are welcome if you have them.
The two works in the References below are free to read and optional. If you cite a source in your explainer, give its author or publisher, its title, and its date, with a link if it's online. That's enough for this purpose.
Common Pitfalls
Quoting the assistant's explanation of itself as evidence. If you ask an assistant how it produced a reply, the answer may describe the general mechanism accurately. It still isn't a report from inside the model. A 2023 study by researchers at New York University and two AI companies, Cohere and Anthropic, found that models' step-by-step explanations of their answers "can be plausible yet misleading" (Turpin et al. 2023, abstract). Build your account from what you know and what you checked.
Describing a specific product's hidden instructions as fact. Most products don't publish their system instructions. You can say that a feature of the reply "is consistent with an instruction to keep answers brief." You can't say what the instructions contain. Put any such claim in your uncertain list.
Saying the model "looked up" or "knew" something. A model doesn't consult its training text when it replies. What training left behind is a set of parameters, the adjustable numbers whose values training sets, and the reply was generated from those. "Looked up" is accurate only for text a search or other tool fetched during the conversation. For the rest, write "the model produced" or "the reply states."
Leaving out the check. The check is the only part of the explainer that tests the reply against the world. It also answers the second half of your colleague's question. Capable systems may still generate "non-existent citations, biographies, or facts" (Bengio and others 2026, sec. 1.2), and the tone of a reply is no guide to which details are sound. One checked detail doesn't show that the whole reply is reliable or unreliable. Report it as one result.
References
- Bengio, Yoshua, and others. 2026. International AI Safety Report 2026. DSIT 2026/001. Published February 3, 2026.
- Turpin, Miles, Julian Michael, Ethan Perez, and Samuel R. Bowman. 2023. "Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting." arXiv:2305.04388. NeurIPS 2023.