KnowledgeInSight
AI Literacy
0% of Course 1 complete

Module 1 · Lesson 3

Machine Learning Basics: Tokens, meaning, and next-token prediction

You'll follow text into a language model and back out: how it's broken into tokens, how meaning is represented as position, and how a reply is produced by predicting one token after another. You'll be able to explain why the same question can get different answers.

What you will be able to do

  • Explain how a language model turns text into tokens and generates a response by predicting what comes next.

0% of this lesson · 10 items · 1h 33m total · 1h 18m without the optional activity

Contents of this lesson10 items
  1. ReadingOne Piece at a Time: How a Chatbot Composes a Reply3 min
  2. ReadingTokens and Tokenization: How Text Becomes Pieces a Model Can Process4 min
  3. ReadingEmbeddings: Words as Positions in a Space of Meaning4 min
  4. ReadingNext-Token Prediction as a Ranking of Possible Continuations4 min
  5. ReadingSampling, Temperature, and Why Identical Prompts Give Different Replies4 min
  6. Guided ReadingWorked Example: One Sentence from Prompt to Finished Reply7 min
  7. Guided ConversationPredict the Next Word Yourself12 min
  8. Hands-on Activity · optionalAsk the Same Question Five Times15 min
  9. Knowledge CheckTokens, meaning, and next-token prediction10 min
  10. Graded QuizMachine Learning and Language Models30 min

Reading 3 min

One Piece at a Time: How a Chatbot Composes a Reply

When you send a question to a chatbot, the reply usually arrives a few words at a time, as if someone were typing quickly. It's easy to take this for a visual effect, like the animated dots that show a friend is writing a text message.

In most products it's a fair picture of what the model is doing. The model doesn't compose a whole answer and then reveal it. It produces one small piece of text, then another, and each piece is chosen after the ones before it are already fixed. The pieces are called tokens, and a token can be a word or part of a word.

Stephen Wolfram, a computer scientist and the founder of the software company Wolfram Research, wrote a long explanation of one chatbot in 2023. He describes what it's doing as always trying to produce a "reasonable continuation" of the text so far. On his account the whole process comes down to asking the same question again and again: given the text up to this point, what should come next (Wolfram 2023, "It's Just Adding One Word at a Time").

That description is the source of a common dismissal. If a chatbot only predicts the next piece of text, then it's "just autocomplete," a scaled-up version of the suggestion bar on a phone keyboard.

The phrase is accurate about the mechanism. A language model does predict what comes next, and nothing else happens at the moment of writing. People who use the phrase are right that there's no separate fact-checking stage and no hidden draft.

The phrase misleads about what the mechanism can do. Consider what predicting the next word requires in different places.

  • After "The opposite of hot is," a good prediction needs word meanings.
  • After "She put the keys in her bag and later took them out of her," it needs to track what happened earlier in the sentence.
  • After the last line of a detective story, "And so the murderer was," it needs to have followed the plot.

A system that predicts well across all the text people write has to pick up a great deal about grammar, facts, and how arguments and stories go. A phone keyboard's suggestions come from a far smaller system that looks at a few recent words.

How far this goes is disputed. Some researchers hold that good enough prediction amounts to a kind of understanding, and others hold that it remains imitation of the surface of text. The mechanism itself isn't in dispute, and it explains several things you may have noticed. A reply can begin well and drift. A chatbot can state something false in a fluent, confident sentence. The same question can get two different answers on two occasions.

Each of these is easier to explain once you know that the reply was written one piece at a time, with every piece chosen as a continuation of the text before it.

References

  • Wolfram, Stephen. 2023. "What Is ChatGPT Doing … and Why Does It Work?" Stephen Wolfram Writings, February 14, 2023.

Report an issue with this item

Reading 4 min

Tokens and Tokenization: How Text Becomes Pieces a Model Can Process

Introduction

Chatbots that write fluent essays have been known to miscount the letters in a common word. AI companies also describe their models' limits and prices in "tokens," a unit most people have never needed before. Both facts come from the same place.

This reading explains what tokens are, why models use them in place of whole words, and what follows for you as a user.

Why Not Whole Words

A model works with numbers, so text has to be converted into numbers before the model can do anything with it. The obvious plan is to give every word its own number. That plan runs into trouble quickly.

  • New words appear all the time, and a fixed word list wouldn't contain them.
  • Names of people, places, and products are effectively endless.
  • People make typing mistakes, and a misspelled word matches nothing on a word list.
  • A model that serves many languages would need a word list for each one.

A model limited to whole words would have to treat every unfamiliar word as a blank.

Pieces Smaller Than Words

The solution is to work with smaller pieces. A token is a piece of text, often a word or part of a word, that a language model handles as one unit. Lee and Trott note that language models operate on "fragments of words called tokens" (Lee and Trott 2023, footnotes), and Wolfram says a token "could be just a part of a word" (Wolfram 2023, "It's Just Adding One Word at a Time").

Tokenization is the splitting of text into tokens. It happens before the model sees anything, and it follows a fixed procedure. Common words usually get a token of their own. Rarer words are split into several pieces. A piece of a word that serves as a token is called a subword.

Here is an illustration. The exact splits differ from one model to another, so treat these as typical and not as any particular model's output.

TextPossible tokensCount
thethe1
househouse1
unhelpfullyun, help, fully3
ZwickauZ, wick, au3

A word the model has never met can still be handled, because it can be built from familiar pieces. Punctuation marks and spaces are tokens or parts of tokens as well.

A Fixed List, and a Number for Each Entry

Every model has a vocabulary: the fixed list of all the tokens it can read and produce. The list is settled before training and doesn't change afterward. Wolfram reports a vocabulary of about 50,000 tokens for the model he describes (Wolfram 2023, "Inside ChatGPT"). Other models use larger or smaller lists.

Each token in the vocabulary has a number, which is its place on the list. When you send a message, your text is split into tokens and each token is replaced by its number. The model receives a sequence of numbers. When it replies, it produces numbers, and those are turned back into text for you to read.

The model never handles letters as such. It handles token numbers.

What You Can Observe

Two consequences show up in ordinary use.

Trouble with letters. A model receives a word as one token or a few, and the letters inside a token aren't visible to it as separate items. Counting the letters in a word, spelling it backward, or finding its third letter all require information the model doesn't directly have. Models often get these right anyway, because their training text included a great deal about spelling. When they get them wrong, tokenization is a large part of the reason.

Limits measured in tokens. The amount of text a model can take in at once is stated in tokens, and so is the price that companies charge for using a model through their developer services. In English a token averages somewhat less than a word, so a passage has more tokens than words. Text in a language that was less common in the model's training data tends to split into more tokens for the same content, which makes that text cost more and fill the limit sooner.

Conclusion

A language model works on tokens, which are pieces of text that are often smaller than words. Tokenization splits your text into entries from a fixed vocabulary and turns each into a number. This lets a model handle any text, including words it has never met, and it accounts for the model's occasional trouble with letters and for limits counted in tokens.

Key Terms

  • Token: A piece of text, often a word or part of a word, that a language model handles as one unit.
  • Tokenization: The splitting of text into tokens.
  • Subword: A piece of a word that serves as a token.
  • Vocabulary: The fixed list of all the tokens a model can read and produce.

References

  • Lee, Timothy B., and Sean Trott. 2023. "Large Language Models, Explained with a Minimum of Math and Jargon." Understanding AI, July 27, 2023.
  • Wolfram, Stephen. 2023. "What Is ChatGPT Doing … and Why Does It Work?" Stephen Wolfram Writings, February 14, 2023.

Report an issue with this item

Reading 4 min

Embeddings: Words as Positions in a Space of Meaning

Introduction

A chatbot treats "physician" and "doctor" as closely related even though the two words share almost no letters. Something inside the model must represent what words mean and not just how they're spelled.

This reading explains how a model represents each token as a position, how those positions come about, and what they do and don't capture.

Meaning from Use

Linguists have long observed that you can learn a lot about a word from the words that surround it. "Coffee" and "tea" turn up in similar sentences, near "cup," "drink," and "morning." "Coffee" and "carburetor" don't.

Language models use this observation. Each token is represented by an embedding: a list of numbers that works as the token's position in a space of meaning. Tokens that are used in similar ways get similar lists, so they sit near each other.

The mathematical name for such a list is a vector, a list of numbers that gives a position. Two numbers are enough to give a position on a map. A dimension is one of the numbers in a vector, and so one of the directions in the space. Embeddings have hundreds or thousands of dimensions. Lee and Trott describe a word vector 300 numbers long (Lee and Trott 2023, "Word vectors"). Nobody can picture a space like that, though the idea of nearness carries over from a map unchanged.

Neighbors

Semantic similarity is closeness in meaning, and in a model it shows up as nearness between embeddings. One way to see what a set of embeddings has captured is to pick a word and list its nearest neighbors.

WordWords found nearby
catdog, kitten, pet
TuesdayWednesday, Monday, Thursday
doctorphysician, nurse, surgeon
quicklyrapidly, swiftly, slowly

The first row is reported by Lee and Trott from a real set of word vectors (Lee and Trott 2023, "Word vectors"). The other rows are illustrations of the kind of result such systems give. The last row shows something useful: "slowly" is near "quickly" although it means the opposite, because the two words are used in the same kinds of sentences. Nearness reflects similar use, which often matches similar meaning and sometimes doesn't.

Directions Carry Relationships

In 2013, Tomas Mikolov and three colleagues at Google reported something further. The positions of words encode relationships between them, in addition to similarity (Mikolov et al. 2013). The paper is a preprint, meaning it was posted publicly without a journal's review, and later work by other researchers has found the same kind of structure.

Their best-known example concerns "king," "man," "woman," and "queen." Start at the position of "king." Move in the direction that leads from "man" to "woman." The word nearest to where you arrive is "queen" (Mikolov et al. 2013, sec. 1.1). The same holds for other pairs. The step from "big" to "biggest" matches the step from "small" to "smallest," and the step from a country to its capital is similar across many countries (Lee and Trott 2023, "Word vectors").

So a direction in the space can stand for a relationship, such as male to female or country to capital.

Learned in Training

Nobody assigned these positions. No one decided that "cat" belongs near "kitten" or measured the distance from "king" to "queen."

Embeddings are parameters, which are the adjustable numbers in a model, and training sets their values. They begin as random numbers. As training adjusts the model to predict text better, tokens used in similar ways are pushed toward similar positions, because treating them alike improves the predictions. The structure appears as a side effect of learning to predict.

What Nearness Reflects

Embeddings record how words are used in the training text. They carry whatever that text carries.

Lee and Trott point out that word vectors end up "reflecting many of the biases that are present in human language." Their example is that in some word vector models, starting at "doctor" and taking the step from "man" to "woman" arrives at "nurse" (Lee and Trott 2023, "Word vectors"). No one programmed that association. It was in the patterns of the text.

A second limit concerns words with several meanings. A single position for "bank" has to serve both the riverside and the financial institution. A language model handles this by adjusting each token's representation according to the surrounding words as the text is processed (Lee and Trott 2023, "Word meaning depends on context").

Conclusion

A model represents each token as an embedding, a position in a space with many dimensions. Tokens used in similar ways end up near each other, and directions in the space can stand for relationships. The positions are learned from text during training, so they reflect how words are used in that text, including its biases.

Key Terms

  • Embedding: A list of numbers that represents a token and works as the token's position in a space of meaning.
  • Vector: A list of numbers that gives a position.
  • Dimension: One of the numbers in a vector, and so one of the directions in the space.
  • Semantic similarity: Closeness in meaning, which in a model shows up as nearness between embeddings.

References

  • Lee, Timothy B., and Sean Trott. 2023. "Large Language Models, Explained with a Minimum of Math and Jargon." Understanding AI, July 27, 2023.
  • Mikolov, Tomas, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. "Efficient Estimation of Word Representations in Vector Space." arXiv:1301.3781.
    • Free: arXiv 1301.3781
    • Free companion: Lee and Trott 2023, section on word vectors

Report an issue with this item

Reading 4 min

Next-Token Prediction as a Ranking of Possible Continuations

Introduction

Ask a chatbot a question and you get one answer, stated plainly. That presentation hides what the model computed. What the model computed was a judgment about every possible next piece of text, and the reply was built from a long series of such judgments.

This reading explains what a model outputs, how it was trained to do so, and how single predictions add up to a full reply.

What the Model Produces

A language model is a model trained to predict what comes next in a piece of text. It works on tokens, the pieces of text, often words or parts of words, that a model handles as units.

Give a language model some text and it returns a likelihood for every token in its vocabulary, which is the fixed list of tokens it can use. This full set of likelihoods is a probability distribution: one likelihood for each possible outcome, which together account for all the possibilities. Wolfram describes the result as "a ranked list of words that might follow," each with a probability (Wolfram 2023, "It's Just Adding One Word at a Time").

Suppose the text so far is "I like my coffee with cream and," an example Lee and Trott use (Lee and Trott 2023, "How language models are trained"). An illustrative ranking might put "sugar" far ahead, with "milk," "a," and "honey" well behind, and tens of thousands of other tokens sharing a tiny remainder. The model's output is the whole ranking. Choosing one token from it is a separate step.

Next-token prediction is the name for this task: given the text so far, assign a likelihood to each token that could come next.

How a Model Learns to Predict

Training uses ordinary text. Take a passage, hide the next token at some point, and have the model produce its ranking. Then reveal the token that was really there. If the model gave that token a high likelihood, its error is small. If it gave it a low likelihood, its error is large. The model's parameters, its adjustable numbers, are nudged so that the real token would be ranked a little higher the next time.

This is repeated at every position in a very large body of text. No person labels anything, since the text supplies the right answer at each position. Lee and Trott describe the result as weights that are "gradually adjusted to make better and better predictions" (Lee and Trott 2023, "How language models are trained").

What Good Prediction Requires

Predicting the next token sounds like a narrow skill. Doing it well across all kinds of text demands a lot.

Text so farWhat a good prediction draws on
"The children was" or "The children were"Grammar
"Water boils at 100 degrees"Facts about the world
"Dear Sir or Madam, I am writing to"The conventions of a formal letter
"Roses are red, violets are"A familiar rhyme

The model isn't given grammar rules or a list of facts. Wherever a regularity in text helps predict what comes next, training pushes the parameters toward capturing it. Grammar, common knowledge, and style all help, so the model picks up a working version of each.

This account says what the training rewards. It doesn't settle what the model has inside as a result. Whether the result deserves to be called knowledge or understanding is argued over by researchers, and the mechanism described here is common ground for both sides.

A second point follows from the training. The model is rewarded for likelihood and not for truth. A false statement that appears often in text can be ranked above a true one that appears rarely.

From One Token to a Whole Reply

A single prediction gives one token. A reply is made by repeating the step.

  1. The model takes the text so far and produces its ranking.
  2. One token is chosen from the ranking.
  3. That token is added to the end of the text.
  4. The lengthened text goes back in as the new input, and the cycle repeats.

This is autoregressive generation: producing text one token at a time, with each new token added to the text so far before the next is predicted. Wolfram's summary is that the system is asking, over and over, what the next word should be given the text so far (Wolfram 2023, "It's Just Adding One Word at a Time"). Generation ends when the model's chosen token is a special one that means the reply is finished.

Two things follow. The model has no finished reply in hand while it writes, only the text so far. And every token it has already produced becomes part of what it predicts from, so an early choice shapes everything after it.

Conclusion

A language model's output is a ranking of every possible next token by likelihood. Training rewards rankings that favor the token that really came next, and predicting well requires picking up grammar, facts, and style from text. A full reply is the same prediction repeated, each chosen token added to the text before the next prediction is made.

Key Terms

  • Language model: A model trained to predict what comes next in a piece of text.
  • Probability distribution: A full set of likelihoods, one for each possible outcome, which together account for all the possibilities.
  • Next-token prediction: The task of assigning a likelihood to each token that could come next, given the text so far.
  • Autoregressive generation: Producing text one token at a time, with each new token added to the text so far before the next is predicted.

References

  • Lee, Timothy B., and Sean Trott. 2023. "Large Language Models, Explained with a Minimum of Math and Jargon." Understanding AI, July 27, 2023.
  • Wolfram, Stephen. 2023. "What Is ChatGPT Doing … and Why Does It Work?" Stephen Wolfram Writings, February 14, 2023.

Report an issue with this item

Reading 4 min

Sampling, Temperature, and Why Identical Prompts Give Different Replies

Introduction

Ask a chatbot the same question twice and you'll often get two different replies. A calculator never behaves this way, and people reasonably wonder whether the variation is a fault.

This reading explains where the variation comes from, the setting that controls it, and what it means for how much weight one reply can bear.

The Trouble with Always Taking the Top Choice

A language model writes one token at a time, where a token is a piece of text such as a word or part of a word. At each point the model ranks every possible next token by likelihood. Something then has to pick one.

The simplest rule is to take the top-ranked token every time. This makes the output predictable, and it makes for poor writing. Wolfram reports that always picking the highest-ranked word typically gives a "flat" essay, one that is dull and inclined to repeat itself (Wolfram 2023, "It's Just Adding One Word at a Time"). Good writing often uses a word that was likely but wasn't the single likeliest.

A rule that gives the same output every time for the same input has a name. Determinism is the property of producing the same output whenever the input is the same. A calculator is deterministic. A model that always takes the top token is deterministic too, or very nearly.

Sampling

Chatbots usually work differently. They use sampling: choosing the next token by chance, with likelier tokens more likely to be picked.

A lottery wheel is a fair picture. Each candidate token gets a slice of the wheel in proportion to its likelihood. A token the model rates at about one-half gets half the wheel. One rated at one in a hundred gets a thin sliver. The wheel is spun, and the token it lands on is chosen. Likely tokens win most of the time, and unlikely ones win now and then.

This is where the variation comes from. Randomness is the element of chance in a process, which lets the same starting point lead to different results. The model's ranking for a given text is the same each time. The spin isn't.

Temperature

How adventurous the choice is can be adjusted. The setting is called temperature, and it controls how often sampling picks less likely tokens. Wolfram describes it as a setting that "determines how often lower-ranked words will be used" (Wolfram 2023, "It's Just Adding One Word at a Time").

TemperatureEffect on the wheelTypical result
LowestThe top token takes the whole wheelThe same reply each time, plain and repetitive
LowLikely tokens get even bigger slicesSteady, predictable wording
MediumSlices match the model's likelihoodsVaried and mostly sensible
HighUnlikely tokens get bigger slicesSurprising, and increasingly incoherent

Wolfram notes that the choice of setting rests on trial and error. In his words there is "no 'theory' being used here," only what has been found to work (Wolfram 2023, "It's Just Adding One Word at a Time").

The company that builds a chatbot product sets its temperature. In most consumer products you can't see the setting or change it. Developers who build on a model through a company's technical services can usually set it themselves.

One Early Choice Changes What Follows

Each chosen token is added to the text, and the model predicts the next token from the text as it now stands. So a different draw at one point changes the input for every later point.

Suppose two runs of the same prompt differ at the fourth word. One reply begins "There are three main reasons" and the other begins "There are several factors." From then on the model is continuing two different texts. The first will go on to list three reasons. The second may list five factors, in another order, with other examples. A small difference near the start grows into two different replies.

The difference can be one of wording only, and it can also be one of substance. Where the model's ranking strongly favors one continuation, such as a well-known date, replies tend to agree. Where several continuations are about equally likely, replies can disagree on the content itself.

One Reply Is One Draw

A reply you receive is one draw from the many replies the model could have given to the same prompt. Asked again, it might say the same thing in other words, or it might say something else.

This has a practical use. If you ask the same question several times and the replies agree on the substance, the model's ranking is firm on that point. That is a statement about the model, and a firm ranking can still be firmly wrong. If the replies disagree, the model's ranking was spread across several options, and no single reply should be taken as its settled answer.

Conclusion

Chatbots choose each token by sampling, which favors likely tokens and leaves room for chance. Temperature sets how much room. Since each choice feeds into the next prediction, small differences grow, and the reply you see is one of many the model could have produced.

Key Terms

  • Determinism: The property of producing the same output whenever the input is the same.
  • Sampling: Choosing the next token by chance, with likelier tokens more likely to be picked.
  • Randomness: The element of chance in a process, which lets the same starting point lead to different results.
  • Temperature: The setting that controls how often sampling picks less likely tokens.

References

  • Wolfram, Stephen. 2023. "What Is ChatGPT Doing … and Why Does It Work?" Stephen Wolfram Writings, February 14, 2023.

Report an issue with this item

Guided Reading 7 min

Worked Example: One Sentence from Prompt to Finished Reply

Introduction

"The model predicts the next token" is easy to say and hard to picture. Following a single short prompt through every stage shows what the phrase covers and where a wrong answer can enter.

This example traces the prompt "The capital of Australia is" through an invented small language model. The model, its token splits, and its likelihoods are made up for illustration and weren't measured from any real system.

The Problem

The prompt is five words long:

The capital of Australia is

The task is to follow how the model turns this prompt into a finished reply, and then to run it a second time and see whether the reply comes out the same. The correct completion is Canberra. Sydney is Australia's largest and best-known city, and it's a common wrong answer among people.

Working It Through

Step 1: Split the prompt into tokens

The model can't take in letters or words directly. The prompt is first split into tokens, the pieces of text a model handles as units. In this invented model every word of the prompt is common enough to have a token of its own, so the split is simple.

PositionToken
1The
2capital
3of
4Australia
5is

Each token is then replaced by its number on the model's fixed list of tokens. From here on the model is working with five numbers.

Step 2: Turn each token into a position

Each token number is swapped for that token's embedding, a list of numbers that works as a position in a space of meaning. Tokens that are used in similar ways get nearby positions (Lee and Trott 2023, "Word vectors"). In this invented model the embedding for "Australia" sits near those for other countries, and also near tokens such as "Sydney," "Canberra," and "Melbourne," which turn up in the same kinds of sentences.

The model's layers then combine the five positions, so that "capital" is read in light of "Australia" and the sentence as a whole points toward a place name. How the layers do that is a subject of its own. What matters here is the result, which is one set of numbers summing up the text so far.

Step 3: Rank the candidates for the next token

From that summary the model produces a likelihood for every token on its list. Most get almost nothing. Here are the top five.

RankCandidate tokenRough likelihood
1CanberraAbout half
2SydneyAbout a quarter
3aAbout one in ten
4theAbout one in twenty
5locatedAbout one in twenty

All the other tokens on the list share the small remainder.

"Sydney" is on this list for a reason. Suppose the model's training text mentions Sydney alongside Australia far more often than it mentions Canberra, and that some of that text states the mistake outright. The model's ranking reflects the text it learned from, so a frequent error earns a real share.

Step 4: Sample one token, add it, and repeat

The model now samples, which means it picks one candidate by chance, with likelier candidates more likely to be picked. On this run the draw lands on "Canberra." That token is added to the text.

The text so far is now "The capital of Australia is Canberra," and the whole of it goes back in as input. The model produces a fresh ranking, a token is drawn, and the cycle repeats. Wolfram describes this loop as asking over and over what should come next given the text so far (Wolfram 2023, "It's Just Adding One Word at a Time").

Text so far ends withLeading candidatesToken drawn
… is Canberraa comma, a period, "which",
… is Canberra,"a," "which," "not"a
… is Canberra, a"city," "planned," "small"planned

Two more cycles add "city" and a period. On the cycle after that, the model's draw is a special token that means the reply is finished, and generation stops. The first reply reads: "The capital of Australia is Canberra, a planned city."

Step 5: Run it again with a different draw

Start over from Step 3 with the same prompt. The model's ranking is identical, since the input is identical. The draw is a new one. This time it lands on "Sydney," which had about one chance in four.

"Sydney" is added to the text, and the model continues from there. It doesn't go back and reconsider. Its input now ends "is Sydney," and its job is to predict what follows that text.

Text so far ends withLeading candidatesToken drawn
… is Sydneya comma, a period, "which",
… is Sydney,"the," "a," "which"the
… is Sydney, the"largest," "country," "most"largest

Two more cycles add "city" and a period, and then the stop token is drawn. The second reply reads: "The capital of Australia is Sydney, the largest city."

The second reply is false, and it's as fluent as the first. Nothing in its wording marks it as the less likely draw.

Key Considerations

The likelihoods in this example are illustrative. A real, large model would probably rank "Canberra" far higher on a question this well covered, and the error would be rare. The pattern becomes common on questions where the training text is thin, mixed, or often wrong.

A common mistake is to think the model looks the answer up. It has no table of capitals to consult. It ranks tokens by how well they continue the text, and "Sydney" can outrank "Canberra" wherever the text the model learned from pairs Sydney with Australia often enough. A correct answer and an incorrect one come out of the same process.

A second point concerns the moment of the error. The second reply went wrong at a single token. Everything after it was a sensible continuation of a false start, and the phrase "the largest city" is even true of Sydney. The model built on its own earlier output as it would on any other text.

Summary

The same prompt, the same model, and the same ranking produced two replies that split at the first generated token.

First runSecond run
PromptThe capital of Australia isThe capital of Australia is
Token 1CanberraSydney (the runs split here)
Token 2,,
Token 3athe
Token 4plannedlargest
Tokens 5 and 6city .city .
Finished replyThe capital of Australia is Canberra, a planned city.The capital of Australia is Sydney, the largest city.

Check: both runs used the ranking from Step 3 unchanged. "Canberra" had about half the likelihood and "Sydney" about a quarter, so across many runs roughly twice as many replies would begin with "Canberra" as with "Sydney." Getting one of each in two runs is an ordinary outcome.

References

  • Lee, Timothy B., and Sean Trott. 2023. "Large Language Models, Explained with a Minimum of Math and Jargon." Understanding AI, July 27, 2023.
  • Wolfram, Stephen. 2023. "What Is ChatGPT Doing … and Why Does It Work?" Stephen Wolfram Writings, February 14, 2023.

Report an issue with this item

Guided Conversation 12 min

Predict the Next Word Yourself

In this conversation you'll do a language model's job by hand: you'll be given unfinished sentences and asked what's likely to come next and what you drew on to guess. You'll leave with your own statement of what it means that a reply is one draw from many.

You'll have this conversation with an AI assistant, using your own account. Choose a button to open a new chat with the prompt already filled in, then press send to start. If the chat opens empty, copy the prompt and paste it in.

Run this conversation in whichever assistant you already use:

Claude desktop app

To use another LLM, simply copy and paste the prompt into its chat window.

Show the full prompt (it lists misreadings to watch for, so skip it if you would rather come to the conversation fresh)
Guided Conversation: Predict the Next Word Yourself (about 12 minutes)

Note to the learner: press send to start. Everything below is facilitator guidance for the AI. It lists misconceptions to watch for, so skip it if you'd rather come to the conversation fresh.

Please facilitate a coached problem session with me. I'm an adult with no technical background who has used AI chatbots for everyday tasks, and I'm studying tokens and how a language model generates text by predicting what comes next. Follow this guidance for the whole conversation.

GOAL
I can explain how a language model generates a response by predicting what comes next, having done the prediction by hand.

HOW TO RUN THE CONVERSATION
- Ask one question at a time, then wait for my reply. Keep each of your turns under about 120 words.
- Don't lecture. Explain a point only when I need it to continue, then return to the sentence we're working on.
- Be curious and collegial. Use plain words and define any technical term briefly on first use. No formulas and no code. Welcome disagreement when I give a reason.
- This is a coached problem. Ask for my guesses and my reasoning before you give any hint. Give one hint at a time. Don't tell me what a model would do until I've made my own attempt.
- Plain conversation only: don't search the web or create files or documents.
- Don't ask for anything confidential or personal.
- Aim for about 12 minutes. Spend most of the time on topics 1 and 2. If my replies are brief, offer one concrete prompt, such as "Give me two words that could come next and say which is likelier," and move on. If I seem uncertain, shorten the conversation to 5-7 minutes. Always reach the final topic.
- Start now. Open with one or two warm sentences: this is a conversation, not a quiz; my reasons for a guess matter more than the guess; I can ask you to clarify anything. Then give me the first unfinished sentence.

TOPICS, IN ORDER
1. Three unfinished sentences, one at a time. Use these or close equivalents: "She poured herself a cup of"; "The children were late for school because the bus"; "Dear hiring manager, I am writing to". For each, ask me for two or three likely next words, roughly ranked, and what knowledge each guess used (word meaning, grammar, facts about the world, the conventions of a kind of writing). After the third, ask what my answers have in common: I produced a ranked set of options, not one answer.
2. A likely continuation that is false. Give me: "The Great Wall of China is the only human-made structure visible from". Ask what word most often follows in what people write, and whether that makes the sentence true. (The common claim that it is visible from space with the naked eye is false.) Ask what a system trained to predict likely text would tend to do here and why.
3. Step by step, yet planned-looking. Ask how a system that only picks the next piece of text can produce a paragraph that looks planned. Draw out that each new piece is added to the text and the whole text is used for the next prediction, so earlier choices constrain later ones.
4. Closing. Explain briefly that a model usually picks among likely options with some chance involved, so one reply is one draw from many possible replies. Ask me to state what that means for how much I should trust a single reply. Tell me I can take my statement into a short optional activity where I ask one question several times and compare the replies.

KEY POINTS TO KEEP ACCURATE
- Method for each sentence: list candidates, rank them by likelihood, name the knowledge behind each. There is no single right word.
- A language model works on tokens, pieces of text that are often words or parts of words.
- Given the text so far, the model assigns a likelihood to every token it knows. Its output is a ranking, not one answer.
- A reply is made by repeating the step: choose a token, add it to the text, predict again.
- The choice usually involves sampling: chance weighted by likelihood. A setting called temperature controls how adventurous it is. This variation is designed in.
- Likely is not the same as true. Training rewards predicting what text says, and text contains errors.
- If I ask what you are doing as you write, explain the general mechanism and say plainly that you can't observe your own token choices or settings, so your statements about yourself are not evidence.

MISCONCEPTIONS TO CORRECT GENTLY
When one appears, name the accurate version briefly, then return to the sentence.
- "It retrieves stored answers": it generates text token by token and has no table of answers to consult.
- "It writes the whole reply and then shows it": each token is chosen after the earlier ones are fixed.
- "Variation means it's broken": variation comes from sampling and is intended.
- "If it sounds confident, it's the likeliest answer": fluent wording doesn't show how likely or how true a statement is.

LIMITS
- Don't introduce attention, transformers, or how models are trained at scale.
- Don't state your own temperature, settings, or internals as fact.
- Don't settle whether prediction amounts to understanding. If I raise it, say researchers disagree and return to the mechanism.
- Don't recommend or compare products.

TO FINISH
After my closing statement, close in one short turn:
- Affirm one specific thing I worked out, in my own words where possible.
- Suggest one or two next steps that fit how the conversation went. Possible steps: ask one question several times in fresh conversations and compare; write two unfinished sentences of my own where the likely continuation is false; reread a definition of sampling and check my statement against it.
- Restate my statement on its own line, labeled "What one draw from many means for me", so I can copy it.

Report an issue with this item

Hands-on Activity 15 minOptional

Ask the Same Question Five Times

Overview

A chatbot's reply is one of many it could have given. In this activity you'll see that for yourself by asking two questions five times each and comparing what comes back.

The activity is optional. Your notes are for you, and nobody collects them.

What You'll Need

  • An AI assistant you already use
  • Somewhere to paste or write ten short replies

Your Task

Put one factual question and one open question to an AI assistant five times each, in fresh conversations, and compare the replies.

Steps

  1. Choose two questions. The first should be factual, with one answer you can check, such as the year a well-known event took place. The second should be open, such as "Suggest a name for a neighborhood café." Keep both short, and ask for a short reply. Don't include anything personal or confidential.
  2. Ask each question five times, starting a new conversation each time. Use exactly the same wording every time. A new conversation matters, because inside one conversation the assistant can see its earlier replies and will take them into account. If your assistant has a memory feature or a temporary-chat mode, use the mode that doesn't carry anything over.
  3. Mark what stayed the same and what changed. Lay the five replies to each question side by side. Look at wording and at substance separately. Two replies can say the same thing in different words, or say different things.
  4. Write three sentences on which kind of question varied more and why. Use what you know about how a model picks each token.

What to Expect

The factual question will probably get the same answer five times, in wording that varies a little. The model's ranking strongly favors one continuation, so chance has little room to change the substance.

The open question will probably get several different answers, perhaps with one or two repeats. Many continuations are about equally likely, so chance decides among them.

If your factual answers disagree with each other, you've found a point where the model's ranking is spread out. Check the fact somewhere reliable. Five matching answers show that the model is consistent on the point, and they don't show that it's right.

Self-Check

When you're done, check that:

  • You have ten replies, five for each question
  • You started a new conversation for every one
  • You noted differences in wording and differences in substance separately
  • Your explanation uses the idea of sampling or of likelihood

Nothing is uploaded. Write in your own notebook or document and keep it.

Report an issue with this item

Knowledge Check 10 min

Tokens, meaning, and next-token prediction

This ungraded knowledge check assesses your understanding of how a language model takes in text and produces a reply. You'll be asked about tokens and tokenization, embeddings, next-token prediction, and sampling and temperature.

Note: Use this to test yourself, review the feedback on any questions you miss, and retry until you feel confident before moving forward.

5 questions · ungraded · retry as often as you like

Report an issue with this item

Graded Quiz 30 min

Machine Learning and Language Models

This graded quiz assesses your understanding of how machine learning systems and language models work. You'll be asked about rules and learned patterns, kinds of learning, generalization and overfitting, neural networks and their parameters, loss and gradient descent, tokens and embeddings, and next-token prediction.

Note: Aim for a score of 80 percent or higher. If you score lower, use the feedback to review the topics you missed, then retake the quiz.

10 questions · target score 80% · 3 forms, rotated on each attempt

Report an issue with this item