KnowledgeInSight
AI Literacy
0% of Course 1 complete

Module 2 · Lesson 1

Building a Model: The transformer and attention

You'll look at the design behind today's language models and the one idea at its center: letting every word in a passage weigh every other word. You'll be able to explain why this design made very large models practical to build.

What you will be able to do

  • Explain what attention does in a transformer and why it made large language models practical.

0% of this lesson · 9 items · 1h 3m total · 48m without the optional activity

Contents of this lesson9 items
  1. ReadingThe 2017 Paper That Today's Chatbots Are Built On3 min
  2. ReadingThe Context Problem: Why a Word's Meaning Depends on the Words Around It4 min
  3. ReadingAttention as Weighing Which Words Matter to Each Other4 min
  4. ReadingLayers and Parallel Processing: Why Transformers Could Be Scaled Up4 min
  5. ReadingThe Context Window: How Much Text a Model Can Take In at Once4 min
  6. Guided ReadingGuided Walkthrough: Tracing Attention Through One Ambiguous Sentence7 min
  7. Guided ConversationFind the Words That Carry the Meaning12 min
  8. Hands-on Activity · optionalTest How Added Context Changes an Answer15 min
  9. Knowledge CheckThe transformer and attention10 min

Reading 3 min

The 2017 Paper That Today's Chatbots Are Built On

The letter "T" turns up in the names of many language models. In OpenAI's GPT series it's the last letter, and the name stands for "generative pre-trained transformer." Google's research models BERT and T5 have it too. In each case the T stands for the same word, transformer.

A transformer is a design for a neural network, which is a model made of many simple units arranged in layers. The design comes from a paper posted in June 2017 under the title "Attention Is All You Need." It lists eight authors. Six give Google's research groups as their affiliation, and the paper notes that the other two did the work while at Google (Vaswani et al. 2017).

The paper wasn't about chatbots. Its subject was machine translation, and its tests were English-to-German and English-to-French. Translation systems of the time read a sentence one word after another, and the best of them added a mechanism called attention, which let the system look back at particular words in the sentence it was translating. The authors proposed keeping that mechanism and removing the rest. They describe their design as "based solely on attention mechanisms" (Vaswani et al. 2017, abstract).

Two results were reported. The translations scored higher than those of earlier systems on standard tests. And the new model was faster to train: the paper says its large model reached its best score after training for "3.5 days on eight GPUs," where GPUs are chips that can do many calculations at once (Vaswani et al. 2017, abstract). The authors call that a small fraction of what the best earlier models had cost.

Within a few years the same design was being used for much more than translation. The language models behind chatbots from OpenAI, Google, Anthropic, Meta, and other developers are transformers, scaled up enormously from the 2017 original.

That history supports a claim you'll meet in accounts of how current AI came about. On this view, no new theory of language opened the way to today's systems. One design change did, by letting each word in a passage draw on every other word directly, and by letting a computer process all the words at the same time.

The claim deserves a careful look, because it leaves things out. The design alone produced a better translation system. Chatbots also needed far more text, far more chips, and the money to pay for both, and they needed further training to make a model behave as an assistant. A fair version of the claim is narrower: the transformer made it practical to train models on amounts of text that earlier designs couldn't get through in a reasonable time.

Two questions follow. One is what attention does, and why a model predicting the next word needs it. The other is why this design could be scaled up when earlier ones couldn't.

References

  • Vaswani, Ashish, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. "Attention Is All You Need." arXiv:1706.03762.
    • Free: arXiv 1706.03762
    • Free companion: Lee and Trott 2023, sections on transformers and attention

Report an issue with this item

Reading 4 min

The Context Problem: Why a Word's Meaning Depends on the Words Around It

Introduction

Ask a chatbot about "the bank" in a message about fishing and it will talk about a riverside. Ask in a message about loans and it will talk about money. The word is the same both times, and the chatbot's reading of it changes with the words nearby.

This reading explains why a model can't predict the next word without settling what the earlier words mean, why earlier designs struggled to do that over long passages, and what a workable solution has to provide.

Words That Change with Their Surroundings

A language model writes by predicting what comes next, one piece of text at a time. To predict well, it has to work out what each earlier word means in this particular passage.

That's harder than it sounds, because many words have more than one meaning. Ambiguity is the property of having more than one possible meaning. Timothy B. Lee, a journalist, and Sean Trott, a cognitive scientist, give the standard example in their explainer on language models: "bank" can refer to a financial institution or to the land next to a river (Lee and Trott 2023, "Word meaning depends on context").

Pronouns are ambiguous in a different way. "It," "they," and "his" have almost no meaning of their own. They point at something mentioned elsewhere. Lee and Trott's example is "the customer asked the mechanic to fix his car," where "his" could be the customer or the mechanic.

What settles each case is context: the other words in the passage that bear on what a given word means. Lee and Trott observe that people resolve such cases from context, and that there are "no simple or deterministic rules" for doing it (Lee and Trott 2023, "Word meaning depends on context"). Nobody could write the rules out for a program to follow.

When the Deciding Word Is Far Away

Sometimes the word that settles a meaning sits right next to the ambiguous one, as in "river bank." Often it doesn't.

Take this invented passage: "The contract Dana drafted last spring, after two rounds of comments from the client's lawyers and a long delay over the insurance terms, was finally signed. She celebrated." The word "She" points back to "Dana," more than twenty words earlier. To predict what comes after "She," a model has to connect those two words across everything in between.

A link of this kind is a long-range dependency: a connection between words that are far apart in a passage, where one is needed to interpret the other. Long documents are full of them.

Reading in Order, and Losing Track

Before 2017, the leading language systems handled text in strict order. A sequence model is a model that takes in a series of items, such as the words of a sentence, and produces an output that depends on their order. The dominant kind read one word at a time and kept a running summary of everything read so far. Each new word updated the summary, and the summary was all the system had to go on.

The 2017 paper that introduced a different design describes these earlier systems as inherently sequential: each step had to wait for the one before it (Vaswani et al. 2017, sec. 1).

The running summary was the weak point. It had a fixed size, so every new word had to be squeezed in alongside what was already there. Details from early in a passage faded as later words arrived. By the time such a system reached "She," its trace of "Dana" might be faint, mixed with the lawyers, the client, and the insurance terms.

Reading in order with a running summaryWhat the task needs
How an early word reaches a later oneOnly through the summary, updated at every stepDirectly
Effect of distanceThe further back a word is, the fainter its traceNone: a far word counts as much as a near one if it's relevant
Which earlier words get usedWhatever survived in the summaryThe ones that matter for this word

What a Solution Needs

The right-hand column describes the requirement. A model needs a way for any word in a passage to draw on any other word, at any distance, and to draw most heavily on the words that are relevant to it.

That is two requirements. The first is a direct route between every pair of words. The second is selectivity. "She" should take a lot from "Dana" and very little from "insurance." And since no person can write rules for which words matter to which, the model has to learn that selectivity from text.

Conclusion

Predicting the next word requires knowing what the earlier words mean here, and that depends on context that may be far away. Systems that read strictly in order carried context in a running summary and lost track over long passages. The requirement they failed to meet was direct, selective access from each word to every other.

Key Terms

  • Ambiguity: The property of having more than one possible meaning.
  • Context: The other words in a passage that bear on what a given word means.
  • Long-range dependency: A connection between words that are far apart in a passage, where one is needed to interpret the other.
  • Sequence model: A model that takes in a series of items, such as the words of a sentence, and produces an output that depends on their order.

References

  • Lee, Timothy B., and Sean Trott. 2023. "Large Language Models, Explained with a Minimum of Math and Jargon." Understanding AI, July 27, 2023.
  • Vaswani, Ashish, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. "Attention Is All You Need." arXiv:1706.03762.
    • Free: arXiv 1706.03762
    • Free companion: Lee and Trott 2023, sections on transformers and attention

Report an issue with this item

Reading 4 min

Attention as Weighing Which Words Matter to Each Other

Introduction

News coverage of AI often mentions "attention" as the idea behind modern language models, usually without saying what it is. The everyday meaning of the word suggests a mind concentrating, which is a poor guide to the mechanism.

This reading explains what attention does to each piece of text a model handles, how the strength of each link is set, and why a model runs many attention patterns at once.

What Attention Does

A language model works on tokens, which are pieces of text that are often words or parts of words. Inside the model, each token is represented by a list of numbers. At the start, that list reflects the token alone, with no account of its neighbors. The list for "it" is the same in every sentence.

Attention is the mechanism that lets each token collect information from other tokens in the passage, taking more from the ones that are relevant to it. After an attention step, the token's list of numbers has been updated with what it collected. The list for "it" now carries information about the thing "it" refers to.

When the tokens doing the collecting and the tokens being collected from belong to the same passage, the mechanism is called self-attention. The paper that introduced the transformer design defines it as attention "relating different positions of a single sequence" (Vaswani et al. 2017, sec. 2). In writing about language models, "attention" nearly always means self-attention.

Asking and Offering

Lee and Trott, authors of a plain-language explainer on language models, describe attention as "a matchmaking service for words" (Lee and Trott 2023, "Can I have your attention please"). In their account each token produces two things. One is a description of what it's looking for. The other is a description of what it has to offer. The model compares every token's request against every token's offer.

Their example uses the partial sentence "John wants his bank to cash the." The token "his" is looking for something like a noun describing a male person. The token "John" offers exactly that. The two match, and information about "John" moves into the representation of "his" (Lee and Trott 2023, "Can I have your attention please").

How well a request matches an offer is expressed as an attention weight: a number that sets how much one token takes from another. A high weight means a strong link and a lot of information transferred. A weight near zero means the token is nearly ignored.

One Token's Links

Here is an invented sentence: "The dog chased the ball across the yard until it rolled under the fence." The table shows the kind of weights a trained model might give the links from "it" to the other words. The strengths are illustrative and weren't measured from a real model.

WordLink from "it"Why
ballStrongA ball is a thing that rolls
rolledStrongIt says what "it" did, which narrows what "it" can be
dogModerateA possible referent, though dogs seldom roll under fences
yardWeakA place, and an unlikely thing to roll
the, across, untilVery weakThese carry little about what "it" is

After this step, the representation of "it" is mostly information about a ball. Every other token gets the same treatment at the same time.

Many Patterns at Once

One set of weights can track one kind of relationship. Language has many kinds: pronouns and what they refer to, verbs and their subjects, adjectives and the nouns they describe.

So a transformer runs several attention patterns side by side. Each one is an attention head: one set of attention weights computed within a layer, alongside other sets that can track different relationships. Lee and Trott suggest that one head might match pronouns with nouns while another works on words with two meanings, such as "bank" (Lee and Trott 2023, "Can I have your attention please"). The original transformer used eight heads in each layer (Vaswani et al. 2017, sec. 3.2.2). Lee and Trott report that a large model from 2020 had 96 layers with 96 heads each.

Learned in Training

Nobody tells a head what to track. The numbers that produce each token's request and offer are parameters, the adjustable numbers in a model whose values are set by training. They begin as random values. Training adjusts them because useful links between tokens lead to better predictions of the next token.

When researchers inspect a trained model, they find heads that have taken on recognizable jobs. The transformer's authors reported that individual heads "clearly learn to perform different tasks," and that many seemed to follow the grammar and meaning of sentences (Vaswani et al. 2017, sec. 4). Other heads do things nobody has managed to describe.

Conclusion

Attention updates each token's representation with information from other tokens, weighted by relevance. The weights come from comparing what each token seeks with what each token offers, and many heads do this side by side. All of the numbers behind those comparisons are set by training.

Key Terms

  • Attention: The mechanism that lets each token collect information from other tokens in the passage, taking more from the ones that are relevant to it.
  • Self-attention: Attention in which the tokens doing the collecting and the tokens being collected from belong to the same passage.
  • Attention weight: A number that sets how much one token takes from another.
  • Attention head: One set of attention weights computed within a layer, alongside other sets that can track different relationships.

References

  • Lee, Timothy B., and Sean Trott. 2023. "Large Language Models, Explained with a Minimum of Math and Jargon." Understanding AI, July 27, 2023.
  • Vaswani, Ashish, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. "Attention Is All You Need." arXiv:1706.03762.
    • Free: arXiv 1706.03762
    • Free companion: Lee and Trott 2023, sections on transformers and attention

Report an issue with this item

Reading 4 min

Layers and Parallel Processing: Why Transformers Could Be Scaled Up

Introduction

Language models grew from millions of adjustable numbers to hundreds of billions in a few years. Earlier designs for handling language were never scaled up that far, and the reason has to do with how they read.

This reading explains how a transformer is put together, what its stacked layers do, and why processing a whole passage at once made very large models possible to train.

A Stack of Layers

Architecture is the overall design of a neural network: what parts it has and how they're connected. The transformer is an architecture in which every token, a piece of text such as a word or part of a word, draws on every other token in the passage through attention, a mechanism that weighs how relevant the tokens are to one another.

A transformer is built as a stack. A layer is a group of artificial neurons that all work on the same incoming numbers at the same stage. In a transformer, each layer does two things. First comes an attention step, in which tokens collect information from one another. Then comes a step in which each token's representation is processed on its own. Lee and Trott, authors of a plain-language explainer, put it this way: in the first step words "look around" for relevant context, and in the second each word "thinks about" what it gathered (Lee and Trott 2023, "Can I have your attention please").

The output of one layer is the input of the next. Each layer receives a list of numbers for every token and passes on an updated list. By the top of the stack, a token's representation reflects much of the passage around it.

The original transformer had six layers (Vaswani et al. 2017, sec. 3.1). Lee and Trott describe a model from 2020 with 96.

What Different Layers Do

Researchers who study trained models have found a rough division of labor. In Lee and Trott's summary, "Research suggests that the first few layers focus on understanding the syntax of the sentence" and on resolving ambiguities, while later layers work toward "a high-level understanding of the passage as a whole" (Lee and Trott 2023, "Transforming word vectors into word predictions").

Treat that as a tendency. Nobody assigned jobs to layers, and the boundaries are blurry. The division appears because training adjusts all the layers together to improve predictions.

The Whole Passage at Once

Earlier language systems read one word at a time, and each step needed the result of the step before. A computer running such a system had to do its work in a fixed order. The transformer's authors note that this ordering rules out doing the steps side by side, which matters more as passages get longer (Vaswani et al. 2017, sec. 1).

A transformer has no such chain. Within a layer, every token's update can be computed at the same moment, because each depends only on the previous layer's output. This is parallel processing: doing many calculations at the same time instead of one after another.

Earlier designsTransformer
How a passage is readOne word at a time, in orderAll tokens together, layer by layer
Can the work be split across chips?Poorly, since each step waits for the lastWell, since each token's update is separate
Effect of more chipsLittle gainTraining speeds up

One qualification matters. Parallel processing describes how a transformer reads a passage it already has. When a model writes a reply, it still produces one token at a time, because each new token depends on the ones before it.

Why That Made Large Models Practical

The chips used for this work, called GPUs, are built to do a very large number of simple calculations at once. A design whose work can be split up suits them. Lee and Trott say the approach lets language models "take full advantage of the massive parallel processing power of modern GPU chips" (Lee and Trott 2023, "Can I have your attention please").

The 2017 paper reported the difference directly. Its authors describe the transformer as "more parallelizable and requiring significantly less time to train" than the systems it replaced, and their large translation model finished training in three and a half days on eight chips (Vaswani et al. 2017, abstract).

"Practical" here means trainable in a reasonable time. Training involves running a model over its training text again and again while adjusting it. If a design can't spread that work across many chips, adding chips doesn't help, and a bigger model or a bigger collection of text means a longer wait with no way to shorten it. With a transformer, a developer who wants a larger model can add hardware. That is what developers went on to do.

Conclusion

A transformer is a stack of layers, each combining an attention step with a step that processes every token separately. Lower layers tend to settle local questions and higher layers broader ones. The whole passage is processed in parallel within each layer, which let training be spread across many chips and let models and their training text grow.

Key Terms

  • Architecture: The overall design of a neural network: what parts it has and how they're connected.
  • Transformer: An architecture in which every token draws on every other token in the passage through attention.
  • Layer: A group of artificial neurons that all work on the same incoming numbers at the same stage.
  • Parallel processing: Doing many calculations at the same time instead of one after another.

References

  • Lee, Timothy B., and Sean Trott. 2023. "Large Language Models, Explained with a Minimum of Math and Jargon." Understanding AI, July 27, 2023.
  • Vaswani, Ashish, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. "Attention Is All You Need." arXiv:1706.03762.
    • Free: arXiv 1706.03762
    • Free companion: Lee and Trott 2023, sections on transformers and attention

Report an issue with this item

Reading 4 min

The Context Window: How Much Text a Model Can Take In at Once

Introduction

In a very long conversation, a chatbot can seem to forget an instruction you gave near the start. Paste the instruction in again and it complies. The model hasn't changed in between, so something about what it can see must have.

This reading explains what a model has in front of it when it writes, why that amount is limited, and what the limit means for long conversations and pasted documents.

What the Window Holds

A language model writes by predicting the next token, a piece of text such as a word or part of a word. It predicts from a body of text in front of it. The prompt is the text a model is given as input and continues from.

In a chatbot, the prompt is more than your most recent message. It usually contains any instructions the product's maker placed at the start, the whole conversation so far, including the model's own earlier replies, and any documents you've supplied. Each time you send a message, all of it goes in again.

The context window is the maximum amount of text a model can take in at one time. The text that is in the window at a given moment is the model's working context. Attention, the mechanism by which tokens draw information from one another, operates only among tokens in the window. Lee and Trott's explainer on language models describes words looking around for other words with relevant context (Lee and Trott 2023, "Can I have your attention please"). The window marks the limit of where they can look.

A Finite Size, Counted in Tokens

Window sizes are stated in tokens. They've grown a great deal. A model described in 2020 took in 2,048 tokens at a time, roughly a few pages (Brown et al. 2020, sec. 2.1). Later models from many developers take in far more, and the figure differs from one model and one product to another.

Every window has an edge, though. The model's reply has to fit in the window too, since each token it writes becomes part of what it predicts from.

When a conversation grows past the limit, something has to give. Truncation is the dropping of text that no longer fits in the context window. Products handle this in different ways. Some drop the oldest messages. Some replace older messages with a short summary. Some stop and ask you to start a new conversation. In the first two cases the model carries on with no sign that anything is missing, because from its side nothing is: it only ever has the text in front of it.

Two Places Information Can Come From

When a model writes, it can draw on two sources.

SourceWhat it holdsWhen it was fixed
Parameters, the adjustable numbers set by trainingGeneral patterns and knowledge from the training textWhen training ended
Context windowThe instructions, conversation, and documents in front of the modelAssembled fresh for each reply

Anything in neither place is unavailable to the model. Yesterday's conversation is a common case. Unless a product stores notes from past conversations and inserts them into the window, a new conversation starts with nothing from the old one. Your organization's internal handbook is another: if it wasn't in the training text and you haven't supplied it, the model has no access to it. It may still produce a confident answer about the handbook, built from general patterns.

Why Long Conversations Drift

Two things can go wrong as a conversation lengthens.

  • Text falls out. An instruction from the start may have been truncated or reduced to a line in a summary. The model can't follow what it can't see.
  • Text gets diluted. Even when everything still fits, an early instruction is now a few tokens among many thousands. Attention has to pick it out from a lot of competing material, and it's commonly observed that models follow early instructions less reliably in very long contexts.

Restating an instruction puts it back in the window, near the end, where there's less to compete with it.

Why Pasted Documents Help

When you ask about a document without supplying it, the model answers from its parameters. That means a compressed impression of whatever its training text contained on the subject, which may be nothing.

When you paste the document, its words are in the window. Attention can link your question to the exact passage that answers it. The model is now working from text and not from recall, and replies about specifics tend to be more accurate. The limit still applies: a document longer than the window can't be taken in whole, and a product may cut it or pass along only parts.

Conclusion

A model writes from the text in its context window plus whatever training left in its parameters. The window is finite and counted in tokens, and text that doesn't fit is truncated, usually without notice. Keeping important material in the window, by restating it or pasting it, is the main way a user affects what the model has to work with.

Key Terms

  • Prompt: The text a model is given as input and continues from.
  • Context window: The maximum amount of text a model can take in at one time.
  • Working context: The text that is in the context window at a given moment.
  • Truncation: The dropping of text that no longer fits in the context window.

References

  • Brown, Tom B., Benjamin Mann, Nick Ryder, and 28 others. 2020. "Language Models Are Few-Shot Learners." arXiv:2005.14165.
  • Lee, Timothy B., and Sean Trott. 2023. "Large Language Models, Explained with a Minimum of Math and Jargon." Understanding AI, July 27, 2023.

Report an issue with this item

Guided Reading 7 min

Guided Walkthrough: Tracing Attention Through One Ambiguous Sentence

Introduction

"Attention weighs which words matter to each other" is a fair summary of the mechanism, and it's hard to picture. Following one pronoun through a few layers of a model shows what the summary means in practice.

This walkthrough traces the word "it" in a sentence where its meaning flips when a single later word changes. The link strengths are invented to illustrate the mechanism and weren't measured from any real model.

The Starting Point

Here are two sentences that differ in their last word.

The trophy didn't fit in the suitcase because it was too big.

The trophy didn't fit in the suitcase because it was too small.

The pair comes from the Winograd Schema Challenge, a set of such puzzles proposed as a test for machines. You read "it" as the trophy in the first sentence and as the suitcase in the second, without effort. Nothing in the grammar tells you which. You use what you know about fitting things into containers.

A language model works on tokens, pieces of text that are often words or parts of words. For simplicity, treat each word here as one token. Inside the model, each token is represented by a list of numbers. Before any processing, the list for "it" is the same in both sentences and says nothing about trophies or suitcases. The task is to follow how that representation changes as the sentence passes through a transformer, the layered design used in current language models.

Walking Through the Sentence

Step 1: Mark the ambiguous token and its candidates

The ambiguous token is "it." Two earlier tokens could be what it refers to: "trophy" and "suitcase." Both are nouns, both are things, and both come before "it," so grammar leaves the choice open.

Other tokens bear on the choice even though they aren't candidates. "Fit" and "in" set up a relationship in which one thing goes inside another. "Big" or "small" names a property that explains why the fitting failed. A reader uses all of these, and a model has to as well.

Step 2: Describe the links from "it" in the first sentence

In an attention step, each token collects information from the other tokens, taking more from those relevant to it. Lee and Trott, authors of a plain-language explainer on language models, describe this as a matchmaking service in which each word looks for other words with relevant context (Lee and Trott 2023, "Can I have your attention please"). How much one token takes from another is set by an attention weight.

Take the first sentence, ending in "big." Suppose the weights from "it" come out like this in one layer.

TokenLink from "it"
trophyStrong
bigStrong
fitModerate
suitcaseWeak
The, in, the, because, was, tooVery weak

After the step, the representation of "it" has been updated. It now holds a good deal of information about a trophy, some about fitting, and some about bigness. The token is still "it." Its list of numbers is now a blend of information from several tokens, in proportion to the weights.

Step 3: Change one word and watch the links shift

Now take the second sentence. Every token is the same until the last one, which is "small."

TokenLink from "it"
suitcaseStrong
smallStrong
fitModerate
trophyWeak
The, in, the, because, was, tooVery weak

The strong link has moved from "trophy" to "suitcase." The representation of "it" now holds mostly information about a suitcase.

Notice where the deciding word sits. "Small" comes after "it." A system that read strictly left to right and settled each word as it went would have to commit on "it" before reaching the word that decides the matter. In a transformer reading a passage it already has, every token's update is computed with the whole sentence available, so a later word can shape the reading of an earlier one.

Nobody wrote a rule saying that a thing too big to fit must be the thing going in. The weights come from parameters, the adjustable numbers set by training. The model was trained on a great deal of text about objects, containers, and sizes, and weights that link these words usefully made its predictions better.

Step 4: Follow the updated representation through two more layers

A transformer is a stack of layers, and each layer runs attention again on the output of the layer below (Vaswani et al. 2017, sec. 3.1). Go back to the first sentence and follow it upward.

In the next layer, the tokens attend to one another again, and this time "it" already carries trophy information. So when "big" collects from "it," what it picks up is, in effect, trophy. The representation of "big" comes to hold "big, said of the trophy." Likewise "fit," which already linked to "didn't," picks up more about what failed to fit into what.

In the layer above that, tokens near the end of the sentence can draw on all of this. The representation at the final position comes to reflect something like the whole situation: a trophy, a suitcase, a failure to fit, and the trophy's size as the cause. Lee and Trott report research suggesting that lower layers tend to settle questions like which noun a pronoun refers to, while higher layers build a broader reading of the passage (Lee and Trott 2023, "Transforming word vectors into word predictions").

Step 5: State what the model can now predict

The point of all this is prediction. Suppose each sentence continues with "So we bought a."

For the first sentence, a good continuation is "bigger suitcase" or "smaller trophy." For the second, "bigger suitcase" still works, but "smaller trophy" would be odd, since the trophy's size was never the stated problem. A model that hadn't resolved "it" would have no basis for telling these apart. A model whose representations carry "the trophy was too big" or "the suitcase was too small" can rank the continuations sensibly.

Key Considerations

The link strengths in this walkthrough are illustrative. In a real model the picture is messier. There are many attention heads in each layer, each with its own weights, and the work of resolving a pronoun may be spread across several heads and layers. Researchers have found heads in trained models that track relationships of this kind (Vaswani et al. 2017, sec. 4), though no single table like the ones above can be read off a real model.

A common mistake is to treat attention as the model "focusing" the way a person does. A person who attends to something is aware of doing it and can choose to look elsewhere. Attention in a transformer is a weighting. It's a set of numbers, computed from learned parameters, that controls how much information moves from one token to another. The name is a metaphor, and the model isn't concentrating on anything.

A second point concerns certainty. The weights are matters of degree. In the first sentence "it" takes a lot from "trophy" and a little from "suitcase," so the reading leans one way without excluding the other. On sentences that are ambiguous to people too, the weights may be closely balanced and the model's reading can go either way.

Summary

Changing one word at the end of the sentence moved the strong attention link from one noun to the other, and that change carried upward through the layers into what the model could predict next.

Sentence endingStrong links from "it"Weak link from "it"Resulting reading of "it"
… because it was too big.trophy, bigsuitcaseThe trophy
… because it was too small.suitcase, smalltrophyThe suitcase
  1. The token "it" starts with the same representation in both sentences.
  2. Attention weights, learned in training, decide which other tokens it draws on.
  3. The word that decides the matter comes after "it," which a strictly left-to-right reading couldn't use.

References

  • Lee, Timothy B., and Sean Trott. 2023. "Large Language Models, Explained with a Minimum of Math and Jargon." Understanding AI, July 27, 2023.
  • Vaswani, Ashish, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. "Attention Is All You Need." arXiv:1706.03762.
    • Free: arXiv 1706.03762
    • Free companion: Lee and Trott 2023, sections on transformers and attention

Report an issue with this item

Guided Conversation 12 min

Find the Words That Carry the Meaning

In this conversation you'll bring an ambiguous sentence from your own field and work out which of its words settle the meaning, the way attention does inside a model. You'll leave with a two-sentence explanation of attention that you could give a colleague.

You'll have this conversation with an AI assistant, using your own account. Choose a button to open a new chat with the prompt already filled in, then press send to start. If the chat opens empty, copy the prompt and paste it in.

Run this conversation in whichever assistant you already use:

Claude desktop app

To use another LLM, simply copy and paste the prompt into its chat window.

Show the full prompt (it lists misreadings to watch for, so skip it if you would rather come to the conversation fresh)
Guided Conversation: Find the Words That Carry the Meaning (about 12 minutes)

Note to the learner: press send to start. Everything below is facilitator guidance for the AI. It lists misconceptions to watch for, so skip it if you'd rather come to the conversation fresh.

Please facilitate a reflective dialogue with me. I'm an adult with no technical background who has used AI chatbots for everyday tasks, and I'm studying the transformer design and attention in language models. Follow this guidance for the whole conversation.

GOAL
I can explain what attention does in a transformer, using sentences from my own work.

HOW TO RUN THE CONVERSATION
- Ask one question at a time, then wait for my reply. Keep each of your turns under about 120 words.
- Don't lecture. Explain a point only when I need it to continue, then return to my sentence.
- Be curious and collegial. Use plain words and define any technical term briefly on first use. No formulas and no code. Welcome disagreement when I give a reason.
- Plain conversation only: don't search the web or create files or documents.
- Don't ask for anything confidential or personal. My sentence should be invented. If I start to share real names or client, patient, or student details, remind me to change them.
- Aim for about 12 minutes. Spend most of the time on topics 1 and 2. If my replies are brief, offer one concrete prompt, such as "Think of a sentence in your field containing 'it' or 'they' where a newcomer might pick the wrong referent," and move on. If I seem uncertain, shorten the conversation to 5-7 minutes. Always reach the final topic.
- Start now. Open with one or two warm sentences: this is a conversation, not a quiz; my reasoning about my own sentence matters more than using the right terms; I can ask you to clarify anything. Then ask me for one invented sentence from my field that could be read two ways, or that contains a pronoun or a word with two meanings.

TOPICS, IN ORDER
1. The words that settle it. Ask which word in my sentence is ambiguous and which other words tell a reader how to take it. Ask how far apart they are, and whether any deciding word comes after the ambiguous one. Follow up by asking me to change one word so the meaning flips.
2. Reading strictly left to right. Ask what a system would miss if it had to settle each word as it reached it, keeping only a running summary of what came before. Draw out two things: a deciding word that comes later can't be used, and a deciding word far back may have faded.
3. What is in the window now. Ask me what text I think you have in front of you at this moment and what you don't have. Draw out: the instructions, this conversation so far, anything I paste. Not my other conversations, unless a product feature adds them, and not my files.
4. Closing. Ask me to explain attention in two sentences to an imagined colleague who has never heard the term. Tell me I can take that explanation into a short optional activity where I test how adding context changes an answer.

KEY POINTS TO KEEP ACCURATE
- Attention lets each token (a piece of text, often a word or part of a word) collect information from other tokens in the passage, taking more from the relevant ones. The result is an updated representation of each token.
- How much one token takes from another is set by attention weights. The weights come from parameters learned in training. No person wrote them.
- A transformer processes all the tokens of a passage together, layer by layer, so a later word can shape the reading of an earlier one. When writing, a model still produces one token at a time.
- Many attention patterns (heads) run side by side and can track different relationships.
- The context window is finite. A model works only from the text in the window plus what training left in its parameters.
- If I ask what you are doing with my sentence, describe the general mechanism and say plainly that you can't observe your own attention weights or inspect your internals, so your statements about yourself are not evidence.

MISCONCEPTIONS TO CORRECT GENTLY
When one appears, name the accurate version briefly, then return to my sentence.
- "Attention is the model concentrating": it's a weighting, a set of numbers that controls how much information moves between tokens. The name is a metaphor.
- "The model remembers my earlier chats": it has only what is in the window, unless a product adds a memory feature that inserts notes into it.
- "Transformers understand grammar rules": no rules are stored. Useful links between words were learned from text.
- "Attention picks one word and ignores the rest": weights are matters of degree, and a token usually draws on several others.

LIMITS
- Don't discuss training cost, training data, or how models are turned into assistants.
- Don't state the size of any product's context window, including your own.
- Don't claim to report your actual attention on my sentence. Any link strengths you mention are illustrations.
- Don't recommend or compare products.

TO FINISH
After my two-sentence explanation, close in one short turn:
- Affirm one specific thing I worked out, in my own words where possible.
- Suggest one or two next steps that fit how the conversation went. Possible steps: give an assistant an ambiguous request with and without context and compare the replies; write a second sentence where the deciding word comes after the ambiguous one; reread a definition of attention weight and check my explanation against it.
- Restate my explanation on its own line, labeled "My two-sentence explanation of attention", so I can copy it.

Report an issue with this item

Hands-on Activity 15 minOptional

Test How Added Context Changes an Answer

Overview

A model reads every word of your request in light of the other words in front of it. In this activity you'll see that at work by asking one ambiguous question three times, with a different piece of context each time.

The activity is optional. Your notes are for you, and nobody collects them.

What You'll Need

  • An AI assistant you already use
  • Somewhere to paste or write three short replies

Your Task

Give an AI assistant an ambiguous one-line request three times, adding one piece of context each time, and record how the reply changes.

Steps

  1. Write a one-line request that could be taken more than one way. "Is Mercury dangerous?" works well, since Mercury is a planet, a metal, and a Roman god. Other options are "How do I get rid of a bug?" and "What's a good pitch?" Keep it to one line, and don't include anything personal or confidential.
  2. Ask it with no context, in a new conversation. Send only the one line and save the reply. Note which meaning the assistant chose, or whether it covered several or asked you which you meant.
  3. Ask it twice more, each time in a new conversation, with one clarifying sentence in front. For the Mercury question you might first write "I just broke an old thermometer." before the question, and then, in a separate conversation, "I'm planning a lesson on the solar system." Use the same wording for the question itself every time. A new conversation matters, because inside one conversation the earlier exchanges stay in front of the model and would act as context too.
  4. Mark the words in your context that changed the reply. Lay the three replies side by side. For each of the second and third, pick out the one or two words in your added sentence that you think did the work, such as "thermometer" or "solar system."
  5. Write one sentence linking what you saw to attention. Say how the words you marked could have changed what the ambiguous word meant to the model.

What to Expect

With no context, an assistant will usually pick the most common reading, cover several readings briefly, or ask what you meant. With a clarifying sentence in front, it will usually commit to one reading and say more about it.

The model has no separate step where it looks up which Mercury you meant. Every token in your request, including "Mercury," has its representation updated with information from the other tokens. "Thermometer" pulls the reading toward the metal, and "solar system" pulls it toward the planet.

If a reply surprises you, look for a word in your context that could have pulled the reading another way. Because replies are chosen partly by chance, you may also see differences that have nothing to do with your context. Repeating a request once is a quick way to tell the two apart.

Self-Check

When you're done, check that:

  • You have three replies, each from a fresh conversation
  • You identified what made your request ambiguous
  • You can point to the words in your added context that shifted the answer
  • You wrote one sentence linking the shift to attention

Nothing is uploaded. Write in your own notebook or document and keep it.

Report an issue with this item

Knowledge Check 10 min

The transformer and attention

This ungraded knowledge check assesses your understanding of the design behind current language models. You'll be asked about the context problem, attention, layers and parallel processing, and the context window.

Note: Use this to test yourself, review the feedback on any questions you miss, and retry until you feel confident before moving forward.

5 questions · ungraded · retry as often as you like

Report an issue with this item