Introduction
"Attention weighs which words matter to each other" is a fair summary of the mechanism, and it's hard to picture. Following one pronoun through a few layers of a model shows what the summary means in practice.
This walkthrough traces the word "it" in a sentence where its meaning flips when a single later word changes. The link strengths are invented to illustrate the mechanism and weren't measured from any real model.
The Starting Point
Here are two sentences that differ in their last word.
The trophy didn't fit in the suitcase because it was too big.
The trophy didn't fit in the suitcase because it was too small.
The pair comes from the Winograd Schema Challenge, a set of such puzzles proposed as a test for machines. You read "it" as the trophy in the first sentence and as the suitcase in the second, without effort. Nothing in the grammar tells you which. You use what you know about fitting things into containers.
A language model works on tokens, pieces of text that are often words or parts of words. For simplicity, treat each word here as one token. Inside the model, each token is represented by a list of numbers. Before any processing, the list for "it" is the same in both sentences and says nothing about trophies or suitcases. The task is to follow how that representation changes as the sentence passes through a transformer, the layered design used in current language models.
Walking Through the Sentence
Step 1: Mark the ambiguous token and its candidates
The ambiguous token is "it." Two earlier tokens could be what it refers to: "trophy" and "suitcase." Both are nouns, both are things, and both come before "it," so grammar leaves the choice open.
Other tokens bear on the choice even though they aren't candidates. "Fit" and "in" set up a relationship in which one thing goes inside another. "Big" or "small" names a property that explains why the fitting failed. A reader uses all of these, and a model has to as well.
Step 2: Describe the links from "it" in the first sentence
In an attention step, each token collects information from the other tokens, taking more from those relevant to it. Lee and Trott, authors of a plain-language explainer on language models, describe this as a matchmaking service in which each word looks for other words with relevant context (Lee and Trott 2023, "Can I have your attention please"). How much one token takes from another is set by an attention weight.
Take the first sentence, ending in "big." Suppose the weights from "it" come out like this in one layer.
| Token | Link from "it" |
|---|
| trophy | Strong |
| big | Strong |
| fit | Moderate |
| suitcase | Weak |
| The, in, the, because, was, too | Very weak |
After the step, the representation of "it" has been updated. It now holds a good deal of information about a trophy, some about fitting, and some about bigness. The token is still "it." Its list of numbers is now a blend of information from several tokens, in proportion to the weights.
Step 3: Change one word and watch the links shift
Now take the second sentence. Every token is the same until the last one, which is "small."
| Token | Link from "it" |
|---|
| suitcase | Strong |
| small | Strong |
| fit | Moderate |
| trophy | Weak |
| The, in, the, because, was, too | Very weak |
The strong link has moved from "trophy" to "suitcase." The representation of "it" now holds mostly information about a suitcase.
Notice where the deciding word sits. "Small" comes after "it." A system that read strictly left to right and settled each word as it went would have to commit on "it" before reaching the word that decides the matter. In a transformer reading a passage it already has, every token's update is computed with the whole sentence available, so a later word can shape the reading of an earlier one.
Nobody wrote a rule saying that a thing too big to fit must be the thing going in. The weights come from parameters, the adjustable numbers set by training. The model was trained on a great deal of text about objects, containers, and sizes, and weights that link these words usefully made its predictions better.
Step 4: Follow the updated representation through two more layers
A transformer is a stack of layers, and each layer runs attention again on the output of the layer below (Vaswani et al. 2017, sec. 3.1). Go back to the first sentence and follow it upward.
In the next layer, the tokens attend to one another again, and this time "it" already carries trophy information. So when "big" collects from "it," what it picks up is, in effect, trophy. The representation of "big" comes to hold "big, said of the trophy." Likewise "fit," which already linked to "didn't," picks up more about what failed to fit into what.
In the layer above that, tokens near the end of the sentence can draw on all of this. The representation at the final position comes to reflect something like the whole situation: a trophy, a suitcase, a failure to fit, and the trophy's size as the cause. Lee and Trott report research suggesting that lower layers tend to settle questions like which noun a pronoun refers to, while higher layers build a broader reading of the passage (Lee and Trott 2023, "Transforming word vectors into word predictions").
Step 5: State what the model can now predict
The point of all this is prediction. Suppose each sentence continues with "So we bought a."
For the first sentence, a good continuation is "bigger suitcase" or "smaller trophy." For the second, "bigger suitcase" still works, but "smaller trophy" would be odd, since the trophy's size was never the stated problem. A model that hadn't resolved "it" would have no basis for telling these apart. A model whose representations carry "the trophy was too big" or "the suitcase was too small" can rank the continuations sensibly.
Key Considerations
The link strengths in this walkthrough are illustrative. In a real model the picture is messier. There are many attention heads in each layer, each with its own weights, and the work of resolving a pronoun may be spread across several heads and layers. Researchers have found heads in trained models that track relationships of this kind (Vaswani et al. 2017, sec. 4), though no single table like the ones above can be read off a real model.
A common mistake is to treat attention as the model "focusing" the way a person does. A person who attends to something is aware of doing it and can choose to look elsewhere. Attention in a transformer is a weighting. It's a set of numbers, computed from learned parameters, that controls how much information moves from one token to another. The name is a metaphor, and the model isn't concentrating on anything.
A second point concerns certainty. The weights are matters of degree. In the first sentence "it" takes a lot from "trophy" and a little from "suitcase," so the reading leans one way without excluding the other. On sentences that are ambiguous to people too, the weights may be closely balanced and the model's reading can go either way.
Summary
Changing one word at the end of the sentence moved the strong attention link from one noun to the other, and that change carried upward through the layers into what the model could predict next.
| Sentence ending | Strong links from "it" | Weak link from "it" | Resulting reading of "it" |
|---|
| … because it was too big. | trophy, big | suitcase | The trophy |
| … because it was too small. | suitcase, small | trophy | The suitcase |
- The token "it" starts with the same representation in both sentences.
- Attention weights, learned in training, decide which other tokens it draws on.
- The word that decides the matter comes after "it," which a strictly left-to-right reading couldn't use.
References
- Lee, Timothy B., and Sean Trott. 2023. "Large Language Models, Explained with a Minimum of Math and Jargon." Understanding AI, July 27, 2023.
- Vaswani, Ashish, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. "Attention Is All You Need." arXiv:1706.03762.
- Free: arXiv 1706.03762
- Free companion: Lee and Trott 2023, sections on transformers and attention