KnowledgeInSight
AI Literacy
0% of Course 1 complete

Module 2 · Lesson 2

Building a Model: Pretraining at scale: data, compute, and cost

You'll see what it takes to pretrain a large language model: the text, the chips, the time, and the money. You'll be able to say what a finished model keeps from all that text and what it doesn't.

What you will be able to do

  • Describe what pretraining requires in data, computing power, and cost, and what a trained model retains.

0% of this lesson · 9 items · 1h 3m total · 48m without the optional activity

Contents of this lesson9 items
  1. ReadingWhy Only a Few Organizations Train the Largest Models3 min
  2. ReadingPretraining Data: What Large Text Collections Contain and Leave Out4 min
  3. ReadingCompute: Specialized Chips, Data Centers, and Training Runs4 min
  4. ReadingScale and the Bitter Lesson: Why Larger Models Trained on More Data Improved4 min
  5. ReadingWhat a Trained Model Retains: Knowledge Cutoffs, Compression, and Memorization4 min
  6. Guided ReadingGuided Walkthrough: A Training Run from Raw Text to Finished Parameters7 min
  7. Guided ConversationWork Out What a Model Could Know12 min
  8. Hands-on Activity · optionalTest a Model's Knowledge Cutoff15 min
  9. Knowledge CheckPretraining at scale: data, compute, and cost10 min

Reading 3 min

Why Only a Few Organizations Train the Largest Models

This content reflects the field as of October 2026.

Follow AI news for a few months and the same names keep coming up: OpenAI, Google, Anthropic, Meta, xAI, Microsoft, and Chinese firms such as Alibaba and DeepSeek. Thousands of companies sell AI products. The large models underneath those products come from a short list of developers.

The pattern shows up in the numbers. Each year, Stanford University's Institute for Human-Centered Artificial Intelligence (Stanford HAI) publishes a survey of the field that includes a count of "notable" AI models, using a list kept by the research group Epoch AI. For 2025 the count was 93 notable models from industry and two from universities. By country, 59 came from the United States, 35 from China, and 8 from South Korea. The three organizations with the most were OpenAI with 20, Google with 14, and Alibaba with 11 (Stanford HAI 2026, chap. 1).

The balance has shifted over time. The same survey notes that industry's share has grown steadily over the past decade, and universities now produce very few of the models on the list.

The reason is material. The method for training a language model is published and widely taught. What most organizations lack is the means to carry it out at the largest scale. Training a leading model takes three things in very large amounts.

  • Text. Leading models have been trained on tens of trillions of tokens, the word-sized pieces of text a model handles.
  • Computing power. Training runs on many thousands of specialized chips working together for months.
  • Electricity. The survey reports that the most demanding training efforts in its data drew upward of 100 million watts (Stanford HAI 2026, chap. 1).

Each of these costs money on a scale that few organizations can raise, and a university research budget doesn't come close. An organization that can't pay for the chips can still build on a model someone else trained. It can't produce a comparable one from scratch.

The survey also reports that less is being disclosed. Several of the most resource-intensive models of recent years, including models from OpenAI, Anthropic, and Google, were released without public figures for their size, the amount of text they were trained on, or how long training took (Stanford HAI 2026, chap. 1).

The concentration raises questions that are argued over in courts, legislatures, and the press.

  • Whose text were the models trained on, and on what terms?
  • How much electricity and water does this work consume, and who bears the cost?
  • What follows when a technology many people rely on is produced by a handful of firms?

Those debates turn on facts about how the models are made: what goes into one, what the process consumes, and what the finished model keeps from the text it was trained on.

References

  • Stanford Institute for Human-Centered Artificial Intelligence. 2026. The 2026 AI Index Report. Stanford University.

Report an issue with this item

Reading 4 min

Pretraining Data: What Large Text Collections Contain and Leave Out

This content reflects the field as of October 2026.

Introduction

A chatbot can discuss Roman history, write computer code, and explain a tax form. It tends to be weaker on your town's local politics or on a language with few speakers. Both facts trace back to the text the model was trained on.

This reading explains where that text comes from, what developers do to it, and what it over-represents and leaves out.

Where the Text Comes From

Pretraining is the first and largest stage of training, in which a model learns to predict the next piece of text across a very large and varied collection. It gives a model its general knowledge of language and the world. The collection is called a corpus: a body of text gathered for training a model.

Most of a pretraining corpus comes from the public web. A web crawl is a collection of web pages gathered by software that follows links from page to page and saves what it finds. Developers add other sources to it, commonly books, reference works such as encyclopedias, and computer code.

One of the few detailed public accounts comes from 2020, when researchers at OpenAI described the corpus for a model called GPT-3. Its parts were weighted like this in training (Brown et al. 2020, sec. 2.2).

SourceShare of the training mix
A filtered web crawlAbout three-fifths
A second collection of web pagesAbout one-fifth
Two collections of booksAbout one-sixth
English-language WikipediaA few percent

The filtered crawl alone came to about 410 billion tokens, where a token is a piece of text such as a word or part of a word. Corpora have grown since. The 2026 survey of the field from Stanford University's Institute for Human-Centered Artificial Intelligence (Stanford HAI) reports that leading models were trained on tens of trillions of tokens (Stanford HAI 2026, chap. 1).

What Developers Do to It

A raw crawl contains spam, duplicated pages, navigation menus, and machine-generated junk. Data filtering is the removal or down-weighting of text a developer judges unsuitable for training. The GPT-3 authors describe three steps: keeping crawled pages that resembled known high-quality text, removing near-duplicate documents, and adding trusted collections to the mix (Brown et al. 2020, sec. 2.2).

Each step is a choice. Someone decides what counts as high quality, which sites to include, and which to block. Two developers starting from the same crawl can end up with different corpora, and their models will differ accordingly.

What Is Over- and Under-Represented

Representation is the share of a corpus that comes from or concerns a given language, period, group, or viewpoint. A corpus built from the web inherits the web's unevenness.

  • Languages. The GPT-3 authors report that 93 percent of the words in their corpus were English (Brown et al. 2020, sec. 3.3). Languages with a large web presence are well covered. Languages spoken by millions of people who publish little online are thinly covered.
  • Periods. The web holds far more text from the last two decades than from earlier ones. Older material is present mostly where someone has digitized it.
  • Viewpoints. People who write a lot online are over-represented. So are institutions that publish heavily, and so are the topics that draw argument. Knowledge that lives in private documents, in speech, or behind a login is largely absent.

A model's strengths follow these proportions. It has seen a great deal about widely discussed subjects in English and little about a small community's affairs.

How Much Developers Disclose

The GPT-3 account is unusual for its detail. Stanford HAI's survey reports that several of the most resource-intensive models of recent years, including models from OpenAI, Anthropic, and Google, were released without public figures for the size of their training datasets (Stanford HAI 2026, chap. 1). Exact contents are disclosed even less often.

So when you ask what a given model was trained on, the answer usually available is a general description: public web text, licensed material, books, code. The model can't fill in the rest. It has no list of its training documents, and what it says about its own training isn't evidence.

Whose Text It Is

Much of the text in these corpora was written by people who weren't asked. Whether training on it requires permission or payment is a live dispute, with lawsuits by authors and publishers, licensing agreements between some publishers and developers, and proposed laws in several countries. The dispute isn't settled, and it isn't argued here.

Conclusion

A model's general knowledge comes from its pretraining corpus, which is mostly web text plus books, reference works, and code. Developers filter that text according to their own judgments and disclose little about the result. The corpus over-represents English, recent years, and prolific sources, and a model's knowledge is uneven in the same ways.

Key Terms

  • Pretraining: The first and largest stage of training, in which a model learns to predict the next piece of text across a very large and varied collection.
  • Corpus: A body of text gathered for training a model.
  • Web crawl: A collection of web pages gathered by software that follows links from page to page and saves what it finds.
  • Data filtering: The removal or down-weighting of text a developer judges unsuitable for training.
  • Representation: The share of a corpus that comes from or concerns a given language, period, group, or viewpoint.

References

  • Brown, Tom B., Benjamin Mann, Nick Ryder, and 28 others. 2020. "Language Models Are Few-Shot Learners." arXiv:2005.14165.
  • Stanford Institute for Human-Centered Artificial Intelligence. 2026. The 2026 AI Index Report. Stanford University.

Report an issue with this item

Reading 4 min

Compute: Specialized Chips, Data Centers, and Training Runs

This content reflects the field as of October 2026.

Introduction

News about AI is full of chip shortages, new data centers, and power deals. A chatbot is software, and it may seem odd that software should need its own power supply.

This reading explains what kind of computing a language model needs, what a training run involves, where the money goes, and how training differs from running a finished model.

Chips Built for Many Small Calculations

Compute is the field's word for computing power: the chips, and the time on them, needed to train or run a model.

Training a model means adjusting its parameters, the numbers that determine what it does. A large model has many billions of them. Each small adjustment requires an enormous number of simple calculations, mostly multiplications and additions, and none of them is hard. The difficulty is the quantity.

The ordinary processor in a laptop is built to do varied tasks quickly, a few at a time. A GPU, short for graphics processing unit, is a chip that does a very large number of simple calculations at the same time. GPUs were designed for drawing video-game images, which poses the same kind of problem, and they turned out to suit neural networks well. Today's AI chips are descendants of them, built for this work.

Supply is concentrated. The 2026 survey of the field from Stanford University's Institute for Human-Centered Artificial Intelligence (Stanford HAI) estimates that the world's stock of AI chips has more than tripled each year since 2022, and that chips from one company, Nvidia, account for over 60 percent of the total (Stanford HAI 2026, chap. 1).

A Training Run

One chip is nowhere near enough. Developers link many thousands of them in a data center, a building that houses computers along with the power and cooling they need. The survey counted 5,427 data centers of all kinds in the United States in 2025, more than ten times as many as in any other country (Stanford HAI 2026, chap. 1).

A training run is one continuous stretch of training that takes a model from its starting parameters to its finished ones. The model works through its training text, predicting and being adjusted, around the clock. According to the survey, leading models have trained for periods longer than 100 days. The most demanding runs in its data, for xAI's Grok 3 and Meta's Llama 4 Behemoth, drew upward of 100 million watts (Stanford HAI 2026, chap. 1). Figures for many other leading models, including models from OpenAI, Anthropic, and Google, haven't been published.

A run can fail. Hardware breaks, and a design choice can turn out badly weeks in. Developers also run many smaller experiments before and alongside the main run, and those use compute too.

Where the Money Goes

CostWhat it covers
ChipsBuying or renting many thousands of AI chips
ElectricityPower for the chips throughout the run
CoolingRemoving the heat the chips produce, which takes more power and often water
StaffResearchers and engineers who design the model and keep the run going

The 2026 International AI Safety Report, written by an international panel of experts chaired by the computer scientist Yoshua Bengio, puts the cost of acquiring the resources to develop a leading system from scratch at hundreds of millions of US dollars (Bengio and others 2026, sec. 1.1). Treat such figures as dated estimates. Developers rarely publish their own costs, and the amounts have risen with each generation of models.

Training and Running Are Separate Costs

Training happens once per model. Using the finished model is a different activity. Inference is the running of a trained model to produce output, as happens each time someone sends it a message.

TrainingInference
When it happensOnce, before releaseEvery time the model is used
What changesThe parameters are adjustedNothing: the parameters stay fixed
What drives the costThe size of the model and of its training textThe number of users and the length of their requests

A single reply takes a tiny amount of compute next to a training run. Multiplied across many millions of users a day, inference becomes a large and continuing expense. That's one reason developers charge for heavy use and offer smaller, cheaper models alongside their largest ones.

Conclusion

Pretraining is a months-long computation on many thousands of specialized chips in a data center. Chips, electricity, cooling, and staff make it expensive, and once a model is in use, inference adds a running cost on top. The published figures are partial, since many developers disclose little about their own runs.

Key Terms

  • Compute: Computing power: the chips, and the time on them, needed to train or run a model.
  • GPU: A chip that does a very large number of simple calculations at the same time.
  • Data center: A building that houses computers along with the power and cooling they need.
  • Training run: One continuous stretch of training that takes a model from its starting parameters to its finished ones.
  • Inference: The running of a trained model to produce output.

References

  • Bengio, Yoshua, and others. 2026. International AI Safety Report 2026. DSIT 2026/001. Published February 3, 2026.
  • Stanford Institute for Human-Centered Artificial Intelligence. 2026. The 2026 AI Index Report. Stanford University.

Report an issue with this item

Reading 4 min

Scale and the Bitter Lesson: Why Larger Models Trained on More Data Improved

Introduction

Each new generation of language models has been described as bigger than the last: more parameters, more training text, more chips. A common explanation of AI progress holds that size itself did most of the work.

This reading explains that argument, the result most often cited for it, what it leaves out, and why its future is disputed.

Sutton's Argument

In 2019 Rich Sutton, a computer scientist at the University of Alberta known for his work on reinforcement learning, published a short essay called "The Bitter Lesson." Its opening claim is that across seventy years of AI research, "general methods that leverage computation are ultimately the most effective, and by a large margin" (Sutton 2019).

A general method is a technique that isn't built around human knowledge of a particular task and that keeps improving as it's given more computation. Sutton names two, search and learning. He contrasts them with the approach researchers have often preferred, which is to build their own understanding of a problem into the system.

His evidence is historical. In chess, the program that beat the world champion in 1997 relied on searching through a vast number of positions, and the programs built around human chess knowledge lost out. He describes the same pattern in the game of Go and in speech recognition. Each time, in his telling, building in human knowledge helped at first and then stalled, and the breakthrough came from methods that used more computation (Sutton 2019).

The lesson is "bitter" because researchers' hard-won insight into a problem turned out to matter less than computing power. Sutton's essay is an argument from selected cases, and other researchers dispute how far it generalizes.

The GPT-3 Result

A year later, a team at OpenAI reported a result that is often read as support for Sutton's view. They are reporting on their own company's model, and the paper is a preprint, posted publicly without a journal's review.

In this field, scale means the size of a model, of its training text, and of the computing power used to train it. Scaling is the practice of increasing those three in the expectation that capability will rise. The team trained GPT-3 with 175 billion parameters, the adjustable numbers in a model. They describe that as ten times more than any comparable earlier model (Brown et al. 2020, abstract). The design was the transformer, already known. The main change was size.

The larger model could do something smaller ones did poorly. Few-shot learning is a model's carrying out of a new task after seeing only a few examples of it in the prompt, with no further training. Give the model three English sentences with their French translations, then a fourth English sentence, and it continues with a French translation. The paper stresses that this happens with no change to the model's parameters. The examples are just text in front of it (Brown et al. 2020, sec. 2).

GPT-3 did this across translation, question answering, and simple arithmetic, and the authors also list tasks where it struggled. Nobody had trained it for these tasks. The ability appeared as the model grew, and the authors report that it kept improving with size across the models they compared.

Other developers, among them Google, Meta, and Anthropic, went on to train very large models of their own.

What Scale Doesn't Explain

Size is part of the story. Three other things contributed to what current assistants can do.

  • Training methods. A model fresh from pretraining continues text. Turning it into something that answers questions took additional training on examples of the wanted behavior and on people's judgments of its outputs.
  • Tools. Much of what assistants do now depends on connecting a model to search, calculators, and other software.
  • Data choices. Which text a model is trained on, and how it's filtered, affects quality independently of quantity.

A fair summary is that scale was necessary for the recent gains and didn't produce them alone.

Whether Scaling Keeps Paying Off

Whether further scaling will keep producing large gains is an open question. One camp, which includes leaders of several AI companies with a commercial stake in the answer, expects it to. Another camp expects the gains to shrink as high-quality training text runs short and the cost of each step rises, and holds that new ideas will be needed. Sutton's essay is evidence about the past. It can't settle a forecast, and neither side's forecast has been confirmed.

Conclusion

Sutton argued that methods which make use of more computation have repeatedly beaten methods built on human knowledge. GPT-3 fit that pattern: a known design made much larger could handle new tasks from a few examples. Later progress also drew on training methods, tools, and data choices, and whether scaling will continue to pay off is disputed.

Key Terms

  • General method: A technique that isn't built around human knowledge of a particular task and that keeps improving as it's given more computation.
  • Scale: The size of a model, of its training text, and of the computing power used to train it.
  • Scaling: The practice of increasing a model's size, training text, and computing power in the expectation that capability will rise.
  • Few-shot learning: A model's carrying out of a new task after seeing only a few examples of it in the prompt, with no further training.

References

  • Brown, Tom B., Benjamin Mann, Nick Ryder, and 28 others. 2020. "Language Models Are Few-Shot Learners." arXiv:2005.14165.
  • Sutton, Rich. 2019. "The Bitter Lesson." March 13, 2019.

Report an issue with this item

Reading 4 min

What a Trained Model Retains: Knowledge Cutoffs, Compression, and Memorization

Introduction

Two claims about language models circulate widely. One says a model is a giant database that stores everything it was trained on. The other says a model stores nothing and only learns general patterns. Both are wrong, and the evidence shows where.

This reading explains what stays in a model after training: a fixed point in time, a compressed summary, and some passages kept word for word.

A Fixed Point in Time

Training ends on a particular day, and the training text was collected before that. The knowledge cutoff is the date after which nothing is reflected in a model's parameters, because its training text ends there. Parameters are the adjustable numbers in a model, set by training and fixed afterward.

A model released months after its cutoff, and used for a year or two after that, has nothing in its parameters about events in between. Asked about them, it may say it doesn't know, or it may produce a plausible answer from older patterns. Many chatbot products now attach a search tool that fetches current pages and places them in front of the model. That supplies new text for the model to read. The parameters stay as they were.

Models are also unreliable about their own cutoff. The last months before a cutoff are thinly covered, since the web hasn't finished writing about them yet, so a model may place its cutoff earlier than it is.

A Compressed Summary

Compression is the reduction of a large body of information to a smaller form that keeps its main patterns and loses detail. Training has this effect on text.

The mechanism explains why. Lee and Trott, authors of a plain-language explainer on language models, describe a model's parameters as being gradually adjusted to make better predictions of the next word (Lee and Trott 2023, "How language models are trained"). Each passage nudges the numbers slightly, and each passage is typically seen once or a few times. A nudge that helps with one sentence only is soon overwritten by nudges from others. What accumulates is whatever helps across many passages: grammar, common facts, typical ways of arguing and explaining.

Size points the same way. Lee and Trott report that one model from 2020 had 175 billion parameters and was trained on about 500 billion words (Lee and Trott 2023). The numbers that make up a model occupy less storage than its training text, and they also have to carry every skill the model has.

The result resembles a well-read person's memory more than a library. The model can tell you what a famous novel is about and how its plot goes. It usually can't recite page 212.

Some Passages Survive Word for Word

Usually isn't never. Memorization means getting training examples right by retaining them individually, without capturing a pattern that carries over to new cases. In a language model it shows up as the ability to reproduce a passage exactly.

In 2020 a team of twelve researchers from Google, OpenAI, Apple, and four universities tested this on GPT-2, an earlier and smaller model whose training text was known. Training data extraction is the recovery of passages from a model's training text by prompting the model and examining what it writes. The team generated a large number of outputs, picked the 1,800 that looked most likely to be memorized, and checked them against the training text. Of those, 604 were word-for-word copies (Carlini et al. 2020, sec. 6).

The copies included news text, software licenses, and people's names with contact details. The authors report that their method worked on some passages that appeared in "just one document in the training data" (Carlini et al. 2020, abstract). They also found that larger models memorized more than smaller ones.

A few hundred passages is a tiny fraction of that model's training text, and the researchers went looking for them with a method built for the purpose. The finding still matters. Text that is repeated many times across a corpus, such as famous quotations, standard legal wording, and song lyrics, is the likeliest to be retained. Unusual strings can be retained too.

What the Two Slogans Get Wrong

ClaimWhat the evidence shows
"It stores everything"Most training text can't be reproduced. The model keeps patterns and loses detail
"It stores nothing"Some passages can be extracted word for word, including text that appeared once
"It knows what's happening now"Its parameters reflect text up to the cutoff and nothing later

Conclusion

A trained model holds a compressed summary of its training text, fixed at its knowledge cutoff. Most of that text is retained as patterns only, and some passages are memorized exactly. How much a given model has memorized can be established only by testing it, as the extraction study did.

Key Terms

  • Knowledge cutoff: The date after which nothing is reflected in a model's parameters, because its training text ends there.
  • Compression: The reduction of a large body of information to a smaller form that keeps its main patterns and loses detail.
  • Memorization: Getting training examples right by retaining them individually, without capturing a pattern that carries over to new cases.
  • Training data extraction: The recovery of passages from a model's training text by prompting the model and examining what it writes.

References

  • Carlini, Nicholas, Florian Tramèr, Eric Wallace, and 9 others. 2020. "Extracting Training Data from Large Language Models." arXiv:2012.07805.
  • Lee, Timothy B., and Sean Trott. 2023. "Large Language Models, Explained with a Minimum of Math and Jargon." Understanding AI, July 27, 2023.

Report an issue with this item

Guided Reading 7 min

Guided Walkthrough: A Training Run from Raw Text to Finished Parameters

Introduction

Accounts of how a language model is made tend to compress the whole process into a phrase such as "trained on the internet." The phrase skips every decision and every stage in between.

This walkthrough follows one training run from start to finish. The developer is invented and stands for no particular company. Quantities are given in rough orders of size, with dated examples from published sources where they exist.

The Starting Point

A developer has two things.

The first is a very large collection of raw text: a crawl of public web pages, plus books, reference works, and computer code.

The second is a transformer, the layered design used in current language models, with billions of parameters. Parameters are the adjustable numbers that determine what a model does. At the start they're set to random values, so the model's output is noise. Ask it to continue "The capital of France is" and it produces a jumble of unrelated word pieces.

The task is to get from those two starting materials to a finished model, and to be clear about what the finished model is.

Walking Through the Training Run

Step 1: Collect, filter, and tokenize the text

The raw collection is too messy to use as it is. It holds duplicate pages, spam, menus, and text generated by machines. The developer removes duplicates, filters out pages judged to be low quality, and decides how much weight each source gets.

These are choices, and they shape the model. When researchers at OpenAI described the text behind a model called GPT-3 in 2020, they reported filtering their web crawl by its similarity to known high-quality text, removing near-duplicates, and adding books and Wikipedia. The filtered crawl made up about three-fifths of the training mix (Brown et al. 2020, sec. 2.2).

The cleaned text is then split into tokens, the pieces of text, often words or parts of words, that a model handles as units. Each token is replaced by its number on the model's fixed list. What goes into training is a very long sequence of numbers. For GPT-3 the filtered crawl alone came to about 410 billion tokens (Brown et al. 2020, sec. 2.2). The 2026 survey of the field from Stanford University's Institute for Human-Centered Artificial Intelligence (Stanford HAI) reports that leading models since then have trained on tens of trillions (Stanford HAI 2026, chap. 1).

Step 2: Start the run: predict, score, adjust

Training repeats one cycle.

  1. Take a stretch of text from the collection.
  2. At each position, hide the next token and have the model rank every possible token by how likely it is to come next.
  3. Reveal the real token. Score the model by how low it ranked that token.
  4. Adjust every parameter slightly, so that the real token would rank a little higher next time.

No person supplies answers. The text provides its own, since the right answer at each position is the token that was there. Lee and Trott, authors of a plain-language explainer, describe the parameters as being gradually adjusted to make better and better predictions (Lee and Trott 2023, "How language models are trained").

The cycle runs on many thousands of specialized chips in a data center, with the work split among them. Each cycle handles many stretches of text at once, and the run consists of an enormous number of cycles.

Step 3: Look in midway

Developers check a model as it trains. What follows is a typical progression, described loosely. The details vary from one run to another.

Very early, the model's continuations are noise. After a small share of the text, it has picked up the most frequent regularities. Common words appear, spaces fall in the right places, and short runs of words look like language. Asked to continue "The capital of France is," it might write "the of and to the."

Further in, grammar settles. Sentences have subjects and verbs, and paragraphs stay loosely on a topic. The model might now write "The capital of France is a city in the north of the country," which is fluent and says little.

Later still, widely repeated facts arrive, and the model writes "Paris." Rarer facts take longer and some never arrive. Throughout, the score from Step 2 falls quickly at first and then more and more slowly.

Step 4: End the run

The run stops when the developer decides to stop it. The usual reasons are that the model has been through the planned amount of text, the score has nearly stopped improving, or the budget is spent. The 2026 survey reports that leading models have trained for periods longer than 100 days (Stanford HAI 2026, chap. 1).

Two things become fixed at this moment.

The parameters are frozen. From here on, using the model doesn't change them. Each conversation reads the same numbers.

The knowledge cutoff is set. It's the date after which nothing is reflected in the parameters. In practice it falls somewhat before the end of the run, on the date the text collection was closed. Nothing written after that date had any effect on the model.

Step 5: Take stock of what exists

What the developer now has is a file of numbers. Loaded onto chips, it does one thing: given some text, it ranks what's likely to come next. It was trained on web pages, books, and code, so it continues text the way such documents continue.

Give it "The capital of France is" and it continues "Paris." Give it a question such as "What should I cook tonight?" and the result is less predictable. It may answer. It may also continue with more questions, as a forum post would, or with a paragraph of a cooking blog. It has no settled manner, no habit of declining anything, and no sense that it's in a conversation with you.

A model at this stage is called a base model or a pretrained model. Making it behave as an assistant is further work, done afterward with additional training and written instructions.

Key Considerations

The quantities here are rough orders of size, "billions" and "months." The specific figures are dated examples: the GPT-3 numbers describe one model from 2020, and the survey figures describe the field as reported in 2026. Developers of many leading models haven't published comparable figures.

A common mistake is to believe that a deployed model keeps learning from the web, or from its conversations, after training. It doesn't. The parameters are frozen at the end of the run. When a chatbot gives you today's news, a search tool has fetched pages and placed them in front of the model, and the model has read them the way it reads anything you paste in. A developer may later use collected conversations to help train a new version. That is a separate training run producing a different set of parameters.

A second point concerns Step 3. The order in which abilities appear, with frequent patterns first and rare facts last, follows from the method. Each adjustment is small, so patterns that recur across many passages build up fastest.

Summary

A training run turns a text collection and a randomly set transformer into a fixed set of parameters that continues text.

StageWhat goes inWhat comes out
1. Prepare the textRaw web pages, books, reference works, codeA filtered collection, split into tokens
2. Start the runToken sequences and random parametersParameters adjusted slightly, cycle after cycle
3. MidwayMore text and more cyclesA model with grammar and common facts, still improving
4. End the runThe developer's decision to stopFrozen parameters and a fixed knowledge cutoff
5. Take stockThe finished parametersA base model that continues text, with no assistant manner

References

  • Brown, Tom B., Benjamin Mann, Nick Ryder, and 28 others. 2020. "Language Models Are Few-Shot Learners." arXiv:2005.14165.
  • Lee, Timothy B., and Sean Trott. 2023. "Large Language Models, Explained with a Minimum of Math and Jargon." Understanding AI, July 27, 2023.
  • Stanford Institute for Human-Centered Artificial Intelligence. 2026. The 2026 AI Index Report. Stanford University.

Report an issue with this item

Guided Conversation 12 min

Work Out What a Model Could Know

In this conversation you'll take five pieces of information and reason about whether a language model could hold each one in its parameters, and in what form. You'll leave with your own rule of thumb for when to distrust a model's recall.

You'll have this conversation with an AI assistant, using your own account. Choose a button to open a new chat with the prompt already filled in, then press send to start. If the chat opens empty, copy the prompt and paste it in.

Run this conversation in whichever assistant you already use:

Claude desktop app

To use another LLM, simply copy and paste the prompt into its chat window.

Show the full prompt (it lists misreadings to watch for, so skip it if you would rather come to the conversation fresh)
Guided Conversation: Work Out What a Model Could Know (about 12 minutes)

Note to the learner: press send to start. Everything below is facilitator guidance for the AI. It lists misconceptions to watch for, so skip it if you'd rather come to the conversation fresh.

Please facilitate a coached problem session with me. I'm an adult with no technical background who has used AI chatbots for everyday tasks, and I'm studying what pretraining a language model requires and what a trained model retains. Follow this guidance for the whole conversation.

GOAL
I can describe what a trained model retains, by reasoning about specific pieces of information.

HOW TO RUN THE CONVERSATION
- Ask one question at a time, then wait for my reply. Keep each of your turns under about 120 words.
- Don't lecture. Explain a point only when I need it to continue, then return to my five items.
- Be curious and collegial. Use plain words and define any technical term briefly on first use. No formulas and no code. Welcome disagreement when I give a reason.
- This is a coached problem. Ask for my reasoning before you give any hint. Give one hint at a time. Don't give your own verdict on an item until I've made my attempt.
- Plain conversation only: don't search the web or create files or documents.
- Don't ask for anything confidential or personal. For my organization's handbook I should name it only in general terms and share none of its contents.
- Aim for about 12 minutes. Spend most of the time on topics 1 and 2. If my replies are brief, offer one concrete prompt, such as "How many times do you think that text appears on the public web?" and move on. If I seem uncertain, shorten the conversation to 5-7 minutes. Always reach the final topic.
- Start now. Open with one or two warm sentences: this is a conversation, not a quiz; my reasoning matters more than the verdicts; I can ask you to clarify anything. Then give me the problem: ask me to name one example of each of these five items. (a) the plot of a famous novel, (b) a news event from last month, (c) my own organization's internal handbook, (d) a common recipe, (e) an obscure fact about my town or neighborhood.

TOPICS, IN ORDER
1. In the parameters or not. Take my five items one at a time. For each, ask whether a general-purpose language model could have it in its parameters and why. Draw out the two tests: was the text public and collected before training ended, and how often did it appear?
2. In what form. For each item I judged could be there, ask whether the model would hold a summary, word-for-word text, or nothing reliable. Follow up on one where I'm unsure.
3. Filling the gaps. Ask how a chatbot product could supply what the parameters lack. Draw out: a search tool that fetches current pages, or documents I paste or upload, both of which put text in front of the model without changing its parameters.
4. Closing. Ask me to state a rule of thumb for when to distrust a model's recall. Tell me I can take it into a short optional activity where I test where an assistant's built-in knowledge ends.

KEY POINTS TO KEEP ACCURATE
- Method: for each item ask (1) was it public text collected before the cutoff, (2) how frequent was it, (3) so is it likely held as a summary, verbatim, or not at all.
- Expected answers. Famous novel's plot: yes, as a summary, since it's widely discussed; exact passages are uncertain. Last month's news: not in the parameters if it falls after the knowledge cutoff, which is fixed when training ends. Internal handbook: no, unless it was public; only if I supply it. Common recipe: yes, repeated across many pages, so typical versions are well retained, though amounts can still be wrong. Obscure local fact: thin or absent, and a fluent wrong answer is likely.
- Frequent text is retained better than rare text. Most training text is kept as patterns only. Some passages are memorized word for word, and researchers have extracted such passages from a model.
- A model can't list its training data. Its own statement of its cutoff may be wrong.
- If I ask what you know or what you were trained on, explain the general mechanism and say plainly that you can't inspect your own internals or training data, so your statements about yourself are not evidence.

MISCONCEPTIONS TO CORRECT GENTLY
When one appears, name the accurate version briefly, then return to my items.
- "It searches the internet when I ask": only if a product adds a search tool. Otherwise it answers from fixed parameters.
- "It has read my documents": only if they were public before training or I supply them.
- "It can tell me exactly what it was trained on": it can't. It has no list of its training text.
- "It keeps learning from the web": the parameters are frozen when training ends.

LIMITS
- Take no position on data rights or copyright. If I raise them, say they are disputed and return to the mechanism.
- Give no figures you can't source, such as corpus sizes, costs, or cutoff dates. Don't state your own cutoff as fact.
- Don't recommend or compare products.

TO FINISH
After my rule of thumb, close in one short turn:
- Affirm one specific thing I worked out, in my own words where possible.
- Suggest one or two next steps that fit how the conversation went. Possible steps: ask an assistant about dated public events to find where its knowledge ends; look up the knowledge cutoff a developer documents for a model I use; reread a definition of memorization and apply it to my novel example.
- Restate my rule on its own line, labeled "My rule of thumb for a model's recall", so I can copy it.

Report an issue with this item

Hands-on Activity 15 minOptional

Test a Model's Knowledge Cutoff

Overview

A model's built-in knowledge ends at a fixed date. In this activity you'll look for that date by asking an AI assistant about events you can date yourself, then compare what you find with what the assistant says and with what its developer documents.

The activity is optional. Your notes are for you, and nobody collects them.

What You'll Need

  • An AI assistant you already use, with web search turned off if the product allows it
  • The developer's public documentation for the model, which you can find with an ordinary search engine
  • Somewhere to write a few notes

Your Task

Find where an AI assistant's built-in knowledge ends by asking about events at known dates, then compare the result with the assistant's own claim and with the cutoff its developer states.

Steps

  1. List five public events with known dates, spread across the past three years. Choose things with one checkable answer: who won a major sports final, which film took a top award, the result of a national election. Include one from the last two or three months. Write down the date and the correct answer for each before you start.
  2. Ask about each event in a new conversation and record the reply. Keep the question plain, such as "Who won the men's final at Wimbledon in 2025?" Mark each reply as accurate, vague, or wrong. Note whether the assistant says it doesn't know or shows that it searched the web. If you can't turn search off, add "Please answer without searching the web" and record whether it complied.
  3. Ask the assistant for its knowledge cutoff. Use a new conversation and write down exactly what it says.
  4. Look up the cutoff the developer documents. Search for the model's name with "knowledge cutoff," and prefer the developer's own documentation to other sites. If you can't tell which model your assistant is running, or the developer publishes no cutoff, write that down.
  5. Compare the three and write one sentence. Set the date where the replies stopped being accurate beside the assistant's claim and the documented cutoff. Then write one sentence on why the assistant's own claim isn't evidence.

What to Expect

Replies about older events will probably be accurate. Somewhere along your list they'll turn vague or wrong, or the assistant will say the event is after its knowledge. That turning point is your estimate of the cutoff. With five events it's a rough estimate, good to within several months.

The three dates often disagree. An assistant may name a cutoff earlier than the documented one, since the last months before a cutoff are thinly covered in its training text. It may also name a later one, or decline to say.

If every reply is accurate, including the most recent, search was probably on. In that case you've tested the product's search tool and learned nothing about the model's parameters.

A wrong answer about an event after the cutoff tells you something as well. The model produced a fluent reply where it had no information, and its wording didn't signal the difference.

Self-Check

When you're done, check that:

  • You tested five dated events, each in a new conversation
  • You recorded the assistant's own claim about its cutoff
  • You found the documented cutoff, or noted that none is published
  • You wrote one sentence on why the assistant's claim isn't evidence

Nothing is uploaded. Write in your own notebook or document and keep it.

Report an issue with this item

Knowledge Check 10 min

Pretraining at scale: data, compute, and cost

This ungraded knowledge check assesses your understanding of what pretraining a large language model takes and what it leaves behind. You'll be asked about pretraining data, compute and training runs, scale and the bitter lesson, and what a trained model retains.

Note: Use this to test yourself, review the feedback on any questions you miss, and retry until you feel confident before moving forward.

5 questions · ungraded · retry as often as you like

Report an issue with this item