Introduction
Accounts of how a language model is made tend to compress the whole process into a phrase such as "trained on the internet." The phrase skips every decision and every stage in between.
This walkthrough follows one training run from start to finish. The developer is invented and stands for no particular company. Quantities are given in rough orders of size, with dated examples from published sources where they exist.
The Starting Point
A developer has two things.
The first is a very large collection of raw text: a crawl of public web pages, plus books, reference works, and computer code.
The second is a transformer, the layered design used in current language models, with billions of parameters. Parameters are the adjustable numbers that determine what a model does. At the start they're set to random values, so the model's output is noise. Ask it to continue "The capital of France is" and it produces a jumble of unrelated word pieces.
The task is to get from those two starting materials to a finished model, and to be clear about what the finished model is.
Walking Through the Training Run
Step 1: Collect, filter, and tokenize the text
The raw collection is too messy to use as it is. It holds duplicate pages, spam, menus, and text generated by machines. The developer removes duplicates, filters out pages judged to be low quality, and decides how much weight each source gets.
These are choices, and they shape the model. When researchers at OpenAI described the text behind a model called GPT-3 in 2020, they reported filtering their web crawl by its similarity to known high-quality text, removing near-duplicates, and adding books and Wikipedia. The filtered crawl made up about three-fifths of the training mix (Brown et al. 2020, sec. 2.2).
The cleaned text is then split into tokens, the pieces of text, often words or parts of words, that a model handles as units. Each token is replaced by its number on the model's fixed list. What goes into training is a very long sequence of numbers. For GPT-3 the filtered crawl alone came to about 410 billion tokens (Brown et al. 2020, sec. 2.2). The 2026 survey of the field from Stanford University's Institute for Human-Centered Artificial Intelligence (Stanford HAI) reports that leading models since then have trained on tens of trillions (Stanford HAI 2026, chap. 1).
Step 2: Start the run: predict, score, adjust
Training repeats one cycle.
- Take a stretch of text from the collection.
- At each position, hide the next token and have the model rank every possible token by how likely it is to come next.
- Reveal the real token. Score the model by how low it ranked that token.
- Adjust every parameter slightly, so that the real token would rank a little higher next time.
No person supplies answers. The text provides its own, since the right answer at each position is the token that was there. Lee and Trott, authors of a plain-language explainer, describe the parameters as being gradually adjusted to make better and better predictions (Lee and Trott 2023, "How language models are trained").
The cycle runs on many thousands of specialized chips in a data center, with the work split among them. Each cycle handles many stretches of text at once, and the run consists of an enormous number of cycles.
Step 3: Look in midway
Developers check a model as it trains. What follows is a typical progression, described loosely. The details vary from one run to another.
Very early, the model's continuations are noise. After a small share of the text, it has picked up the most frequent regularities. Common words appear, spaces fall in the right places, and short runs of words look like language. Asked to continue "The capital of France is," it might write "the of and to the."
Further in, grammar settles. Sentences have subjects and verbs, and paragraphs stay loosely on a topic. The model might now write "The capital of France is a city in the north of the country," which is fluent and says little.
Later still, widely repeated facts arrive, and the model writes "Paris." Rarer facts take longer and some never arrive. Throughout, the score from Step 2 falls quickly at first and then more and more slowly.
Step 4: End the run
The run stops when the developer decides to stop it. The usual reasons are that the model has been through the planned amount of text, the score has nearly stopped improving, or the budget is spent. The 2026 survey reports that leading models have trained for periods longer than 100 days (Stanford HAI 2026, chap. 1).
Two things become fixed at this moment.
The parameters are frozen. From here on, using the model doesn't change them. Each conversation reads the same numbers.
The knowledge cutoff is set. It's the date after which nothing is reflected in the parameters. In practice it falls somewhat before the end of the run, on the date the text collection was closed. Nothing written after that date had any effect on the model.
Step 5: Take stock of what exists
What the developer now has is a file of numbers. Loaded onto chips, it does one thing: given some text, it ranks what's likely to come next. It was trained on web pages, books, and code, so it continues text the way such documents continue.
Give it "The capital of France is" and it continues "Paris." Give it a question such as "What should I cook tonight?" and the result is less predictable. It may answer. It may also continue with more questions, as a forum post would, or with a paragraph of a cooking blog. It has no settled manner, no habit of declining anything, and no sense that it's in a conversation with you.
A model at this stage is called a base model or a pretrained model. Making it behave as an assistant is further work, done afterward with additional training and written instructions.
Key Considerations
The quantities here are rough orders of size, "billions" and "months." The specific figures are dated examples: the GPT-3 numbers describe one model from 2020, and the survey figures describe the field as reported in 2026. Developers of many leading models haven't published comparable figures.
A common mistake is to believe that a deployed model keeps learning from the web, or from its conversations, after training. It doesn't. The parameters are frozen at the end of the run. When a chatbot gives you today's news, a search tool has fetched pages and placed them in front of the model, and the model has read them the way it reads anything you paste in. A developer may later use collected conversations to help train a new version. That is a separate training run producing a different set of parameters.
A second point concerns Step 3. The order in which abilities appear, with frequent patterns first and rare facts last, follows from the method. Each adjustment is small, so patterns that recur across many passages build up fastest.
Summary
A training run turns a text collection and a randomly set transformer into a fixed set of parameters that continues text.
| Stage | What goes in | What comes out |
|---|
| 1. Prepare the text | Raw web pages, books, reference works, code | A filtered collection, split into tokens |
| 2. Start the run | Token sequences and random parameters | Parameters adjusted slightly, cycle after cycle |
| 3. Midway | More text and more cycles | A model with grammar and common facts, still improving |
| 4. End the run | The developer's decision to stop | Frozen parameters and a fixed knowledge cutoff |
| 5. Take stock | The finished parameters | A base model that continues text, with no assistant manner |
References
- Brown, Tom B., Benjamin Mann, Nick Ryder, and 28 others. 2020. "Language Models Are Few-Shot Learners." arXiv:2005.14165.
- Lee, Timothy B., and Sean Trott. 2023. "Large Language Models, Explained with a Minimum of Math and Jargon." Understanding AI, July 27, 2023.
- Stanford Institute for Human-Centered Artificial Intelligence. 2026. The 2026 AI Index Report. Stanford University.