KnowledgeInSight
AI Literacy
0% of Course 3 complete

Module 3 · Lesson 3

Information and Rights: Copyright, training data, and creative work

You'll examine the dispute over whether AI developers may train models on copyrighted work without permission. You'll be able to state each side's position and say what US and UK courts have ruled so far.

What you will be able to do

  • Compare the positions of creators and AI developers on training data, using the state of the law.

0% of this lesson · 10 items · 1h 33m total · 1h 18m without the optional journal

Contents of this lesson10 items
  1. ReadingLearning from Books or Taking Them: The Question Courts Are Now Answering3 min
  2. ReadingThe Training-Data Dispute: What Creators Object To and What Developers Argue4 min
  3. ReadingFair Use in US Courts: The Bartz, Kadrey, and Ross Rulings4 min
  4. ReadingThe Copyright Office Report and the UK Getty Ruling: Case-by-Case Analysis and Territorial Limits4 min
  5. ReadingLicensing, Settlements, and Effects on Creative Work4 min
  6. Guided ReadingGuided Close Reading: Judge Chhabria on What the Kadrey Ruling Does Not Decide7 min
  7. Guided ConversationArgue for the Author and for the Developer12 min
  8. Journal · optionalWhere You'd Draw the Line15 min
  9. Knowledge CheckCopyright, training data, and creative work10 min
  10. Graded QuizInformation, Fairness, and Creative Rights30 min

Reading 3 min

Learning from Books or Taking Them: The Question Courts Are Now Answering

This content reflects the field as of October 2026.

Authors and AI developers describe the same act in different words. To build a chatbot, a company runs a model through an enormous quantity of text, and that text includes books still under copyright. Authors' groups call this taking, and developers call it learning.

The Authors Guild, a professional organization for writers in the United States, says generative AI tools were "built illegally on vast amounts of copyrighted works without licenses," leaving authors without "any compensation or control over the use of their works." It wants consent and payment, and it says that "licensing is the path forward" (Authors Guild n.d.). The Guild speaks for authors, who would be paid if its view prevails.

OpenAI set out a developer's view in a March 2025 submission to the White House Office of Science and Technology Policy. It said its training uses "existing works to create something wholly new and different without eroding the commercial value of those existing works." It argued that this is protected by fair use, a doctrine in US law that permits some uses of a copyrighted work without the owner's permission (OpenAI 2025b). OpenAI's business depends on training data, and it would bear the cost if licenses were required. Other developers, including Anthropic and Meta, have made the fair use argument in court.

For several years the argument ran without a court decision. In 2025 and 2026 US courts began to rule, and the rulings don't all point one way.

  • In June 2025 a federal judge in California ruled that training Anthropic's models on books was a fair use, and that keeping a library of pirated copies was not (Bartz v. Anthropic 2025).
  • Two days later another judge in the same court ruled for Meta on the record before him. He wrote that his ruling "does not stand for the proposition" that Meta's training is lawful (Kadrey v. Meta 2025).
  • In February 2025 a federal judge in Delaware ruled that Ross Intelligence's copying of a legal publisher's case summaries, to build a competing research tool, was not fair use (Thomson Reuters v. Ross 2025).

Each ruling decided one dispute on its own evidence. Read together, they settle less than either side's public statements suggest.

That makes it worth separating two kinds of statement. "A court ruled that this company's use of these books was fair use" is a settled point about one case. "AI training is legal" and "AI training is theft" are claims about an open question. The Supreme Court has not ruled, other cases are still in progress, and courts in other countries apply different laws.

This is a description of public events as of October 2026. It isn't legal advice.

References

  • Authors Guild. n.d. "Artificial Intelligence." Accessed October 3, 2026.
  • Bartz v. Anthropic PBC. 2025. Order on Fair Use, No. C 24-05417 WHA, US District Court, Northern District of California, June 23, 2025.
  • Kadrey v. Meta Platforms, Inc. 2025. No. 23-cv-03417-VC, US District Court, Northern District of California, June 25, 2025.
  • OpenAI. 2025b. Response to the Office of Science and Technology Policy request for information on an AI Action Plan, March 13, 2025.
  • Thomson Reuters Enterprise Centre GmbH v. Ross Intelligence Inc. 2025. No. 1:20-cv-613-SB, US District Court, District of Delaware, February 11, 2025.

Report an issue with this item

Reading 4 min

The Training-Data Dispute: What Creators Object To and What Developers Argue

Introduction

Writers, artists, and news organizations say AI companies used their work without asking. The companies say they did nothing the law forbids. Underneath the accusations is one question, and each side has a worked-out answer to it.

This reading states what training does with a work, the creators' position, a developer's position, the legal test both sides appeal to, and what each side says is at stake.

What Training Does with a Work

Training data is the collection of examples a system learns from. To train a language model, a developer makes copies of a very large body of text and runs the model through it, adjusting the model's internal settings so that it gets better at predicting what comes next. The finished model holds those settings and not the text, although a model can sometimes reproduce passages it encountered many times.

Copyright is the legal right of a work's creator to control who may copy, distribute, or adapt it. Copies are made during training. The dispute is over whether making them requires permission and payment.

The Creators' Position

The Authors Guild is a professional organization for writers in the United States. It describes generative AI as "built illegally on vast amounts of copyrighted works without licenses," which it glosses as "without giving authors any compensation or control over the use of their works" (Authors Guild n.d.).

Its position has three parts.

  • Consent. An author should be able to decide whether a book is used for training.
  • Compensation. If it is used, the author should be paid.
  • Licensing as the remedy. A license is permission from a copyright owner to use a work, usually in return for payment. The Guild says, "We believe that licensing is the path forward to converting unlicensed uses to licensed uses."

The Guild also argues that the harm continues after training. It says these tools are used to produce works "that compete with and displace human-authored books."

A Developer's Position

OpenAI stated a developer's position in a submission to the US government dated March 13, 2025. Other developers have made similar arguments in court.

Fair use is a doctrine in US law that permits some uses of a copyrighted work without the owner's permission. OpenAI argues that training falls within it. In its words, training uses "existing works to create something wholly new and different without eroding the commercial value of those existing works" (OpenAI 2025b, sec. 3).

Two ideas carry that sentence. The first is transformative use: a use that gives a copied work a new purpose or character instead of substituting for the original. On this view a model that has learned patterns from a novel does a different job from the novel. The second is the absence of market harm: the company says its models "are trained to not replicate works for consumption by the public."

The Four Factors

US courts decide fair use by weighing four factors set out in the copyright statute. In plain words:

FactorThe question it asks
Purpose and character of the useIs the use commercial? Does it transform the work into something with a new purpose?
Nature of the copyrighted workIs the work highly creative, like a novel, or mainly factual?
Amount usedHow much of the work was copied, relative to the whole?
Effect on the marketDoes the use harm sales of the work or the owner's ability to license it?

No single factor decides a case. Developers put most of their weight on the first: training is transformative. Creators put most of theirs on the fourth: the output of these systems competes with their work, and unlicensed training removes a licensing market they could have had. The second and third factors tend to favor creators, since whole novels are creative works copied in full, and they usually count for less.

Two Interests and Two Sets of Stakes

Copyright law tries to serve two interests at once. One is the incentive to create: people write and paint partly because they can earn from it. The other is the ability to build on existing work, which includes building new technology.

Each side says the other's rule would sacrifice one of these.

  • Creators' stakes. The Guild warns that the "market dilution caused by AI-generated works will ultimately result in a shrinking of the profession." This is a forecast, made by an organization whose members' incomes are involved.
  • Developers' stakes. OpenAI says that if developers in China have unrestricted access to data and American companies lack fair use access, "the race for AI is effectively over" (OpenAI 2025b, sec. 3). This is also a forecast, made by a company whose costs would rise under a licensing rule.

Both forecasts could be partly right. Neither has been tested, because no country has yet run either rule for long.

Conclusion

The dispute is over whether using copyrighted works to train a model requires permission and payment. Creators say it does and propose licensing. Developers say training is a transformative fair use that doesn't harm the market for the originals. In the United States the answer turns on a four-factor test that courts apply to the facts of each case.

Key Terms

  • Training data: The collection of examples a system learns from.
  • Copyright: The legal right of a work's creator to control who may copy, distribute, or adapt it.
  • License: Permission from a copyright owner to use a work, usually in return for payment.
  • Fair use: A doctrine in US law that permits some uses of a copyrighted work without the owner's permission.
  • Transformative use: A use that gives a copied work a new purpose or character instead of substituting for the original.

References

  • Authors Guild. n.d. "Artificial Intelligence." Accessed October 3, 2026.
  • OpenAI. 2025b. Response to the Office of Science and Technology Policy request for information on an AI Action Plan, March 13, 2025.

Report an issue with this item

Reading 4 min

Fair Use in US Courts: The Bartz, Kadrey, and Ross Rulings

This content reflects the field as of October 2026.

Introduction

Headlines about AI and copyright tend to announce that courts have sided with AI companies, or that courts have sided with authors. Three US rulings are usually behind those headlines, and each is narrower than its headline.

This reading describes the three rulings, what each turned on, and how much weight they carry. It reports court decisions as of October 2026 and isn't legal advice.

The Terms

A fair use ruling is a court's decision on whether a particular use of a copyrighted work was a fair use, meaning a use US law permits without the owner's permission. Two ideas recur in these rulings. A use is transformative when it gives a copied work a new purpose or character instead of substituting for the original. Market harm is damage that a use does to the sales or licensing value of the copied work.

Three Rulings

Bartz v. Anthropic. Authors sued Anthropic over the use of their books to train its language models. On June 23, 2025, Judge William Alsup of the federal district court in Northern California divided the company's conduct into parts (Bartz v. Anthropic 2025).

  • Training. Using the books to train the models "was exceedingly transformative" and was a fair use.
  • Scanning purchased books. For print books the company had bought and scanned, "the mere format change was a fair use."
  • Pirated copies. A pirated copy is a copy of a work obtained from an unauthorized source without paying for it. The order states that Anthropic "pirated over seven million copies of books" and "had no entitlement to use pirated copies for its central library." The judge ordered a trial on those copies and on damages.

Kadrey v. Meta. Thirteen authors sued Meta for downloading their books from online "shadow libraries," sites that host pirated copies, and using them to train its language models. On June 25, 2025, Judge Vince Chhabria of the same court granted judgment to Meta on the training claim (Kadrey v. Meta 2025).

His reasons were about the evidence. The authors "presented no meaningful evidence on market dilution," the argument he considered their strongest. He wrote that the ruling "does not stand for the proposition that Meta's use of copyrighted materials to train its language models is lawful," and that "in many circumstances it will be illegal to copy copyright-protected works to train generative AI models without permission." A separate claim, that Meta distributed the books while downloading them, was left for later proceedings.

The two judges disagreed in print. Judge Alsup had compared the authors' complaint about competing works to a complaint about "training schoolchildren to write well." Judge Chhabria called that analogy "inapt."

Thomson Reuters v. Ross. Thomson Reuters publishes Westlaw, a legal research service that includes editors' summaries of court decisions. Ross Intelligence, a legal technology company, used material derived from those summaries to train a competing legal search tool.

On February 11, 2025, Judge Stephanos Bibas, sitting in the federal district court in Delaware, rejected Ross's fair use defense. He found the use commercial and not transformative, and found that Ross "meant to compete with Westlaw by developing a market substitute." He added that "only non-generative AI is before me today" (Thomson Reuters v. Ross 2025).

On September 29, 2026, the US Court of Appeals for the Third Circuit affirmed: "we hold that ROSS's use was not fair." It noted that, unlike the models in Bartz, Ross's platform "cannot generate original expression" (Thomson Reuters v. Ross 2026).

The Rulings Compared

CaseSystemRulingWhat it turned on
Bartz v. Anthropic (2025)Language modelsTraining was fair use; keeping pirated copies was notThe transformative purpose of training; how the copies were obtained
Kadrey v. Meta (2025)Language modelsFair use on that recordThe authors' lack of evidence of market harm
Thomson Reuters v. Ross (2025, affirmed 2026)Non-generative legal search toolNot fair useA competing product with the same purpose as the original

A precedent is an earlier court decision that later courts follow or treat as guidance. The Bartz and Kadrey rulings come from trial courts. They bind the parties and may persuade other judges, and no other court has to follow them. Kadrey affects only its thirteen authors.

The Ross appeal ruling binds lower courts within the Third Circuit, and it concerns a tool that doesn't generate text. The Supreme Court has not ruled on AI training. Other cases were still in progress as of October 2026, among them a group of suits against OpenAI in New York that the appeals court's opinion refers to.

Conclusion

As of October 2026, US trial courts have found training on books to be fair use in two cases involving generative models, one of them with an explicit warning that the result depended on weak evidence from the authors. A trial court and an appeals court have found no fair use where a non-generative tool competed with the works it copied. The rulings turn on their records, and the question of harm to creators' markets remains open.

Key Terms

  • Fair use ruling: A court's decision on whether a particular use of a copyrighted work was a fair use.
  • Transformative: Giving a copied work a new purpose or character instead of substituting for the original.
  • Market harm: Damage that a use does to the sales or licensing value of the copied work.
  • Pirated copy: A copy of a work obtained from an unauthorized source without paying for it.
  • Precedent: An earlier court decision that later courts follow or treat as guidance.

References

  • Bartz v. Anthropic PBC. 2025. Order on Fair Use, No. C 24-05417 WHA, US District Court, Northern District of California, June 23, 2025.
  • Kadrey v. Meta Platforms, Inc. 2025. No. 23-cv-03417-VC, US District Court, Northern District of California, June 25, 2025.
  • Thomson Reuters Enterprise Centre GmbH v. Ross Intelligence Inc. 2025. No. 1:20-cv-613-SB, US District Court, District of Delaware, February 11, 2025.
  • Thomson Reuters Enterprise Centre GmbH v. Ross Intelligence Inc. 2026. No. 25-2153, US Court of Appeals for the Third Circuit, September 29, 2026.

Report an issue with this item

Reading 4 min

The Copyright Office Report and the UK Getty Ruling: Case-by-Case Analysis and Territorial Limits

This content reflects the field as of October 2026.

Introduction

People on both sides of the AI copyright dispute would like a simple rule: training is always allowed, or training always needs permission. The main official analysis in the United States declined to give one, and the first major UK ruling decided a much narrower question than people expected.

This reading describes the US Copyright Office's 2025 report and the UK High Court's 2025 decision in a case brought by Getty Images, and what the two show about why outcomes differ. It reports public documents as of October 2026 and isn't legal advice.

The Copyright Office Report

The US Copyright Office is the federal agency that registers copyrights and advises Congress on copyright law. In May 2025 it released the third part of a study of copyright and AI, on the training of generative models (US Copyright Office 2025). The report is the Office's analysis and doesn't bind any court.

Its central conclusion is that no blanket answer is available. Fair use, the US doctrine that permits some uses of a copyrighted work without permission, has to be judged by case-by-case analysis: deciding each dispute on its own facts instead of applying one rule to all. On the report's approach, some uses of copyrighted works in training will qualify as fair use and others won't, depending on such things as what works were used, how they were obtained, what the model is for, and what its outputs do to the market.

On that last point the report identifies three kinds of harm a court can consider.

Kind of harmWhat it means
Lost salesA model's output substitutes for the original work, so fewer copies are bought
Lost licensingThe owner loses fees that could have been charged for permission to train on the work
Market dilutionLarge volumes of AI-generated works compete with human-made works of the same kind

Market dilution is the weakening of a market for human-made works when large volumes of AI-generated works of the same kind compete with them. Treating it as a harm under copyright law is a newer idea and is disputed.

On remedies, the report favors letting voluntary licensing between rights holders and developers continue to develop, and it does not recommend new legislation for now.

The Report's Status

The document released in May 2025 is labeled a pre-publication version. Its cover states: "A final version will be published in the near future, without any substantive changes expected in the analysis or conclusions" (US Copyright Office 2025).

Whether a final version has been issued is unsettled as of October 2026. The pre-publication version is the one cited here, and the Office's website is the place to check its current status.

Getty Images v. Stability AI

Getty Images licenses stock photographs. Stability AI makes Stable Diffusion, a model that generates images from text. Getty sued Stability in the High Court of England and Wales over the use of its photographs to train the model. Mrs Justice Joanna Smith gave judgment on November 4, 2025 (Getty Images v. Stability AI 2025).

The question most observers were waiting for was never decided. Getty had dropped its claim about training. The judgment records Getty's acknowledgment that "there is no evidence that the training and development of Stable Diffusion took place in the United Kingdom" (para. 9). UK copyright law covers acts done in the UK, so copying done elsewhere fell outside the claim. Getty also dropped a claim about the model's outputs.

What remained was a narrower claim. Getty argued that the model itself was an infringing copy, which in UK law is an article whose making infringed the copyright in a work, and that bringing it into the UK was therefore unlawful. The judge rejected this. Her reasoning was that the model doesn't store or contain copies of the photographs, so it couldn't be an infringing copy of them. Getty succeeded only on a limited trademark point about its watermark appearing in some generated images.

This is a trial-level ruling on one claim and is open to appeal. It doesn't say whether training on copyrighted images is lawful in the UK.

Why Outcomes Differ

Jurisdiction is the territory whose courts and laws apply to a dispute. Set beside US practice, the Getty case shows three things that shape an outcome.

  • The country. The United States has a flexible fair use doctrine. The United Kingdom has no equivalent general defense.
  • The claim. "The model was trained on my work" and "the model is a copy of my work" are different claims. Getty lost the second without the first being tested.
  • Where training took place. Copyright is territorial. A developer that trains in one country and offers the model in another may face different rules for each act.

Conclusion

As of October 2026, the US Copyright Office's analysis, available in a pre-publication version, rejects a blanket rule in either direction and names lost sales, lost licensing, and market dilution as harms courts can weigh. The UK High Court held that an image model was not itself an infringing copy and did not rule on training. Results in this area depend on the country, the claim, and where the training was done.

Key Terms

  • Case-by-case analysis: Deciding each dispute on its own facts instead of applying one rule to all.
  • Market dilution: The weakening of a market for human-made works when large volumes of AI-generated works of the same kind compete with them.
  • Infringing copy: In UK law, an article whose making infringed the copyright in a work.
  • Jurisdiction: The territory whose courts and laws apply to a dispute.

References

  • Getty Images (US) Inc. v. Stability AI Ltd. 2025. [2025] EWHC 2863 (Ch). High Court of England and Wales, November 4, 2025.
  • US Copyright Office. 2025. Copyright and Artificial Intelligence, Part 3: Generative AI Training. Pre-publication version. Washington, DC, May 2025.

Report an issue with this item

Reading 4 min

Licensing, Settlements, and Effects on Creative Work

This content reflects the field as of October 2026.

Introduction

While courts argue over whether AI training needs permission, money has started to change hands. One developer has agreed to pay authors a very large sum, and rights holders and developers have been signing agreements for access to works.

This reading covers that settlement, licensing, the concern about flooded markets, and what is known about creators' incomes. It isn't legal advice.

The Bartz Settlement

A class action is a lawsuit in which a few named people sue on behalf of a large group with the same claim. Bartz v. Anthropic became one, on behalf of authors whose books the company had downloaded from two pirate libraries.

In June 2025 the judge ruled that training on books was a fair use, that keeping pirated copies was not, and that the pirated copies would go to trial (Bartz v. Anthropic 2025). The parties then settled. A settlement is an agreement that ends a lawsuit without a final court ruling on who was right.

The Authors Guild, which represents writers, reported its terms when the court gave final approval on July 20, 2026 (Authors Guild 2026).

TermWhat the Guild reports
Amount1.5 billion dollars, about 3,000 dollars per work
What it coversThe company's past acquisition and copying of the listed works, through August 25, 2025
What the company must doDestroy the files it took from the two pirate libraries
What it leaves openClaims based on what the models produce, and any claims about future conduct

The settlement is an agreement and sets no precedent, so it doesn't establish what the law requires of anyone else. And it followed the part of the ruling the company had lost, on pirated copies. The settlement resolved those claims and did not address the ruling on training.

Suits against other developers continue. The Authors Guild lists its own class action against OpenAI, filed in 2023, with Microsoft added as a defendant, and a suit by authors against Meta (Authors Guild n.d.).

Licensing

Licensing is the practice of granting permission to use copyrighted works in return for payment. The US Copyright Office's 2025 report describes licensing for AI training as already under way in some sectors and favors letting it develop voluntarily (US Copyright Office 2025).

In July 2023 OpenAI agreed to license part of the Associated Press's archive of news stories (O'Brien 2023). In November 2025 Warner Music Group settled its copyright suit against the AI music company Suno, which said it would launch licensed models (Malik 2025). Neither report states what was paid. The two sides read such deals differently.

  • Creators' reading. Licensing shows that a market exists and that developers can pay. The Guild calls it "the path forward" (Authors Guild n.d.). Judge Vince Chhabria, ruling in a suit by authors against Meta, wrote that if copyrighted works are as necessary as the companies say, "they will figure out a way to compensate copyright holders" (Kadrey v. Meta 2025).
  • Developers' reading. Deals with large publishers are possible. Obtaining permission from every author of every work is a different matter, and developers have argued that a rule requiring it would stall the technology. Judge Chhabria called that suggestion "ridiculous" in the same order. Whether licensing can work at that scale hasn't been tested.

Market Dilution

Market dilution is the weakening of a market for human-made works when large volumes of AI-generated works of the same kind compete with them.

Judge Chhabria described it: generative AI "has the potential to flood the market with endless amounts of images, songs, articles, books, and more." The Copyright Office's report names dilution as a harm courts may weigh (US Copyright Office 2025). The Authors Guild warns that it "will ultimately result in a shrinking of the profession" (Authors Guild n.d.).

Judge William Alsup rejected the concern in Bartz. He compared the authors' complaint to objecting that "training schoolchildren to write well would result in an explosion of competing works" (Bartz v. Anthropic 2025). No ruling cited here found market dilution proved.

What Is Known and What Isn't

Creative labor is paid work that consists of making original writing, images, music, or other creative works. Claims about AI's effect on it run ahead of the evidence.

Status as of October 2026
One developer has agreed to pay 1.5 billion dollars to authorsDocumented
Licensing agreements between rights holders and developers existDocumented
The settlement resolves future useIt doesn't; claims about future conduct were kept
AI-generated works have reduced creators' incomesNot established by the sources cited here
Creators' incomes will fall, or will be sustained by licensingForecasts

Long-run effects on how many people can earn a living from creative work haven't been measured.

Conclusion

A market in settlements and licenses is forming beside the litigation. The Bartz settlement paid authors for past copying of pirated books and left future use and model outputs unresolved. Creators hold that dilution of their markets is the main threat. One judge and the Copyright Office have taken that concern seriously, another judge has rejected it, and the effect on creators' incomes over time is unknown.

Key Terms

  • Class action: A lawsuit in which a few named people sue on behalf of a large group with the same claim.
  • Settlement: An agreement that ends a lawsuit without a final court ruling on who was right.
  • Licensing: The practice of granting permission to use copyrighted works in return for payment.
  • Market dilution: The weakening of a market for human-made works when large volumes of AI-generated works of the same kind compete with them.
  • Creative labor: Paid work that consists of making original writing, images, music, or other creative works.

References

  • Authors Guild. 2026. "Court Grants Final Approval of $1.5 Billion Anthropic Copyright Settlement." July 21, 2026.
  • Authors Guild. n.d. "Artificial Intelligence." Accessed October 3, 2026.
  • Bartz v. Anthropic PBC. 2025. Order on Fair Use, No. C 24-05417 WHA, US District Court, Northern District of California, June 23, 2025.
  • Kadrey v. Meta Platforms, Inc. 2025. No. 23-cv-03417-VC, US District Court, Northern District of California, June 25, 2025.
  • Malik, Aisha. 2025. "Warner Music Signs Deal with AI Music Startup Suno, Settles Lawsuit." TechCrunch, November 25, 2025.
  • O'Brien, Matt. 2023. "ChatGPT Creator OpenAI Signs Deal with AP to License News Stories." Associated Press, July 13, 2023.
  • US Copyright Office. 2025. Copyright and Artificial Intelligence, Part 3: Generative AI Training. Pre-publication version. Washington, DC, May 2025.

Report an issue with this item

Guided Reading 7 min

Guided Close Reading: Judge Chhabria on What the Kadrey Ruling Does Not Decide

Introduction

In June 2025 a US judge ruled for Meta in a copyright suit brought by authors, and many reports summarized the result as a court approving AI training. The judge had written, in the order itself, that it meant no such thing.

This reading goes through the passage where he says what his ruling does and doesn't decide, to show how a court can find for one side and warn it in the same document.

Locating the Passage

The document is the order of June 25, 2025, in Kadrey v. Meta Platforms, Inc., by Judge Vince Chhabria of the US District Court for the Northern District of California. Thirteen authors had sued Meta for downloading their books from online "shadow libraries," sites that host pirated copies, and using them to train its language models. The order is free on FindLaw, and the link is in the References.

The passage is the last three paragraphs of the order's opening section, which comes before the detailed background and analysis. This reading also draws on the order's concluding section. The order has no numbered paragraphs, so the reading identifies each quotation by where it sits. All quotations are from the order (Kadrey v. Meta 2025).

Fair use is the US doctrine that permits some uses of a copyrighted work without the owner's permission. Courts weigh four factors, the fourth being the effect of the use on the market for the work. This reading describes a court document. It isn't legal advice.

Walking Through the Passage

Step 1: Read what the judge says the ruling doesn't stand for

The final paragraph of the opening section says: "this ruling does not stand for the proposition that Meta's use of copyrighted materials to train its language models is lawful."

"Stand for the proposition" is lawyers' language for the rule that a decision establishes. The judge is telling future readers which rule his decision does not establish. It doesn't establish that Meta's training was lawful.

The sentence before it narrows the ruling further: "This is not a class action, so the ruling only affects the rights of these thirteen authors." Everyone else whose work Meta used keeps whatever claims they have.

Step 2: Read what he says it does stand for

The next sentence is: "It stands only for the proposition that these plaintiffs made the wrong arguments and failed to develop a record in support of the right one."

A "record" is the body of evidence the parties put before the court. The sentence makes two criticisms of the authors' case. They argued the wrong things, and they brought no evidence for the argument that might have worked.

The paragraph before names the wrong arguments. The authors said the model could reproduce small snippets of their books. They also said that unlicensed training cost them the chance to license their books for training. The judge writes that "both of these arguments are clear losers." On the first, the model couldn't produce enough of their text to matter. On the second, he held that authors aren't entitled to the market for licensing their works as training data.

Step 3: Find his description of market dilution

The same paragraph names the right argument. He calls it "the potentially winning argument": that Meta copied the books to create "a product that will likely flood the market with similar works, causing market dilution."

Market dilution means that a market for human-made works is weakened when large volumes of AI-generated works of the same kind compete with them. Earlier in the opening section the judge explains why he takes it seriously. Generative AI, he writes, "has the potential to flood the market with endless amounts of images, songs, articles, books, and more," made with "a tiny fraction of the time and creativity that would otherwise be required."

On this argument the authors offered almost nothing. In his words, they "barely give this issue lip service, and they present no evidence about how the current or expected outputs from Meta's models would dilute the market for their own works."

Step 4: Ask why a judge would rule for one side and warn it in the same order

The answer is in the paragraph that begins the turn to this case: "Courts can't decide cases based on general understandings. They must decide cases based on the evidence presented by the parties."

The judge holds a general view, and he states it. Just before that turn he writes that "in many circumstances it will be illegal to copy copyright-protected works to train generative AI models without permission." He also holds that a court may act only on what the parties prove. Meta brought evidence on market effects and the authors didn't, so on this record Meta had to win.

The warning serves two audiences. It tells other AI developers not to treat the result as clearance. And it tells future plaintiffs which argument to build. A judge who thought the law favored developers in general would have had no reason to write it.

Step 5: Say what a future plaintiff would need to show

The order's concluding section spells this out. Because Meta's use was highly transformative, the authors "needed to win decisively on the fourth factor." The judge adds that if they had "presented any evidence that a jury could use to find in their favor on the issue, factor four would have needed to go to a jury."

So a future plaintiff would need evidence that the outputs of the defendant's model compete with the plaintiff's own kind of work and reduce its market. The order suggests where that is more and less likely. News articles and typical genre fiction may be vulnerable to AI-generated substitutes. A memoir, which people read because of who wrote it, may not be.

He also predicts how such cases will go: "it seems like the plaintiffs will often win, at least where those cases have better-developed records on the market effects of the defendant's use." That is a forecast by one trial judge, and it binds nobody.

Key Considerations

The order decides motions for summary judgment. That is a ruling made without a trial, on the ground that the evidence each side has presented leaves nothing for a jury to decide. It tests the evidence in the file. A different file could produce a different result under the same law.

The common mistake is to report the case as "court says AI training is legal." The order says the opposite about its own meaning. What it decides is that thirteen authors didn't prove their case against one company.

The reverse mistake is to report it as a finding that training is illegal. The judge's general statements about what "will likely" be unlawful are his reasoning, and they weren't needed to decide the case. They don't bind other courts. Two days earlier another judge of the same court, in a case against Anthropic, had treated the market-dilution concern as outside what copyright protects, and Judge Chhabria's order criticizes that reasoning by name.

The order also left part of the case undecided. A separate claim, that Meta distributed the authors' books in the course of downloading them, wasn't covered by the motions and remained live. A trial-court ruling of this kind can also be appealed.

Summary

The order grants judgment to Meta and, in the same passage, limits what that judgment means. Stated plainly:

The holding. On the evidence these thirteen authors presented, Meta was entitled to judgment that its use of their books to train its language models was a fair use. The authors relied on two arguments the judge rejected and offered no meaningful evidence on market dilution, the argument he considered potentially winning.

The warning. The ruling does not establish that such training is lawful, and a plaintiff who proves that a model's outputs dilute the market for their work could win.

References

  • Kadrey v. Meta Platforms, Inc. 2025. No. 23-cv-03417-VC, US District Court, Northern District of California, June 25, 2025.

Report an issue with this item

Guided Conversation 12 min

Argue for the Author and for the Developer

In this conversation you'll hear the creators' case on AI training argued the way an authors' organization states it, then the developers' case argued the way a developer has stated it to government. You'll leave with a statement of which fact in the dispute matters most to you, and why.

You'll have this conversation with an AI assistant, using your own account. Choose a button to open a new chat with the prompt already filled in, then press send to start. If the chat opens empty, copy the prompt and paste it in.

Run this conversation in whichever assistant you already use:

Claude desktop app

To use another LLM, simply copy and paste the prompt into its chat window.

Show the full prompt (it lists misreadings to watch for, so skip it if you would rather come to the conversation fresh)
Guided Conversation: Argue for the Author and for the Developer (about 12 minutes)

Note to the learner: press send to start. Everything below is facilitator guidance for the AI. It lists misconceptions to watch for, so skip it if you'd rather come to the conversation fresh.

Please facilitate a role play with me. I'm an adult with no technical background who has used AI chatbots for everyday tasks, and I'm studying the dispute over training AI models on copyrighted work. Follow this guidance for the whole conversation.

GOAL
I can compare the positions of creators and AI developers on training data, using the state of the law.

ROLE
You will play two advocates in turn, then step out of role.
- First, an advocate for creators, arguing as the Authors Guild states its position: AI tools were built on copyrighted works without licenses; authors want consent and compensation; licensing is the way forward; AI-generated works will dilute the market for human work. Full strength, first person, never a caricature.
- Second, an advocate for developers, arguing as OpenAI stated its position to the US government in 2025: training is a transformative fair use that creates something new without eroding the commercial value of existing works, and restricting it would hand the lead in AI to other countries. Equal force, equal length.
- Announce each change clearly: "In role as an advocate for...", and later "Stepping out of role."
- What makes this hard: I will push back, and you must answer in role without drifting to a middle position or making one advocate weaker.
- In role, say what your side has at stake financially.

HOW TO RUN THE CONVERSATION
- Ask one question at a time, then wait for my reply. Keep each of your turns under about 120 words.
- Be curious and collegial. Use plain words and define any legal term briefly on first use.
- Plain conversation only: don't search the web or create files or documents.
- Aim for about 12 minutes: four on each advocate and four out of role. If my replies are brief, offer one concrete prompt, such as "Does it matter to you whether the books were bought or pirated?" and move on. If I seem uncertain, shorten to 5-7 minutes. Always reach the final topic.
- Start now. Open with one or two warm sentences: this is a conversation, not a quiz; I don't have to be persuaded; I can ask you to clarify anything. Then ask whether I create anything myself or mainly use AI tools, and begin topic 1.

TOPICS, IN ORDER
1. Creators' advocate. In role, argue the creators' position. Ask for my strongest objection and answer it in role.
2. Developers' advocate. Change roles and argue the developers' position the same way. Ask for my objection and answer it in role.
3. Out of role. Ask me which facts I think courts have treated as decisive. Draw out three: how the copies were obtained, whether the new product competes with the originals, and what evidence of market harm was presented.
4. Closing. Ask me which fact matters most to me and why. Tell me I can take it into a short optional journal entry.

KEY POINTS TO KEEP ACCURATE (as of October 2026)
- Rulings are mixed and turn on specific records. Bartz v. Anthropic (June 2025): training on books was fair use; keeping a library of pirated copies was not. The case later settled for 1.5 billion dollars, covering past conduct only.
- Kadrey v. Meta (June 2025): fair use on that record, because the authors gave no meaningful evidence of market harm. The judge wrote that the ruling does not establish that such training is lawful.
- Thomson Reuters v. Ross (2025, affirmed on appeal September 2026): copying legal summaries to build a competing, non-generative research tool was not fair use.
- Acquiring works by piracy was treated differently from training. Market harm, including dilution by AI-generated works, is the main open issue.
- These are trial rulings and one appeals ruling. The US Supreme Court has not ruled.
- Rulings differ by country. In Getty Images v. Stability AI (UK, 2025) the court held the model was not an infringing copy because it doesn't store the works. Getty had dropped its claim about training.
- You may not know developments after your training. Say so if I ask about recent events.

MISCONCEPTIONS TO CORRECT GENTLY
Correct these out of role, briefly, then return to the conversation.
- "Courts have said training is legal": some rulings, on some records.
- "Courts have said it's theft": one case, about a non-generative tool that competed with the works it copied.
- "The model contains copies of the books": a UK court found the model doesn't store them. Whether a model can reproduce some passages from memory is a separate question.

LIMITS
- No legal advice. If I ask about my own work or situation, tell me to consult the rules and professionals that apply to me.
- Outside the roles, give no view of your own on who should win.
- Don't favor or disparage any company, including the one that built you. If the company that built you is named in this conversation or is a party to anything discussed, say so once when it first comes up, then describe that company as you do every other and take no side.

TO FINISH
After my closing answer, close in one short turn, out of role:
- Affirm one specific thing I worked out, in my own words where possible.
- Suggest one or two next steps that fit how the conversation went. Possible steps: write a journal entry stating where I'd draw the line; read the opening pages of one of the rulings; check the current status of one case.
- Restate my answer on its own lines, labeled "The fact that matters most to me", so I can copy it.

Report an issue with this item

Journal 15 minOptional

Where You'd Draw the Line

Overview

You'll state your own rule for when AI developers may train on copyrighted work, then test it against the facts of two real court rulings. Testing a rule against cases shows whether it says what you meant.

The entry is optional. It's for you, and nobody collects it.

Writing Prompt

State where you'd draw the line on training AI with copyrighted work, and test your line against one ruling that cuts each way. Write 250–400 words.

Steps

  1. State your line in one sentence. Say when training on copyrighted work should be allowed without permission and when it shouldn't. For example, your line might depend on how the copies were obtained, on whether the result competes with the originals, or on whether the creator was paid.
  2. Apply it to the facts of a ruling that favored a developer. In Bartz v. Anthropic (2025), a US court held that training language models on books was a fair use, including books the company had bought in print and scanned. Say what your line would decide on those facts.
  3. Apply it to the facts of a ruling that favored a rights holder. In Thomson Reuters v. Ross (2025), a US court held that copying a legal publisher's case summaries to build a competing legal research tool was not a fair use. In Bartz, the same court that approved the training held that keeping a library of pirated books was not a fair use. Choose one and say what your line would decide.
  4. Say whether your line held and what you'd change. If your line gave an answer you don't accept in either case, revise it. If it held, say why you're satisfied with both results. You can draw on any notes of your own.

Self-Check

Before you finish, check that your entry:

  • States your line in one sentence
  • Uses the facts of two real rulings
  • Says honestly whether the line held
  • Revises the line or defends it

Nothing is uploaded. Write in your own notebook or document and keep it.

Report an issue with this item

Knowledge Check 10 min

Copyright, training data, and creative work

This ungraded knowledge check assesses your understanding of the dispute over training AI models on copyrighted work. You'll be asked about the creators' and developers' positions, the three US rulings, the Copyright Office report and the UK ruling, and licensing, settlements, and market dilution.

Note: Use this to test yourself, review the feedback on any questions you miss, and retry until you feel confident before moving forward.

5 questions · ungraded · retry as often as you like

Report an issue with this item

Graded Quiz 30 min

Information, Fairness, and Creative Rights

This graded quiz assesses your understanding of the evidence and arguments about AI's effects on information, fairness, and creative work. You'll be asked about documented harms from synthetic media, the evidence on elections, the liar's dividend, provenance and labeling, how bias enters AI systems, the COMPAS dispute, competing fairness definitions, the training-data dispute, the US court rulings, and settlements.

Note: Aim for a score of 80 percent or higher. If you score lower, use the feedback to review the topics you missed, then retake the quiz.

10 questions · target score 80% · 3 forms, rotated on each attempt

Report an issue with this item