Resources › Current Affairs

Computing, AI & IT

Science & Technology

$3,000 a Book: The Largest Copyright Settlement in US History and What It Settles

$3,000 a Book: The Largest Copyright Settlement in US History and What It Settles

A court approved $1.5 billion for authors whose pirated books trained an AI model — but the ruling that produced it drew a line between acquisition and use

21 July 2026·Science & TechnologyComputing, AI & IT◆ High Yield·TechCrunch·7 min read

What happened

The temptation is to read this as a verdict on whether AI training infringes copyright. It is not — and the distinction the court actually drew, between how material was acquired and what was done with it afterwards, is both narrower and more useful. For an aspirant, this is the cleanest available illustration of how existing intellectual property doctrine is being applied to a technology that predates none of it.

What the court permitted and what it did not

How the books were obtainedCourt's treatmentReasoning
Purchased and scannedPermittedTraining treated as sufficiently transformative under fair use
Downloaded from pirate repositoriesNot permittedInfringement complete at unauthorised reproduction, whatever the later use
Settlement$1.5 billion~$3,000 per work across ~500,000 works; 91%+ of rights-holders claimed

Source: Rulings of the US District Court for the Northern District of California; TechCrunch, 20 July 2026

Smart Gravity Note

The distinction the court drew is between acquisition and use, and it is the whole of the case.

Training a model on lawfully obtained copies was treated as sufficiently transformative to fall within fair use, the American doctrine codified at Section 107 of the US Copyright Act, which weighs the purpose and character of the use, the nature of the work, the amount used and the effect on the market for the original.

Downloading the same books from pirate repositories was a separate act of reproduction that fair use did not excuse, because the infringement was complete at the moment of unauthorised copying, irrespective of what followed.

India's equivalent framework is not fair use but fair dealing under Section 52 of the Copyright Act, 1957, which is an enumerated list of permitted purposes rather than an open-ended balancing test — a structural difference that matters for how an Indian court would approach the same facts.

The court did not hold that training an AI on copyrighted books infringes copyright — it held that stealing the books in the first place does.

◎ In Simple Words

AI systems learn by reading enormous amounts of text, including books. A group of authors sued a company for using their books without permission. The judge decided that buying books and scanning them was acceptable, but downloading them from piracy websites was not. Rather than go to trial over the downloading, the company agreed to pay about $1.5 billion — roughly $3,000 for each of around 500,000 books — to the authors and publishers who own them.

16PYQs on this sub-topic →SCIENCE & TECHNOLOGY · Computing, AI & IT

Factual Pointers

Practice · 2 questions

1Practice Question

Which one of the following correctly distinguishes fair dealing under Indian law from fair use under American law?

2Practice Question

Consider the following statements regarding the settlement:

1. The court held that training an artificial intelligence model on copyrighted books is itself an infringement of copyright.

2. The infringement found concerned the downloading of books from pirate repositories rather than the subsequent use of the material for training.

3. The settlement amounted to roughly $3,000 per work across an estimated 500,000 works.

Which of the statements given above are correct?

Mains Practice Questions

1

A court has distinguished between how AI training data is acquired and what is done with it thereafter. Examine whether this distinction adequately protects the interests of creators.

2

India's fair dealing regime enumerates permitted purposes while the United States applies an open-ended fair use test. Discuss the implications of this structural difference for regulating artificial intelligence training data in India.

3

Should India adopt a text and data mining exception to copyright? Evaluate the competing interests of the domestic creative economy and domestic AI development.

Frequently Asked

· People also ask
What did the court actually decide in the Anthropic copyright case?

It distinguished between acquisition and use. Judge William Alsup treated training on books the company had purchased and scanned as permissible under fair use, but held that downloading the same books from pirate repositories was unlawful. The company settled rather than take the piracy claim to trial.

GS3 · IPR and technologyThis means the ruling does not establish that AI training on copyrighted material infringes copyright. It establishes that obtaining the material unlawfully does, leaving the larger licensing question substantially open.

SOURCE US District Court, Northern District of California; TechCrunch

How much is the settlement worth per book?

Roughly $3,000 per work, across an estimated 500,000 works, totalling $1.5 billion — reported as the largest copyright settlement in United States history. Counsel stated that over 91 per cent of covered authors and publishers had claimed their share.

GS3 · EconomyThe sum is distributed among whoever holds the rights, which in many cases is a publisher rather than the author. It functions as a penalty for unlawful acquisition rather than as remuneration proportionate to any book's influence on the model.

SOURCE TechCrunch, 20 July 2026

How does Indian fair dealing differ from American fair use?

Section 52 of India's Copyright Act, 1957 lists specific permitted purposes — private use, criticism or review, reporting current events, certain educational and research uses. Section 107 of the US Copyright Act instead applies an open four-factor test weighing purpose, nature of the work, amount used and market effect.

GS3 · IPRThe consequence is adaptive capacity: an American court can find a novel use transformative, while an Indian court applying a closed list has less room to accommodate a technology the drafters never contemplated.

SOURCE Copyright Act, 1957; US Copyright Act, Section 107

Does India have a text and data mining exception for AI training?

No. The European Union and Japan have legislated text and data mining exceptions permitting computational analysis of lawfully accessible material, the EU version subject to a rights-holder opt-out. India has enacted no equivalent, leaving the question to the enumerated exceptions in Section 52.

GS3 · Technology policyWhether to adopt such an exception is a live policy question, and the trade-off is concrete: a permissive regime transfers value out of India's creative economy, while a restrictive one raises the cost of building AI models domestically.

SOURCE EU Directive on Copyright in the Digital Single Market; Copyright Act, 1957

Does this settlement set a binding legal precedent?

No. Because the parties settled rather than litigating to final judgment, no appellate ruling was produced, and a district court's reasoning binds no other court. What the episode establishes is a commercial reality — that unlawfully acquired training data carries very large quantifiable risk.

GS2 · Judicial processThis is why the settlement changes industry procurement behaviour more reliably than it changes the law. The doctrinal question of whether training on lawfully acquired copyrighted material requires a licence remains unresolved.

SOURCE TechCrunch; Kluwer Copyright Blog analysis

Why do many authors not regard the settlement as a win?

Because it compensates for the manner of acquisition rather than for creative contribution. A flat sum per work takes no account of how much any book shaped the model, no method existing to measure that. Where publishers hold the rights, payment flows to them rather than to authors.

GS4 · Ethics · Consent and attributionThe deeper objection is that the reasoning implies a company with enough capital could lawfully buy and scan the same books and train on them — so the economic effect on the author is identical, and only the procurement route differs.

SOURCE TechCrunch, 20 July 2026