Fair use on trial

Lesson 2 of 5 in AI and Intellectual Property: Training Data, Outputs, and the Litigation Wave.

In the United States, everything turns on fair use — the flexible, judge-made-then-codified exception in §107 of the Copyright Act. It is not a checklist but a weighing of four statutory factors, and you need to argue each the way both sides do.

Factor 1: Purpose

The purpose and character of the use — is it transformative, serving a different purpose than the original, and is it commercial?

Developers’ story: training extracts statistical patterns to build something with an utterly different purpose — Authors Guild v. Google (2d Cir. 2015) blessed copying millions of books to build a search index, and Google v. Oracle (2021) blessed reimplementing code APIs. Plaintiffs’ counter, armed with Warhol v. Goldsmith (2023): transformativeness is judged use by use, and where the AI’s output serves the same market purpose as the original (news summaries competing with news; images competing with stock photos), the "different purpose" story collapses. Thomson Reuters v. Ross applied exactly that lens — a legal-research tool trained on headnotes to compete with Westlaw was not transformative.

Factor 2: Nature

The nature of the copyrighted work — creative works get thicker protection than factual ones.

Rarely decisive, but it separates the cases: novels and photographs (Bartz, Andersen, Getty) sit at the creative core; headnotes and news articles carry a factual component the defence can lean on. Courts in the 2025 AI rulings gave this factor little independent weight.

Factor 3: Amount

The amount and substantiality copied — and here every training case starts from the same awkward fact: developers copied entire works, at scale.

The defence answer comes from Google Books: copying the whole work can be fair when necessary to the transformative purpose and when the public never receives the full copy. That is why memorisation and regurgitation are strategically catastrophic — verbatim outputs convert "we only used it internally" into "the public gets the work back out", which is also why NYT’s hundred regurgitated excerpts and Getty’s watermarked outputs headline their complaints.

Factor 4: Market effect

The effect on the potential market — historically the heavyweight factor, and where the war will likely be decided.

Three distinct harm theories: substitution (outputs replace the original — Ross lost here), the licensing market (a market for training licences now exists — every deal signed with news publishers and stock agencies strengthens plaintiffs’ argument that unlicensed training destroys a real market), and market dilution (a flood of AI outputs devalues the whole category, even without copying any single work — the theory Judge Chhabria in Kadrey v. Meta called potentially decisive while ruling for Meta on the evidence actually presented).

Put the 2025 rulings side by side and the emerging pattern is sharper than "courts are split":

Thomson Reuters v. Rossfair use rejected. Non-generative system, trained on a rival’s editorial content, to build a direct substitute for that rival. Factors 1 and 4 aligned against the developer.

Bartz v. Anthropic — training on lawfully purchased books held fair use, in language as strong as any developer could hope for; but assembling and retaining a pirate library was not fair use, and that exposure priced at $1.5 billion in settlement. Acquisition legality became its own battlefield.

Kadrey v. Meta — Meta won, but the opinion reads like a plaintiff’s playbook: the court all but invited future claimants to build the market-dilution record these plaintiffs skipped.

The synthesis to carry into practice: fair use is winnable for training and losable at the edges — substitution-shaped products, pirated sources, and regurgitating outputs are the three edges. Every AI licensing deal, provenance requirement, and output-filter you meet in operations exists to stay off those edges.

Interactive sorting exercise: You are drafting the fair-use brief. Sort each fact by which side it strengthens.

Interactive checkpoint quiz (2 questions) — open this page in a browser to take it.