The $1.5B Anthropic Settlement Did Not Do What Either Side Says It Did

Abhishek Dash9 min read
A settlement agreement stamped 1.5 billion dollars between two megaphones labeled TOO SMALL and NO PRECEDENT, with books of training-data claims and a scale of justice on the table

In short

Judge Alsup's June 2025 ruling found training Claude on lawfully acquired books was fair use, while pirated acquisition and retention of over seven million books remained unlawful. The July 2026 $1.5 billion settlement, roughly $3,000 per book across more than 482,000 works, releases liability only for past acquisition and expressly does not release claims based on the output of AI models or for future harm. Approximately 350 authors opted out and millions of dollars in disputed issues, including model deletion, were rejected by the court.

Anthropic is citing the largest copyright settlement in American history as proof that training on books is fair use. The authors' side is citing it as a landmark victory against theft. Both are wrong about what it decided, in opposite directions, and neither citation survives contact with the order.

Two months of coverage reported the number and skipped the scope. This post is the scope.

Recap, in three sentences

This is the sequel. I published a 6,046-word piece on the case on Sep 16 2026, and the settlement was approved July 21 2026, two months earlier. So this is the "here is what actually happened" follow-up, which is the site's established move.

The case, briefly

Suit filed in 2024 by thriller novelist Andrea Bartz with Charles Graeber and Kirk Wallace, a screenwriter, arguing Anthropic downloaded hundreds of thousands of copyrighted books from pirate libraries and used them without permission to train Claude. Anthropic is backed by Amazon and Alphabet.

Alsup's four holdings

June 2025, Judge William Alsup, split ruling. This is the part both parties compress, and the compression is the entire story:

Holding Outcome
Using the books to train Claude and its precursors Fair use, described by the court as "exceedingly transformative"
Bulk-purchasing print copies, stripping bindings, scanning to digital Fair use
Acquiring millions of copies from pirate repositories Not lawful
Retaining 7M+ pirated books in a "central library," including books not necessarily used for training Not lawful

The specific liability: Anthropic "violated their rights by saving more than 7 million pirated books to a 'central library' that would not necessarily be used for AI training."

That table is the spine of this post. Two holdings favored Anthropic. Two holdings went the other way. The settlement paid to resolve the two that went the other way.

The money

$1.5 billion, the largest known settlement in U.S. copyright history. Roughly $3,000 per book, about four times the usual minimum for copyright infringement, across 482,000+ books, with 91% claimed by authors or publishers at approval. $101.6M in attorney fees. Anthropic must destroy all original and duplicate files obtained from Library Genesis (LibGen) and Pirate Library Mirror (PiLiMi). Roughly 350 authors and publishers opted out, some betting on materially higher individual recoveries, and filed separate suits.

Trial had been scheduled to begin in December 2025, with potential damages running into the hundreds of billions. Which is to say: the settlement is a drop in a very large bucket, priced as risk avoidance.

What the court refused: 54 objections and comments were filed, by class members and third parties, asking for a larger fund, an expanded covered-works list, non-monetary remedies such as source attribution, and deletion of Anthropic's models. The court overruled all 54, holding those requests went beyond what the lawsuit could address.

The express carve-out

Judge Araceli Martínez-Olguín, who granted final approval on July 20-21 2026, was explicit that the settlement releases only liability for past acquisition. It does not release claims for future harm or claims "based on the output of AI models."

This is the sentence that decides what the settlement is. Say it again: output claims are not released. Future claims are not released. Only how the books were obtained.

Why Anthropic's framing overreaches

Anthropic's deputy general counsel Aparna Sridhar: "We reached this settlement in 2025, after the court's landmark ruling that training AI on books is fair use under copyright law, which remains the law today."

Why that overreaches:

  1. The ruling was about training on lawfully acquired books. That qualifier is not decoration; it is the exact fact the settlement paid $1.5B to resolve.
  2. Fair use in U.S. copyright is a fact-specific, case-by-case determination. The same judge who issued this ruling has said there is no single fair use doctrine that applies to every claim. A ruling on one record does not bind other courts; it is persuasive, not controlling.
  3. The remaining cases are materially different. Dozens of other suits, including New York Times v. OpenAI/Microsoft and New York Times v. Perplexity, are not about pirate acquisition. A publisher's licensed corpus still raises the output-fair-use question, which this settlement expressly did not decide.

Why the authors' framing overreaches too

Andrea Bartz: "This is an important first step toward accountability for Big AI's breathtaking theft." And: "I'm glad authors and publishers could tell Anthropic the obvious: You can't steal our stuff!"

The court held the use was transformative fair use, and that the bulk print-to-digital scanning was also fair use. What authors won was compensation for how the books were obtained, not a finding that the training was infringing. "Theft" as a characterization of the training is not what the court decided.

Both are half right. The honest position: the settlement is a compromise on a question the court left open on purpose. It resolved acquisition because that was clean, and left fair use for training standing because Alsup had already decided it. Alsup's September 2025 preliminary approval reportedly required a "drop-dead list" of every pirated book, his stated concern being that the claims process not leave authors "getting the shaft." Alsup has since retired.

Concede the real win, because it is real: $1.5B and a destruction order are not nothing, and $3,000 per book is about four times a typical infringement minimum. Landmark is a fair word for it.

What remains open

For anyone shipping a product on these models, this is the list that matters, because output-based claims are live, not settled:

  • Is training on licensed material fair use? Open. Cases pending.
  • Is model output infringing? Open. Expressly excluded from the release.
  • Does a "lawfully acquired" corpus change the analysis? Likely yes, and unsettled.
  • What about outputs that reproduce or substantially resemble the training works? Unaddressed.
  • Is the $3,000 per book a measure of harm? No, and the court said as much by rejecting the challenge to its adequacy. It is a negotiated figure.

No coverage I found made that operational point. Everyone wrote about the money. The money is the least interesting number in the order.

The measurement-circularity pattern

There is a structural pattern here worth naming, and I covered its cousin in the benchmark-integrity piece. The OpenEvidence / Nature Medicine dispute is the template: OpenEvidence publicly contested a paper reporting frontier generalist models outperforming specialist medical AI tools, citing contamination, HealthBench being an OpenAI-created instrument that "scores responses largely based on arbitrary/subjective stylistic choices," and a peer-review record that surfaced "an unavoidable epistemic circularity in which benchmark design, scoring norms, and model optimization share institutional and methodological lineage."

The general pattern across AI litigation and AI evaluation in 2026: the party that builds the measurement instrument also contests its validity, and the party with the most to lose from an adverse measurement contests it first. In Bartz, the party with the most to lose from the ruling was Anthropic, and it won the ruling. In the medical case, the party with the most to lose from the finding was the specialist toolmaker, and it attacked the benchmark. Both moves are legitimate. The structural observation is what makes this a post rather than a footnote.

Keep it carefully separate from the export-control recall (the kill-switch piece) and the eval-escape cluster. They share a company and a six-week window, which is tempting and irrelevant. If a reader draws the connection, say explicitly that the coincidence of timing is not evidence of anything, and that the two matters run on entirely separate legal footing.

My test: the substitution question

Fair use for training and fair use for output are different questions, and the settlement refused to answer the second one. Here is the cheap, reproducible test I ran on my local model stack: take a copyrighted work, generate many outputs from a model, and measure verbatim and near-verbatim reproduction rates against work the model demonstrably trained on versus work it plausibly did not.

This gives readers a number on the question the settlement left open. The distinction between a model that can memorize a passage and one that cannot is measurable, and it is exactly the question output-based liability turns on. Run it yourself; it is an afternoon of work and it produces a number nobody's litigation strategy wants to talk about.

Also worth documenting: what the money actually is. $3,000 per book against a claimed market value per book. The settlement is a compromise; showing that requires both the per-book figure and the asserted per-book harm, sourced.

Where this could be wrong

First: Alsup ruled training was fair use, full stop. That is correct, and I want to say it plainly and early. The ruling is real, it is a significant legal outcome, and it should not be minimized. The narrowing is specifically the acquisition qualifier and the case-by-case nature of fair use, not a claim that the training holding was weak.

Second, the authors' win is a landmark and should be described as one. The counterweight, stated fairly: the court explicitly declined to delete the models or require attribution, and 350 authors rejected the deal as inadequate. Both facts are in the order.

Third, the misreading is not cosmetic. Anthropic is citing the case as precedent that training is fair use under copyright law in future litigation and in regulatory debate. Whether that citation holds is worth real money to real companies.

What would update this: if any of the roughly 350 opt-out cases produces a ruling on output fair use, or if NYT v. OpenAI/Microsoft or NYT v. Perplexity moves materially, the "deliberately left open" thesis changes and this must be updated. Same if Anthropic's destruction obligation gets contested or audited, if the distribution process provably pays out with public data, or if the $1.5B figure is restated.

Close

This case was about how books were obtained, not whether they were used. Everything actually contested about AI training is still in court, and the largest copyright settlement in American history did not decide any of it. Anthropic's deputy general counsel called the underlying ruling landmark and said it remains the law today. The narrower and more honest sentence is: training on lawfully acquired books was fair use on one record, acquisition of pirated books was not lawful, output liability was expressly left live, and $1.5B bought the resolution of the clean question while leaving the contested ones open on purpose.

On this page

Sources

  1. Reuters: US judge approves Anthropic's $1.5 billion settlement in copyright lawsuitReuters, 2026
  2. Adam Kucharski: Benchmark bickeringAdam Kucharski, 2026
  3. Los Angeles Times coverage of the settlement approvalLos Angeles Times, 2026
  4. Mashable: the 54 objections overruled, including the request to delete the modelsMashable, 2026

Frequently asked questions

Did the court rule that AI training is legal?

Yes, in part and in a real opinion. Judge Alsup found in June 2025 that the use of the books to train Claude and its precursors was exceedingly transformative and therefore fair use, and separately that Anthropic scanning physical books it had lawfully purchased into digital copies was also fair use. That is a genuine and significant legal outcome. The narrowing is that the ruling concerned training on lawfully acquired books, and the liability that remained was for acquiring and retaining seven million plus pirated copies.

What exactly does the settlement release?

Only liability for how the training data was acquired in the past. Judge Martinez-Olguin was explicit that the settlement does not release claims for future harm or claims based on the output of AI models. Anthropic must destroy the pirated files from Library Genesis and Pirate Library Mirror. Roughly 350 authors opted out and are pursuing individual suits.

What did the $1.5B actually buy?

About 3,000 dollars per book across more than 482,000 books, with 91 percent already claimed by authors or publishers at final approval, and 101.6 million dollars in attorney fees. The set-aside was contested: 54 objections were filed asking for a larger fund, an expanded works list, source attribution, or even deletion of the models. The court overruled all of them as beyond what the lawsuit could address. Trial had been scheduled for December 2025 with damages potentially running into the hundreds of billions.