The $1.5bn Anthropic Settlement Is Final, and It Prices Piracy Rather Than Training

A San Francisco judge has signed off on the largest copyright recovery on record, but the deal settles where the books came from, not whether models may learn from them.

2 min read ·

US District Judge Araceli Martínez-Olguín gave final approval on Monday to Anthropic's $1.5bn settlement with authors and publishers, closing the class action that began in 2024 when writers including Andrea Bartz, Charles Graeber and Kirk Wallace Johnson accused the company of building a training library out of books pulled from piracy sites. The deal covers more than 482,000 works at roughly $3,000 apiece. The judge cut the lawyers' fee request from $187.5m to $101m and dismissed objections that the payout was too small, calling them unrealistic about what a trial would have risked.

Claims were already in before approval: Anthropic says more than 91 per cent of eligible rights holders have filed for their share. Payments now move to the settlement administrator. A minority of authors and publishers opted out and are pursuing their own suits, so this is not quite the end of Anthropic's exposure on this catalogue.

What the money actually buys

It is worth being precise about what was settled, because the headline figure invites the wrong reading. The case turned on an earlier ruling by Judge William Alsup, who drew a line that the rest of the industry has been studying ever since. Training a model on lawfully acquired books, he found, could qualify as fair use. Downloading and keeping pirated copies did not. Anthropic was never paying $1.5bn for the right to train; it was paying to make the second half of that ruling go away before a jury could put a per-work number on it.

That distinction matters to anyone building on top of these models. The settlement does not create a licensing market for training data, and it does not tell you that scraping is safe or unsafe. What it does establish is a reference price for provenance failures. Roughly $3,000 per work is now the number every general counsel will plug into a spreadsheet when they ask how their own data lake was assembled.

Why the rest of the docket cares

There are dozens of pending cases brought by authors and news organisations against AI developers, and this is the first big one in the US to settle rather than be decided. That cuts both ways. Plaintiffs now have a concrete benchmark and a demonstration that a class can be certified and paid at scale. Defendants have a template for containing the damage: concede nothing on training, pay for acquisition, and move on.

The weaker spot for defendants is the piracy fact pattern itself. Companies that can show they bought, licensed or lawfully scanned what they trained on are in a materially better position than those whose pipelines touched shadow libraries. Expect discovery in the remaining cases to focus less on model internals and more on download logs, torrent records and procurement emails. The boring paper trail is where the liability now lives.

What to watch

  • The opt-outs. Individual suits from authors and publishers who declined the deal will test whether courts value works above the class rate.
  • Copycat settlements. Other labs facing similar acquisition claims now know what the market clears at. A quiet wave of private deals would not be surprising.
  • Appellate fair use. Alsup's training-is-transformative reasoning is a district court view. Until an appeals court rules on the training question directly, nobody has the certainty the headline number seems to imply.

For builders, the practical takeaway is dull but real: keep records of where your data came from. The industry has just learned that the expensive question was never "did you train on it" but "how did you get it".


Sources

Covers tech policy, chips, and the business underneath the product launches.

Responses (2)

Sign in to leave a response.

  • Agree the price tag is for provenance, not training. The uncomfortable part is how many open datasets still carry the same shadow-library lineage nobody has audited.

  • Download logs as the new liability surface is exactly right. Most teams I talk to cannot reconstruct where a 2023 corpus came from, let alone prove it in discovery.

More from Marcus Tan

Recommended from Horizon