Remember the Napster years? A whole generation discovered that copying a song took two seconds, that the internet made it feel victimless, and that courts eventually disagreed. The music industry spent years untangling what was allowed, what was theft, and what was just a new way of listening. We landed somewhere reasonable, but only after a lot of lawsuits.
We are living through the book version of that story right now, except the copier is an AI model and the question is stranger than “did you download this.” The question is whether reading millions of books to learn patterns from them counts as copying at all.
The short answer, and why it frustrates everyone
Training an AI model on copyrighted books can be legal, according to recent court decisions, if two things are true: the works were acquired legally, and the use is transformative. Pirating books is still illegal, and no amount of clever technology changes that.
That split is the entire ballgame. Courts have been drawing a line not around what AI does with a book but around how the company got the book in the first place. It is a distinction that feels almost old-fashioned, and honestly, that is what makes it easy to explain to non-technical folks. Walking into a bookstore and buying a hundred novels to study is different from grabbing a hundred novels off a pirate site. Same books, very different legal exposure.
What the courts actually said
Two decisions out of the Northern District of California have done most of the heavy lifting here: Kadrey v. Meta Platforms, Inc. and Bartz v. Anthropic PBC. Together they sketch how judges are thinking about copyrighted material in training data.
In Bartz v. Anthropic, the court found that training an AI on copyrighted works could qualify as fair use. But it denied summary judgment for Anthropic on a separate issue: the company’s use of pirated copies to assemble a central library. Training, transformative. Building a warehouse of pirated books to train from, a different problem entirely.
A federal court has since described training on copyrighted books as “quintessentially” transformative fair use, language that legal commentators picked up on in analyses published as recently as March 2026. “Quintessentially” is a strong word for a judge to reach for. It suggests the court did not see this as a close call on the training question itself.
And then there is the number that got everyone’s attention. Judge William Alsup ordered Anthropic to pay a $1.5 billion copyright settlement to a group of writers whose works were used to train the company’s model. That is not a rounding error. It is one of the first rulings of its kind, and it lands squarely on the acquisition side of the line rather than the training side.
Why this matters if you’re not a lawyer
If you use AI agents at work, or you are thinking about building one, this affects you in ways that are easy to miss.
- Provenance is now a product feature. Where the training data came from is no longer a back-office detail. It is a liability question, and increasingly a selling point.
- “The model learned it” is not a defense for how you got it. The transformative-use finding protects the learning step. It does not retroactively clean up an illegal download.
- Vendor questions are fair game. Asking an AI provider about licensing and data sourcing is a normal procurement question now, not an act of paranoia.
For writers, the picture is genuinely mixed. The transformative-use rulings are not what most authors hoped for. But the settlement figure shows that courts are willing to make piracy expensive, and that gives authors real use over how their work enters these systems.
The messy middle we’re in
What makes this hard cleanly is that we have a handful of district court decisions, not settled national law. Different judges, different fact patterns, different outcomes on the details. Appeals are a normal part of this process. New cases are being filed. The rules that apply in 2026 may look different by the time your next model update ships.
So my honest advice, as someone who spends a lot of time translating this stuff: treat the current state as a working understanding rather than a final answer. Check current laws and rulings before you make a decision that depends on them, and be skeptical of anyone, including confident bloggers, who tells you the question is closed.
The Napster era did eventually resolve. It resolved into licensing deals, subscription services, and a music industry that looks nothing like it did in 1999. My guess is that AI and books head somewhere similar, with licensing markets forming around what courts make expensive and what they permit. We are just early enough that the shape of it is still being argued in courtrooms rather than contracts.
🕒 Published:
Related Articles
- Cursor vs Continue: Qual Escolher para Empresas
- Aplicación de Agente de IA en Salud
- Por que uma rodada de seed de $65 milhões para agentes de IA sinaliza algo maior do que dinheiro
- Selbst unsere KI-Agenten könnten betroffen sein: Der Trivy-Hack zeigt, dass Risiken in der Lieferkette überall vorhanden sind.