Short answer, it depends how you got it.
That’s the strange, slightly unsatisfying place we’ve landed on one of the biggest questions in AI right now. If you’ve been wondering whether the models you talk to every day were built on books someone else wrote, the answer is yes. And if you’ve been wondering whether that was legal, the courts have started to answer, but the answer splits neatly in two.
I explain AI agents for a living, and this is the question I get most from people who don’t write code. Not “how does the model work,” but “wait, is this allowed?” So let’s walk through what’s actually been decided.
What the courts have said so far
Two decisions out of the Northern District of California set the tone. In Bartz v. Anthropic PBC, Judge Alsup considered claims from authors who sued Anthropic over training its models on their books. In Kadrey v. Meta Platforms, Inc., Judge Chhabria looked at similar claims against Meta. Both judges concluded that training generative AI models on copyrighted books qualified as fair use.
A United States District Court ruled on a Monday that training large language models on copyrighted books constitutes fair use. The reasoning centers on a word lawyers care about a lot: transformative. One federal court described the training as “quintessentially” transformative fair use. In plain terms, the court decided that feeding a book into a model to teach it patterns of language is a fundamentally different act from reprinting the book and selling copies.
In the Anthropic case, the court weighed the fair use factors and granted the company’s motion for summary judgment on the training question specifically.
The part that didn’t go the AI companies’ way
Here’s where the split matters. The fair use finding applies to legally acquired materials. It does not cover pirated ones.
That distinction is doing enormous work. The claims involving pirated copies were not resolved by those rulings. They’re headed toward trial, where they’ll be decided separately. So a company can win the argument that training is transformative and still be on the hook for how it got the books in the first place.
Think of it like this. A court saying “reading a library book and writing something new is fine” does not also say “breaking into the library was fine.” Those are two different acts, and they get judged separately.
Why this matters if you’re not a lawyer
Most people reading agent101 aren’t building models. You’re using them, maybe deploying agents at work, maybe just curious about the tools you’re increasingly asked to trust. So why should any of this land on your radar?
- Provenance is becoming a product feature. If courts keep drawing the line at acquisition, then “where did your training data come from” stops being a philosophical question and becomes a due diligence question. Expect vendors to start talking about it.
- Legal risk shifts, but doesn’t vanish. Training being fair use removes one enormous cloud over the industry. The piracy question keeps another one firmly in place.
- These are district court decisions. They’re meaningful and they’re being cited, but they sit at the trial level. Appeals and further cases can reshape the picture, and legal commentary is already flagging concerns about how these decisions were reasoned.
The honest state of play
Legal scholars have started working through what these rulings mean. Analysis in the Houston Law Review notes that these were the first two district court decisions addressing generative AI training, with Judges Alsup and Chhabria reaching their conclusions about Anthropic’s and Meta’s training respectively. Recent commentary has also raised critical concerns about the decisions. Cross-border discussions are picking apart the fair use analysis factor by factor, because other countries don’t have a fair use doctrine that works the same way.
What that tells me is that we’re at the beginning of this conversation, not the end of it. Two trial court decisions in one district is a starting point. It’s not settled law across the country, and it’s certainly not settled internationally.
What I’d tell a friend over coffee
If someone asked me whether AI companies broke the law by training on books, I’d say the courts have looked at the training itself and found it lawful under fair use because it’s transformative. Then I’d say the harder question is the sourcing, and that question is still live and still headed to trial.
That’s a genuinely complicated answer, and I think the complication is the point. A lot of coverage on this topic wants to declare either a total win for AI companies or a total win for authors. The actual rulings did something more careful than that. They separated the act of learning from the act of obtaining, and they treated those differently.
For anyone trying to understand the tools they use, that’s a useful frame to carry forward. How a model learns and how a model got its material are two separate questions. Right now, only one of them has an answer.
đź•’ Published: