Two federal rulings have now put actual legal contours around the question that has haunted the AI industry for three years: can companies train models on copyrighted work? One case drew a clear line between lawful training and pirated data, settling for $1.5 billion. The other left the industry's biggest unanswered question intact - whether chatbot outputs that compete with original reporting can survive a copyright challenge.
Bartz v Anthropic: training is fair use, piracy is not
In Bartz v Anthropic, a class of authors alleged their books were illegally copied to train Claude. The court split the case down the middle. Training an AI model on copyrighted books constitutes fair use, the court held - the clearest judicial endorsement yet of the practice every large language model depends on. But storing pirated copies does not qualify as fair use.
The distinction is between what you do with a book and how you obtained it. Learning from a lawfully acquired text is transformative; maintaining a library of illegally downloaded ones is straightforward infringement, regardless of what you subsequently do with it.
The case settled for $1.5 billion - an estimated $3,000 per work. For most companies, that number matters more than the legal reasoning. It establishes a price. Any laboratory that assembled training data from pirated sources now has a rough figure for what that exposure costs per book.
The New York Times case survives: market substitution is alive
New York Times v OpenAI addresses a different, harder question for the industry. The Times sued OpenAI and Microsoft over use of millions of its articles to train GPT models. OpenAI moved to dismiss. The court denied that motion, finding the Times had plausibly alleged that ChatGPT's outputs compete with the Times's own content.
That framing is the threat. The Bartz reasoning protects training as transformative - the model learns from the work rather than reproducing it. But if a model's outputs substitute for the original in the market, the transformation argument weakens considerably. A reader who gets the substance of a Times article from a chatbot has not bought the article.
Denying a motion to dismiss is not a finding of liability. It means the claim is serious enough to proceed. But it establishes that market substitution is a live question AI companies will have to answer on the facts.
What the two rulings establish together
Read side by side, a workable rule emerges. Training on copyrighted material is likely lawful, provided the material was lawfully obtained. Acquiring it through piracy creates liability independent of how it is used. And whether the resulting model's outputs compete with the original in the market remains genuinely open, with the largest test case still to be decided.
For AI companies, the practical implication is about provenance. The legal exposure attaches less to the act of training than to the chain of custody of the data. Knowing where every corpus came from is now a compliance requirement rather than a courtesy.
The rulings land alongside a broader reckoning. AI-related securities class actions were only about 13 per cent of core filings in the first half of 2026, but accounted for nearly three-quarters of all alleged investor losses - fifteen filed in six months, nearly matching all of 2025. Meanwhile the EU AI Act's Article 50 obligations became enforceable on August 2, requiring AI systems to identify themselves and AI-generated content to be labelled. Anthropic responded by embedding invisible watermarks in Claude's text and signed C2PA provenance metadata in generated files, applied worldwide rather than only in Europe.
The pattern across all of it is the same: the period in which AI development ran ahead of the law is closing, and it is closing through ordinary contract disputes, securities suits, and copyright claims - rather than through dedicated AI legislation.
Why this matters for legal professionals
The $1.5 billion figure gives your clients a concrete benchmark for exposure, something that did not exist before. For lawyers advising AI companies or rights holders, provenance tracking is where compliance now starts. The shift alters how you brief a client: before ruling on the legality of training, verify how the data was obtained. The New York Times case means every model that outputs content resembling a source's reporting carries a live market-substitution claim that cannot be dismissed early.
For professionals working at the intersection of law and AI, these rulings also signal a shift in what keeps you current, from abstract debate to concrete data-handling and market-competing scenarios. AI Legal Assistant Courses now cover exactly these kinds of compliance and document-handling issues, and AI for Legal resources track the regulatory changes professionals need to follow.
Your membership also unlocks: