|
Getting your Trinity Audio player ready...
|
The final approval of Anthropic’s $1.5 billion copyright settlement closes one of the most closely watched artificial intelligence cases to date. For corporate legal departments, however, its most important lesson is not the size of the payment.
It is the distinction the court drew between what an AI company does with copyrighted material and how it acquires that material.
On July 20, US District Judge Araceli Martínez-Olguín approved the class settlement in Bartz v. Anthropic, resolving claims involving more than 482,000 books allegedly acquired from online piracy repositories. Approximately 91% of the covered works had been claimed by authors or publishers, with eligible rights holders expected to receive roughly $3,000 per work. The court described the agreement as providing meaningful relief while approving approximately $101 million in attorneys’ fees, well below the roughly $187.5 million requested by class counsel.
The settlement is believed to be the largest copyright recovery in US history. Yet it does not establish a binding appellate rule on whether copyrighted works may be used to train generative AI systems.
For general counsel, that uncertainty is not a reason to delay action. It is a reason to examine AI governance more closely.
The court separated fair use from unlawful acquisition
The case produced a legal outcome that appears contradictory at first glance.
In an earlier ruling, US District Judge William Alsup held that Anthropic’s use of books to train its language models could qualify as fair use. The training process was considered transformative because the models were designed not to reproduce the books but to learn statistical relationships within language and generate new outputs.
That aspect of the ruling represented a significant legal victory for AI developers.
The court reached a different conclusion about Anthropic’s acquisition and retention of millions of books from shadow libraries such as LibGen and PiLiMi. Copying books from piracy sites to create a permanent internal library raised a separate infringement issue, regardless of whether their later use for training qualified as fair use.
Fair use is not a blanket defense covering every stage of AI development.
A company may have a credible fair-use argument for model training while still facing liability arising from the original reproduction, storage or transfer of copyrighted material. The legal analysis may also differ depending on the source of the data, the purpose of each copy, how long the material is retained and whether it is used for activities beyond training.
The case reinforces a longstanding principle of intellectual property law: lawful use and lawful possession are separate questions.
For legal departments, that means an AI governance program focused only on model outputs is incomplete. Counsel must also examine the upstream data supply chain.
AI due diligence must reach the training data
Many companies assess AI vendors by reviewing cybersecurity controls, data privacy terms, service availability and limits on the use of customer information. The Anthropic litigation shows why copyright provenance belongs in that process.
A procurement team may know what an AI system does without knowing what content was used to build it. A contract may promise that customer prompts will not be used for training while saying little about the material used to train the existing model. A vendor may offer broad assurances of legal compliance without identifying the sources, licenses or acquisition methods behind its datasets.
Those gaps can leave an enterprise dependent on technology carrying risks it cannot independently assess.
Legal teams should ask vendors to explain how training and fine-tuning data were sourced, whether copyrighted materials were licensed and what controls prevent the ingestion of pirated or unlawfully obtained content. The answers may not provide a complete inventory, particularly for foundation models trained on enormous datasets, but they can reveal whether a vendor maintains a defensible data governance process.
Contract terms deserve similar scrutiny.
Representations and warranties should expressly address the vendor’s authority to use its training data, not only its authority to provide the finished service. Indemnification provisions should specify whether they cover claims related to training inputs, generated outputs or both. Liability limitations may warrant separate treatment for intellectual property claims. Audit rights, notice obligations and cooperation requirements can become critical if a vendor faces litigation that affects its customers.
Companies developing or fine-tuning their own models face a more direct obligation. They should document where datasets came from, what licenses apply, which copies are retained and who approved their use. Publicly accessible content should not automatically be treated as content that is free to copy.
The Anthropic outcome suggests that provenance records can become litigation records. A company that cannot reconstruct how it acquired its data may struggle to distinguish lawful training activity from unlawful copying.
A settlement can shape conduct without creating precedent
The approval order resolves the claims covered by the agreement, but it does not transform the earlier fair-use ruling into binding law across the country.
Because Anthropic settled rather than taking the piracy claims through trial and appeal, another judge may reach a different conclusion in a future AI copyright dispute. Cases against other technology companies may also involve different datasets, acquisition practices, model architectures and outputs.
The settlement’s release is limited. It covers claims tied to Anthropic’s past acquisition and copying of listed works while preserving certain claims involving AI outputs and conduct occurring on or after Aug. 25, 2025. The $1.5 billion fund is also nonreversionary, meaning unused funds will not return to Anthropic.
Those boundaries make the agreement more useful as a signal of legal risk than as a definitive statement of AI copyright law.
The financial terms establish a visible benchmark for resolving large-scale allegations involving pirated training materials. The settlement may influence negotiations among publishers, authors, model developers and data providers even while the underlying legal questions remain unsettled. The court’s decision to reduce the requested attorneys’ fees also reflects judicial scrutiny of how these landmark AI settlements are structured, not simply whether they are approved.
Corporate counsel should resist drawing an overly broad lesson from Anthropic’s fair-use victory. The stronger takeaway is that courts may divide the AI lifecycle into separate acts and evaluate each one independently.
Acquisition, storage, training, fine-tuning, retrieval, output generation and distribution can each raise distinct legal questions. A favorable argument at one stage may not cure a violation at another.
The case also illustrates why AI oversight cannot remain confined to an innovation committee or technology function. Copyright exposure can affect vendor selection, insurance coverage, contractual liability, disclosure obligations and reputational risk. Those issues belong within a company’s broader governance framework.
The next wave of AI cases may refine the fair-use analysis. They may test whether particular outputs are substantially similar to protected works, whether licensing markets affect the fourth fair-use factor and whether different forms of training require different rules.
General counsel do not need to predict every ruling. They need a governance process that can withstand scrutiny while the law develops.
Anthropic’s settlement offers a practical starting point: know where the data came from, document the right to use it and do not assume that a defensible purpose excuses an unlawful method of acquisition.
Source






