According to Reuters, the U.S. federal court officially approved a $1.5 billion class-action settlement between Anthropic and authors and publishers on Monday. The agreement is intended to resolve litigation against Anthropic over alleged copyright infringement during the training of its AI models, and the company will begin paying compensation to the relevant copyright holders.

Judge Araceli Martinez-Olguin of the U.S. District Court for the Northern District of California signed the final approval document. Previously, retired judge William Alsup had provisionally approved the settlement and ruled that Anthropic had illegally downloaded and stored millions of copyrighted books.

Under the compensation plan, each affected work is expected to receive $3,000 in damages, involving approximately 500,000 works, with the amount to be shared among the authors and publishers who hold the copyrights. This sum is considered one of the largest settlements in the field of U.S. copyright law.

Anthropic, Claude, claude

However, the agreement does not fully resolve the legal disputes surrounding AI training data. Judge Alsup previously supported Anthropic on the core issue of the case, stating that using copyrighted texts to train AI models may fall under "fair use," which was seen as an important development in AI copyright cases. At the same time, he pointed out that some of the ways Anthropic obtained training data were illegal.

Anthropic had primarily obtained training data through two methods: some from books legally purchased and scanned, and others from pirated websites such as Library Genesis and Pirate Library Mirror. The court deemed the latter method as illegal. To avoid further trials and potential high penalties, Anthropic ultimately chose to reach a settlement.

Industry experts note that while this settlement ends the case, it has not established a unified legal standard for the entire AI industry. Because Anthropic chose to settle, the case did not proceed to an appeal, and Judge Alsup's previous ruling on "fair use" in AI training will not become a binding judicial precedent.

Currently, lawsuits over the copyright of training data for AI models are still ongoing, including cases against companies such as Google, Meta, Midjourney, and OpenAI. Recently, publishing houses such as Hachette, Cengage, Elsevier, author Scott Turow, and SCRIBE also jointly filed a lawsuit against Google, accusing it of using copyrighted works without permission to train the Gemini AI platform.

As generative AI continues to develop rapidly, how to balance the needs of model training with copyright protection has become a major legal and business challenge for the global AI industry. While the Anthropic settlement provides one solution path, the long-term rules for AI data compliance in the industry remain to be further clarified.