In 2024, several authors filed a class-action lawsuit accusing Anthropic of downloading millions of pirated ebooks from "shadow libraries" like LibGen and incorporating approximately 482,000 copyrighted works into its Claude training database. The court subsequently separated "model training" from "data acquisition," ruling that AI learning from legally obtained books might constitute fair use; however, downloading and storing works from pirated websites long-term constitutes infringement. Anthropic ultimately agreed to a $1.5 billion settlement, averaging about $3,000 per work, and to destroy the related pirated files.
Anthropic had previously accused some Chinese AI teams of extracting Claude output through APIs for model distillation, claiming this violated intellectual property rules. While the two behaviors are not entirely legally identical, they present a stark contrast: AI companies emphasize the right of their models to learn from human works; yet, they demand stronger exclusive protection for their own model outputs. Therefore, the core boundary is: AI training itself may not be illegal, but the method of obtaining training data must be legal.