Anthropic agreed to a $1.5bln settlement in a 2024 class action by authors
alleging the company downloaded millions of pirated e‑books from LibGen and
other shadow libraries and included about 482,000 copyrighted works in Claude’s
training data. The court split the legal issues, finding that using lawfully
obtained texts for model training can be fair use, but downloading and long‑term
storage of works from pirate sites constitutes infringement. The settlement
averages roughly $3,000 per work and requires destruction of the pirated files.
Anthropic had earlier accused some Chinese AI teams of extracting Claude outputs
via API for model distillation; the court noted those practices are legally
distinct. Core takeaway: model training is not per se illegal, but data
acquisition must comply with copyright law.