Anthropic faces another lawsuit from music publishers, escalating the copyright battle in AI training
Music publishers such as Sony and Warner have sued Anthropic, accusing the company of obtaining sheet music and lyrics from pirated sources through the Seed network to train Claude. The settlement of the book case cannot hide the music copyright dispute.
The lawsuit documents exposed the internal chat records, and the pirated library became the core evidence.
This new lawsuit was jointly initiated by music publishers such as Sony and Warner. The complaint cited the internal employee chat records of Anthropic, with a sentence like "Zlibrary, my beloved", bringing to light the way this AI company obtained training data. According to the lawsuit materials, the entire process of collecting pirated data was initiated in 2021. The co-founder of Anthropic, Benjamin Mann, personally used the BitTorrent seed tool to download millions of pirated books from LibGen.
The download plan was approved by the company's CEO, Dario Amodei. After the FBI took down LibGen, the pirated resources were moved to the PiLiMi mirror site. Mann quickly arranged for employees to keep downloading through seeds. The downloaded materials were not just any ordinary books. Publishers discovered that the file packages contained a large number of sheet music and lyric collections. Publishers emphasized that this was not a scattered and small-scale scraping. Instead, it was an organized piracy acquisition project, with the aim of preparing training materials for the Claude large model.
The settlement of the book case does not mean that the music copyright issue has been resolved.
Before this, Anthropic had reached a $1.5 billion settlement with the group of book authors. At that time, many outside opinions believed that this matter was basically closed. However, the music publishers did not accept this outcome. The plaintiff stated that the $1.5 billion compensation amount was insufficient in terms of its deterrent effect compared to Anthropic's use of this training data to increase the company's valuation to $2 trillion. The ruling of the book case cannot be directly applied to the music copyright issue.
The judge's view on the book case was that large language model training is similar to readers reading books. The model learns the patterns of words and tries to generate new content, not to copy the original work directly. However, the music publishers presented a different argument. The market for songwriting works is more direct. AI-generated songs will directly enter the music charts and compete with human creators for traffic and royalties. Even if AI works are not completely the same as the original works, they will dilute the income pool of the creators.
The controversy over model training: Did pirated data indirectly flow into commercial Claude?
Anthropic's public statement is that the commercial version of Claude did not directly use the pirated materials downloaded from LibGen or PiLiMi for training. However, music publishers countered by pointing out the AI training process. Large model development is not just one pre-training stage. Many teams first train a non-commercial model using a basic dataset, and then use this model to generate synthetic data. Publishers suspect that Anthropic trained the pre-model using pirated datasets and then used the generated synthetic data for Claude's iteration.
Anthropic is still retaining the LibGen dataset to test model outputs and prevent the generation of text that is too similar to the original. That is to say, the pirated library has not been completely destroyed after the training process. When users ask Claude for song chords, the model often outputs complete copyrighted lyrics. Users can also ask Claude to change a certain section of lyrics to a Beyoncé style, and Claude can directly complete the rewrite.
The trial materials also mentioned another way Anthropic collected materials. The team purchased physical sheet music, dictionary books, scanned them to create digital copies, and then destroyed the original physical books. This operation is considered a reasonable use in book cases. On one hand, there is purchasing physical books and scanning them. On the other hand, there is batch downloading of pirated resources from BT seeds. Both are preparing text materials for AI, but the legal risks are completely different. The former is recognized by the court, while the latter is characterized by publishers as large-scale infringement.
The compliance threshold for AI training data has risen.
In the past, many AI teams assumed that as long as the final model did not directly reproduce the original work, the process of reading copyrighted texts during the training stage was considered a reasonable use. If music publishers obtain a favorable judgment, this perception will be overturned. The market attributes of music copyrights and book copyrights are different. The commercial realization chains of songs and lyrics are short, and the streaming chart can directly show the impact of AI works on live music. When judges assess market damages, they are more likely to believe the evidence provided by the distributors. In the future, major AI companies will be more cautious in distinguishing the sources of training materials.