avalw news
William BensonWilliam BensonVIEW PROFILE →

Inside the Sony and Warner Lawsuit That Could Reshape How AI Companies Train Their Models

world2026-08-31 · 1 min read · 1 reads

Two of the biggest names in music publishing are accusing Anthropic of a "brazen campaign" of illegal torrenting, and they want up to $150,000 for every song caught in the crossfire.

Two of the biggest names in music publishing are accusing Anthropic of a "brazen campaign" of illegal torrenting, and they want up to $150,000 for every song caught in the crossfire.

Late on a Friday night, the kind of timing lawyers use when they'd rather not make headlines right away, Sony Music Publishing and Warner Chappell Music quietly filed one of the most aggressively worded copyright lawsuits the AI industry has seen yet. It named Anthropic, the company behind the Claude family of AI models, along with its two co-founders personally, and it didn't hold back on language. The complaint accuses the company of running "one of the largest and most blatant ongoing thefts of intellectual property in history." That's a striking accusation to level at any company, let alone one valued in the hundreds of billions of dollars and already sitting at the center of the broader AI boom. Here's what the lawsuit actually claims, how it connects to a previous, already-massive settlement Anthropic paid out, and why this particular case could end up mattering more than most. What Sony and Warner Are Actually Alleging The lawsuit, filed in the U.S. District Court for the Northern District of California, was first reported by Music Business Worldwide after the filing went largely unannounced. At its core, the 48-page complaint accuses Anthropic and its co-founders, CEO Dario Amodei and Benjamin Mann, of running what the plaintiffs describe as a deliberate, large-scale campaign of illegally torrenting, scraping, and downloading copyrighted musical works specifically to train and commercialize the Claude AI models. The publishers are seeking a jury trial, along with statutory damages of up to $150,000 for every work found to have been willfully infringed, plus an additional $25,000 for each instance in which Anthropic is alleged to have stripped or altered copyright management information from a work. Given that the complaint alleges "thousands upon thousands" of copyrighted musical compositions were involved, the total exposure here could realistically run into the billions of dollars if the case goes the plaintiffs' way.

Inside the Sony and Warner Lawsuit That Could Reshape How AI Companies Train Their Models

The Specific Claims Inside the Complaint

According to the filing, the case rests on four separate legal counts. The first two involve direct and contributory copyright infringement tied specifically to torrenting activity, with the contributory claim aimed personally at Amodei and Mann rather than the company as a whole. The remaining two counts, aimed at Anthropic alone, allege direct infringement more broadly and the unlawful removal or alteration of copyright management information, the metadata that typically identifies a work's rights holder.

Perhaps the most eye-catching specifics in the complaint involve claims about how the underlying data was actually obtained. The lawsuit alleges that in June 2021, Mann personally used BitTorrent to download at least five million pirated books from Library Genesis, a well-known online repository of pirated texts, and that Anthropic employees separately torrented at least two million additional books from a similar source called Pirate Library Mirror the following year.

"Defendants Anthropic and its founders Dario Amodei and Benjamin Mann have conducted a brazen campaign of illegally torrenting, scraping, and downloading copyrighted works on a massive scale in order to develop, operate, and reap enormous profits from Anthropic's 'Claude' series of artificial intelligence models." — From the Sony Music Publishing / Warner Chappell complaint

Photo: Sasun Bughdaryan / Unsplash — the lawsuit brings four separate legal counts, ranging from direct copyright infringement to allegations of unlawfully removing copyright management data.

Why Books Are at the Center of a Music Lawsuit

It might seem strange that a lawsuit filed by music publishers leans so heavily on allegations about pirated books, but the connection makes more sense once you understand how lyrics actually get captured in large text datasets. Books, including songbooks, biographies, and printed lyric collections, frequently contain full song lyrics and sheet music embedded within their pages. The complaint argues that by allegedly pirating millions of books en masse, Anthropic also swept up enormous volumes of copyrighted musical content along with everything else.

Beyond the books angle, the complaint also alleges that Anthropic scraped lyrics directly from licensed lyric databases including MusixMatch and LyricFind, services that pay to license lyrics legally and distribute them through proper channels. The suit further claims Anthropic engaged in what it describes as "destructive scanning" of secondhand physical books, and separately drew on large publicly available datasets including Common Crawl, The Pile, and a collection known as Books3.

The specific songs named in the complaint

To illustrate the alleged harm concretely, the lawsuit cites specific, well-known songs it claims Claude can reproduce lyrics from nearly verbatim, including "Ain't No Mountain High Enough," "All I Want for Christmas Is You," "Eye of the Tiger," "Here Comes Santa Claus," and "Paper Rings." That kind of specificity is a common legal tactic, turning an abstract dataset argument into something a jury can immediately recognize and relate to.

This Isn't Anthropic's First Rodeo With This Exact Issue

What makes this lawsuit particularly pointed is that it isn't emerging out of nowhere. In September 2025, Anthropic agreed to what was, at the time, the largest copyright settlement in U.S. history, a $1.5 billion deal resolving a class action brought by authors and publishers, widely referred to as the Bartz case. That case centered on very similar underlying conduct, a federal court found Anthropic had illegally torrented more than seven million copyrighted books from the same Library Genesis and Pirate Library Mirror sources named again in this new complaint.

Sony and Warner's legal team appears to be leaning directly into that history. The new complaint explicitly references the prior case, noting that another district court had already described Anthropic's conduct as "straightforward piracy but at massive scale." The publishers go further, arguing that the $1.5 billion Bartz settlement wasn't nearly large enough to actually change Anthropic's behavior going forward.

Inside the Sony and Warner Lawsuit That Could Reshape How AI Companies Train Their Models

The line that's drawing the most attention

The complaint includes an unusually blunt accusation about corporate strategy: that Anthropic "clearly considers that to be just the cost of doing business given that its entire business model continues to be built on copyright theft," and that the prior $1.5 billion settlement was "not large enough to deter infringing conduct" from a company the suit describes as now carrying a roughly $2 trillion valuation.

How This Case Differs From the Others Already Filed

Anthropic isn't new to facing lawsuits over its training data, and this isn't even the first case brought specifically by music industry plaintiffs. Universal Music Group, Concord, and ABKCO filed suit back in 2023, and BMG Rights Management brought its own case in March 2026, alleging infringement tied to 493 specific musical compositions. What sets the Sony and Warner complaint apart is its sheer scope: rather than focusing on a defined, relatively narrow list of songs, it alleges infringement across "tens of thousands" of copyrighted compositions, a scale that dwarfs the earlier music-industry filings against the company.

Axios technology reporter Sara Fischer has pointed out an added layer of legal complexity specific to music copyright cases like this one: a single song can carry multiple separate copyrights spread across different rights holders, covering the lyrics, the underlying composition, and the actual sound recording independently. That fragmented ownership structure means AI companies training on music-adjacent content can end up facing overlapping lawsuits from entirely different rights holders over the exact same songs.

Photo: Mick Haupt / Unsplash — music copyright is notoriously
fragmented, with lyrics, composition, and sound recording rights often held by
entirely different parties, a structure that exposes AI companies to multiple
overlapping lawsuits.
Photo: Mick Haupt / Unsplash — music copyright is notoriously fragmented, with lyrics, composition, and sound recording rights often held by entirely different parties, a structure that exposes AI companies to multiple overlapping lawsuits.

What Happens Next

As of the most recent reporting, Anthropic had not been reached for comment before the initial coverage of the lawsuit broke, and the company has not yet filed its formal response in court. Given the pattern set by the earlier Bartz case, legal observers widely expect Anthropic to eventually pursue either a negotiated settlement or a protracted court fight rather than a quick resolution, mirroring how the company handled the previous author and publisher lawsuit before it eventually agreed to that $1.5 billion payout.

Beyond the specific dollar figures at stake, cases like this one are steadily building the body of case law that will define how courts treat AI training data going forward, particularly the crucial legal distinction between using copyrighted material at all, which some rulings have found permissible under fair use in certain contexts, and how that material was actually acquired in the first place. The Bartz ruling drew exactly that line: using copyrighted books to train an AI model wasn't necessarily illegal on its own, but acquiring those books through piracy was a separate and serious problem regardless of what the material was ultimately used for.

Inside the Sony and Warner Lawsuit That Could Reshape How AI Companies Train Their Models

Why this case could set a bigger precedent than it looks

Legal fights over AI training data are still a relatively young area of law, and each new ruling adds real weight to how future cases, against Anthropic and its competitors alike, are likely to be decided. A case this large, alleging infringement across tens of thousands of works rather than a narrow list, gives a court the opportunity to rule on questions of scale and intent that smaller cases haven't fully tested yet, which is part of why this particular lawsuit is being watched so closely across both the music and AI industries.

Inside the Sony and Warner Lawsuit That Could Reshape How AI Companies Train Their Models
William Benson
Stay updated
William Benson
Subscribe to get an email whenever William Benson publishes a new story. No spam, unsubscribe anytime.
William Benson
WRITTEN BY THE AUTHOR
William Benson
2026-08-31 · 1 min read · 1 reads
View profile →
VERIFY THIS STORY
ASK AI
MORE FROM William Benson
Report this articlesupport@avalw.com