The latest front in the fight over how large language models get built is not a novel, a news archive, or a pile of source code. It is the American songbook. On August 28, Sony Music Publishing and Warner Chappell walked into federal court in Northern California and sued Anthropic, the lab behind the Claude family of models, along with CEO Dario Amodei and co-founder Benjamin Mann. The publishers’ complaint does not mince words. It calls Claude’s training “one of the largest and most blatant ongoing thefts of intellectual property in history.”

That sentence is designed to travel. It is also designed to frame a damages theory that, if a jury ever accepted it at the high end of the statute, would be large enough to rearrange the economics of a frontier lab. The filing seeks up to $150,000 per work plus $25,000 per stripped copyright-management tag across “tens of thousands” of compositions — potentially several billion dollars. Anthropic did not immediately comment.

The case sits at a junction the industry has been circling for more than two years: whether ingesting copyrighted text at industrial scale, then answering users who ask for a chorus or a verse, is research, product, or theft. Music publishers have decided they already know the answer. They are asking a court in the nation’s most important technology venue to agree.

The catalog on the complaint

Five songs, and a theory of harm

The works named in the papers are not obscure deep cuts. They are songs that have lived in radios, weddings, stadiums, and streaming playlists for decades. The complaint lists “Ain’t No Mountain High Enough,” the Holland-Dozier-Holland / Ashford & Simpson standard that Marvin Gaye and Tammi Terrell turned into a Motown monument. It lists Bon Jovi’s “Livin’ on a Prayer,” a working-class power ballad that still detonates when a bar band hits the key change. It lists Earth, Wind & Fire’s “September,” a calendar-date singalong that has outlived the decade that produced it. It lists “Hallelujah,” Leonard Cohen’s hymn of doubt and desire, covered so often that many listeners no longer know which version they first loved. And it lists Taylor Swift’s “Paper Rings,” a latter-day catalog entry that reminds the court this is not only a fight about the 1960s and 1980s.

Those titles do work that a damages spreadsheet cannot. They tell a judge, and later perhaps a jury, that the alleged taking is not abstract. It is the stuff people hum. Publishers have learned, in earlier technology fights, that specificity sells. Napster was not defeated by a theory of peer-to-peer networking. It was defeated by Metallica and a list of tracks. The same instinct is on display here.

Sony Music Publishing and Warner Chappell are two of the three global majors that control a huge share of the compositions underneath recorded music. The third, Universal, is already in the fight. The complaint’s choice of famous songs is a way of saying: if Claude was trained on the culture, it was trained on us.

How the publishers say the data was gathered

Torrents, a pirate mirror, and licensed lyric sites

The most explosive factual claims in the filing are not about output. They are about intake. The complaint says Mann torrented more than five million pirated books. It says staff grabbed at least two million more from Pirate Library Mirror. And it says the lab scraped lyrics from licensed sites such as MusixMatch and LyricFind.

Those three alleged pipelines matter for different legal reasons.

Torrenting a pirate library of books is, if proven, a straightforward copyright story: wholesale reproduction of works the lab did not buy, from a source that did not have the right to hand them over. Pirate Library Mirror is the sort of shadow archive that has circulated in machine-learning circles as a convenient, if radioactive, substitute for licensed text. Scraping MusixMatch and LyricFind is a different kind of allegation. Those services exist because lyrics are licensed. They sit behind commercial arrangements with publishers. To scrape them, the complaint implies, is not to wander into the public web. It is to take the one place where the industry already built a legal market and treat it as a free training set.

None of this has been tested at trial. Anthropic has not answered the complaint in public. A lawsuit is a set of accusations, not a verdict. But the publishers have put the lab’s data hygiene — who collected what, from where, and with whose blessing — at the center of the case. That is a more dangerous place for a defendant than a debate about whether a chatbot’s paraphrase of a lyric is transformative. Intake is easier to count. Intake leaves logs.

The arithmetic of statutory damages

Per-work dollars and stripped tags

American copyright law does not require a plaintiff to prove the exact dollar value of every lost license. For registered works, the statute allows a range of statutory damages. The publishers are asking for the top of that range: up to $150,000 per work. They add a second, less familiar number: $25,000 per stripped copyright-management tag.

Copyright-management information is the metadata that travels with a work — credits, identifiers, the quiet machinery that tells a licensee who owns this and how to pay. Federal law treats the removal of that information as its own wrong, separate from copying the song. If a training pipeline strips tags so that a model no longer knows, or no longer reveals, that a lyric belongs to a publisher, the complaint treats each stripped tag as a billable event.

Multiply those figures across “tens of thousands” of compositions and the publishers arrive at a headline they are happy to repeat: potentially several billion dollars. That is not a prediction of what a judge will award. It is a ceiling, and a bargaining chip. In copyright litigation, the ceiling is the point. It changes the settlement math. It changes how boards talk about reserve. It changes whether a company treats music as a rounding error or as an existential docket.

The same math explains why the publishers sued Amodei and Mann by name, not only the corporate entity. Individual defendants concentrate the mind. They also, in a case about alleged data collection, put a human face on decisions the lab might otherwise describe as research practice.

A $1.5 billion book settlement, and a crowded docket

This is not Anthropic’s first expensive argument about training data. The lab already settled a book-publisher case for $1.5 billion. That figure is now part of the public record of the AI copyright wars. It tells every subsequent plaintiff two things: Anthropic will pay when the facts and the forum look bad enough, and the market for settling these cases is no longer theoretical.

The music docket was already filling up before Sony and Warner Chappell filed. Anthropic faces suits from Universal, Concord, ABKCO, BMG, and Round Hill. With Friday’s complaint, all three major publishers are now in court against Claude’s maker. That is a structural fact, not a flourish. The majors do not always move in lockstep. When they do, defendants lose the hope that one house will break ranks and license around the others.

Independent and mid-major publishers — Concord, ABKCO, BMG, Round Hill — add catalog depth and a different political texture. ABKCO’s holdings include pieces of the Rolling Stones and Sam Cooke stories. BMG has spent years presenting itself as a more author-friendly alternative to the old majors. Concord’s catalog is a museum of American pop. Together they make it harder for Anthropic to argue that it is being singled out by a pair of conglomerates. The industry, in this telling, has arrived as a bloc.

Why lyrics are a sharper weapon than books

Books and news articles have been the first wave of generative-AI copyright suits because they are long, easy to quote, and easy to show in a screenshot. Lyrics are shorter, stickier, and in some ways more dangerous for a model maker.

A large language model that has seen a novel can often discuss its themes without reprinting a chapter. A model that has seen a lyric is one prompt away from reprinting the whole work. Pop songs are brief. Their value lives in exact sequence — the hook, the bridge, the last refrain. Users ask chatbots for lyrics the way they once asked search engines. If Claude recites a verse, the publisher does not need a sophisticated theory of substantial similarity. The output is the work.

There is also a licensing market that already exists. Lyric sites pay. Streaming services pay. Radio pays, through performing-rights organizations. Synchronization in film and advertising pays. The publishers’ theory is that Anthropic skipped the queue. It did not license MusixMatch or LyricFind. It did not sit down with Sony or Warner. It trained first and, in the publishers’ telling, hoped the law would treat the result as fair use.

Fair use is the doctrine every AI lab is counting on. It asks whether a use is transformative, how much of the work was taken, and what the use does to the market for the original. Training a model on text in order to teach it language is, labs argue, a transformative intermediate copy, more like reading than like reprinting. Rightsholders argue the opposite: that a commercial chatbot which can emit the work is the market, or a substitute for one. Courts have not settled that fight. They have produced a scatter of early rulings, some friendlier to training, some not, and a growing pile of settlements that look like insurance policies.

Music publishers have a particular reason to reject the reading metaphor. They already license the act of displaying a lyric. The market is not hypothetical. It has rate cards.

Northern California as the chosen field

Filing in Northern California is not an accident of geography. Anthropic is a San Francisco company. The Northern District is where the country’s most closely watched technology cases go to be interpreted by judges who have seen a decade of platform fights. It is also a jury pool that lives among the people who build these systems. That can cut either way. It can produce skepticism of old-industry plaintiffs. It can also produce impatience with labs that talk about safety in public and, if the complaint is right, torrent pirate libraries in private.

The decision to name Dario Amodei and Benjamin Mann keeps the case from becoming a purely corporate morality play. Amodei is the public face of Anthropic’s safety brand — the former OpenAI researcher who left to build a company that would, in its own telling, treat powerful models with more caution than its rivals. Mann is a technical founder whose name, in this complaint, is attached to an alleged torrent of more than five million pirated books. The contrast is the point. A safety lab, the publishers suggest, does not get to be precious about military customers and casual about catalogs.

What “ongoing” is meant to do

The complaint’s most important adjective may be ongoing. The publishers did not describe a historical sin, a training run that ended in 2023 and can be walked back with a settlement and a cleaner dataset. They described a practice. If Claude is still being trained, still being refreshed, still answering lyric prompts from weights that absorbed MusixMatch and LyricFind, then every day is a new count.

That framing matters for injunctions as much as for damages. A court that believes the taking is finished may be content to talk about money. A court that believes it is still happening can be asked to shut something down — a pipeline, a feature, a model’s ability to emit lyrics. Product teams fear injunctions more than they fear checks. Checks can be reserved. Features that vanish on a Friday afternoon are how users leave.

The wider war this suit joins

The music industry has been here before, or believes it has. Home taping. Sampling. Napster. YouTube’s first decade. TikTok’s soundtrack deals. Each time, the argument is that a new machine has made copying too easy and payment too optional. Each time, the industry arrives late, sues loudly, and eventually licenses. The open question is whether generative AI is another licensing cycle or a break in the pattern — a technology that does not need the song as a product, only as nutrition.

If it is the former, this case will end the way expensive copyright cases often end: in a confidential deal, a licensed lyric corpus, a watermark, a revenue share, and a press release about partnership. If it is the latter, the publishers are trying to make the price of nutrition so high that labs have to come to the table or starve the model of music.

The $1.5 billion book settlement suggests Anthropic already knows what the first path looks like. The presence of Universal, Concord, ABKCO, BMG, and Round Hill on the same field suggests the second path is the one the industry is preparing to walk if talks fail.

What the lab has not said

Anthropic did not immediately comment. That silence is ordinary on the day a complaint lands. It is also a vacuum that the publishers’ language will fill. “One of the largest and most blatant ongoing thefts of intellectual property in history” is not a line written for a chambers conference. It is a line written for the week between filing and answer, when the public story hardens.

When the lab does speak, it will have a familiar menu. It can deny the torrenting. It can say any lyrics in the training mix were incidental, filtered, or used in a way the law allows. It can point to output filters that refuse to reprint a full song. It can argue that suing a chief executive and a co-founder is theater. It can note that tens of thousands of works, at $150,000 each, plus tag damages, is a number built to frighten, not to measure harm.

What it cannot easily do is pretend the music industry is fragmented. All three major publishers are now in court against Claude’s maker. The catalog on the complaint — “Ain’t No Mountain High Enough,” “Livin’ on a Prayer,” “September,” “Hallelujah,” “Paper Rings” — is a reminder that the works at issue are not a niche. They are the water the culture swims in.

The next moves are procedural and slow: service, a motion to dismiss, a fight over what discovery will reach inside Anthropic’s data pipeline. Somewhere in that pipeline, if the publishers are right, is a torrent, a mirror, and a scrape of sites that already pay for the words. Somewhere in the damages theory is a per-work number large enough to make a frontier lab flinch. The songs, for their part, will keep playing. The question the Northern District has been asked to answer is who, if anyone, gets paid when a machine learns to sing them back.