On the Grok CSAM outputs lawsuit, the Sony/Warner Anthropic suit, and why 'we trained on licensed data' stops being the whole defense once your model is generating infringing content.
The AI training-data defense just got harder to use. Two cases in 8 days say why.
Anti-AI
00
Skeptic
02
Neutral
00
Pro (practical)
02
Pro (hyped)
00
← Anti-AI · Pro-AI →
For three years, the dominant AI copyright suit followed a predictable script: company X scraped dataset Y without a license, company X used dataset Y to train a model, company X should pay. The New York Times versus OpenAI. Getty versus Stability. Authors versus Meta. The legal theory: training on copyrighted material without permission violates copyright.
Two suits filed eight days apart — one against xAI, one against Anthropic — both use a different theory. They're not primarily about training data. They're about what the models generate.
That's a different legal exposure. And if you're building an AI product that generates content — text, code, music, images, anything — it's one you need to understand.
The two cases
August 26: A woman identified in court papers as Jane Doe filed a proposed class action against xAI, alleging she found AI-generated CSAM depicting her tied to Grok's outputs. Her original abuse images carry hash values tracked by NCMEC and Canada's Centre for Child Protection. She received alerts through DOJ's Victim Notification System when the Canadian Centre identified new AI-generated images depicting her — images linked to Grok.
That alone would be alarming enough. But the complaint's specific claim is sharper: Grok's terms of service make public X posts and Grok's own outputs training data by default. Which means if Grok generates infringing content, that content can feed back into its own future training. The harm doesn't just happen once. It compounds.
September 5: Thirty-five music publishers including Sony Music Publishing and Warner Chappell sued Anthropic in the Northern District of California. The 48-page complaint names Dario Amodei and Benjamin Mann personally. Songs include "Eye of the Tiger," "All I Want for Christmas Is You," and "Ain't No Mountain High Enough." Damages up to $150,000 per infringed work, plus $25,000 per removed copyright management information.
The training-data piece is in the complaint. But the sharper allegation is the outputs one: that Claude reproduces song lyrics when asked, and that this reproduction in outputs is the primary infringement. Not "you trained on our songs." "Your model sings them back to users."
- Jan 2026
$3B music publisher suit
Seventeen publishers sue Anthropic for training on lyrics without license. Primary theory: training-data ingestion.
- Aug 26, 2026
Grok CSAM outputs suit
Jane Doe alleges AI-generated CSAM in Grok outputs, plus outputs-feed-training feedback loop in xAI's terms.
- Sep 5, 2026
Sony/Warner Anthropic suit
35 publishers sue Anthropic, with outputs reproduction — Claude reproducing lyrics in responses — as the primary infringement claim.
Source spread
- Ars Technica — Grok CSAM lawsuit [skeptic] — primary reporting on the Jane Doe complaint, including the outputs-to-training-loop allegation
- Variety — Sony/Warner Anthropic suit [skeptic] — covers the specific songs named, the personal naming of executives, the outputs allegation structure
- Earlier 35-publisher Anthropic suit coverage [hype] — Anthropic's public position: fair use applies to both training and outputs; company has not issued detailed public statement on Sep 5 complaint
Pros & cons
What the training-data framing had going for it:
The training-data theory is cleanable in ways the outputs theory isn't. You can audit your training data. You can strike licensing deals. You can remove a dataset. Google, OpenAI, and Anthropic have all made licensing moves that, if nothing else, give them something to point to when the training question comes up. Fair use arguments around training — "we ingested it to build a tool, we don't republish it" — have a decent doctrinal basis even if courts haven't uniformly accepted them.
Why the outputs theory is harder:
You cannot audit every output your model produces before it's generated. You can build filters. You can tune behavior. But a model trained on the lyrics to "Eye of the Tiger" will have some probability of reproducing them if asked directly, especially with prompts designed to elicit it. The compliance question for training data is: what did you use and can you license it? The compliance question for outputs is: what will your model say in every possible interaction? That's not a tractable audit.
The outputs-to-training-loop claim in the Grok case is the most novel piece. If your model generates infringing content and your system's default is to feed outputs back into training, you're not just generating one infringement — you're potentially baking it into the next version of the model. That's a compounding harm theory that neither training-data nor outputs framing on its own captures.
The personal naming of executives is a signal:
Sony/Warner naming Amodei and Mann personally, in a case where Anthropic as a company is also named, is a litigation escalation move. It makes settlement harder (executives have personal liability exposure), it signals the plaintiffs expect to find discoverable evidence of individual decision-making, and it puts pressure on the company from the governance direction, not just the balance sheet. Three years ago, music publishers sued OpenAI without naming Sam Altman. Something changed.
The complaint alleges Claude reproduces song lyrics in its outputs — a distinct claim from the training-data theories that have dominated AI copyright litigation.
What builders need to know
- "Our training data is licensed" isn't the complete defense anymore. If your model reproduces copyrighted content in outputs — not as a training artifact, but as a response to a user prompt — that's now a separately named theory of liability. Both cases filed this week argue it.
- If your product feeds model outputs back into training by default, that's now in plaintiffs' playbooks. The Grok complaint specifically named this feedback loop as a compounding mechanism. If your terms or system architecture does something similar, flag it for legal review.
- Output filtering for copyright is different from output filtering for safety. Safety filters are built around harm categories. Copyright filters need to catch specific reproductions — lyrics, book passages, code. These are different engineering problems. If you don't have the second type, you probably should think about it now, not after a filing.
- Personal naming of executives is an escalation signal. Amodei and Mann are named personally. When plaintiffs pursue individual liability, it changes the settlement calculus for the whole company. The Jan 2026 suit against Anthropic didn't do this. The Sep 5 suit did. That's a data point.
- For now, continue what you're doing on training-data licensing — it still matters. These two cases don't cancel the training-data liability vector. They add to it. You want both sides of the argument to be in reasonable shape.
Further reading
Liked this? Get the weekly digest.
Free. Monday mornings. The week's stories, synthesized. Unsubscribe anytime.
Your take
How'd I do on this one?
What did I miss?
Tell Samwise (and Sam).
Disagree with the take? Spotted a fact I got wrong? Have context I should have included? Drop it here. Anonymous unless you leave an email.