What happened?
In mid-December 2024, a redacted filing in Kadrey et al. v. Meta Platforms – the authors’ copyright class action over Llama’s training data – pulled back the curtain on internal Meta deliberations that read like regulatory exhibits in search of a crime scene: engineers warning that torrenting LibGen from a corporate laptop “doesn’t feel right,” executives debating a dataset “we know to be pirated,” and a paper trail alleging the company stripped copyright management information from millions of books before feeding them to its models. The filings, and reporting into December as experts and press parsed them, reframe the case from a fair-use test into something harder: whether one of the world’s largest companies knowingly distributed pirated data – because BitTorrent uploads as it downloads – and sanitized it to hide provenance, while its CEO blessed the project. When this post publishes on 25 December 2024, the motion to amend is before Judge Chhabria in the Northern District of California, the DMCA claim is live, and the AI industry’s data-gathering era has its strongest-yet legal stress test.
Quick Answer: Kadrey et al. v. Meta Platforms (N.D. Cal., No. 23-cv-03417) is a class action by authors including Richard Kadrey, Sarah Silverman, and Ta-Nehisi Coates, alleging Meta trained its Llama models on their copyrighted books sourced from shadow libraries – LibGen, Anna’s Archive, Z-Library, Bibliotik. The December 2024 filings (a reply supporting a third amended complaint, unsealed in redacted form mid-month) added two claims: a DMCA violation for allegedly stripping copyright management information (CMI) – copyright notices, acknowledgements, attribution lines – from the training corpus, and a CDAFA claim over the acquisition method. The exhibits allege Meta downloaded and seeded (“torrented”) the LibGen dataset via BitTorrent despite internal emails acknowledging the optics and legal risk (“torrenting from a corporate laptop doesn’t feel right”), that a memo described LibGen as “a dataset we know to be pirated,” and that CEO Mark Zuckerberg approved its use over AI executives’ concerns. A corporate-representative deposition described scripts built to remove copyright-identifying text before training. The case’s larger stakes: US courts deciding what “piracy” means when the pirate is a Fortune 500 platform building foundation models – and whether statutes written for bootleg DVDs scale to computational copying at Library-of-Congress scale.
The legal architecture matters more than the schadenfreude. Original copyright claims in AI training cases lean on fair use – an open, fact-intensive defense that courts have only begun to shape for ML. The December amendment changes the pressure point: the DMCA’s Section 1202 bans removing or altering CMI, and it carries statutory damages that do not care whether the underlying use was fair. If plaintiffs can show Meta deliberately scrubbed copyright notices to make infringement harder to detect – and the deposition testimony describing purpose-built CMI-stripping scripts is exactly that showing – the case detours around fair use entirely and asks a jury to punish the cover-up. That is a very different lawsuit, and every AI company general counsel understood the difference within hours of the filing’s publication.
The paper trail
| Date | Event |
|---|---|
| 2023-07 | Authors including Kadrey, Silverman, and Coates sue Meta over Llama training data; consolidated class action before Judge Chhabria |
| 2023→2024 | Discovery proceeds; plaintiffs allege Meta obtained LibGen, Anna’s Archive, Z-Library, and Bibliotik caches; Books3 lineage mapped in earlier complaints |
| 2024-04→06 | Per later filings, Meta’s torrenting of additional LibGen data occurs in this window, alongside internal debate over legality and optics |
| 2024-12 mid | Plaintiffs’ reply supporting the third amended complaint is unsealed in redacted form: the “we know to be pirated” memo, CMI-stripping deposition, and Fortune-500-torrenting emails become public |
| 2024-12-17 | Zuckerberg, deposed, concedes the conduct would raise “lots of red flags,” while disputing direct knowledge in key parts |
| 2024-12-25 | This post publishes: motion to amend pending, DMCA and CDAFA claims teed up, industry watching what discovery discipline means for Big AI |
Why the torrenting allegation cuts deepest
BitTorrent is symmetric: to download is to upload. Seeding a shadow-library dataset means distributing copyrighted works to strangers – a supply-side act the copyright statute treats as infringement’s engine, and one no fair-use theory launders, because fair use analyzes the defendant’s own conduct, not the swarm’s. According to the exhibits, Meta’s engineers knew this in real time, debating whether torrenting “legally not ok” data from company hardware was prudent, and the resolution was not to stop but to manage risk – allegedly proceeding anyway, with leadership approval. For a company simultaneously negotiating licensing deals with publishers, the juxtaposition is brutal: one division purchasing rights while another distributed the same works free. And for the DMCA claim, intent does the heavy lifting – plaintiffs quote internal rationale that CMI removal would “reduce the chance that the models will memorise this data,” which reads uncomfortably like concealment dressed as ML hygiene.
The discovery shadow-proceeding
Half the story of December 2024 is procedural misconduct: separate litigation brought by authors against Meta in a different district produced hundreds of relevant documents that plaintiffs say Meta failed to surface in Kadrey discovery – surfaced only when cross-checks caught the omissions, prompting the bad-faith allegations and sanctions-adjacent briefing that accompany the amendment request. If the court credits the pattern, the lesson generalizes far beyond Meta: in an era when every AI lab’s data operations generateinternal paper trails, discovery discipline becomes existential, and courts will treat data-room asymmetries as evidence problems rather than clerical noise. Corporate defendants usually survive copyright cases on legal theory; they implode on documents. The December filings are a preview of how much documentary DNA this litigation will emit before trial.
The signals every AI operator should read
- Provenance is now a legal artifact: “where did this corpus come from” is a question with statutory damages attached; data cards and lineage records transition from best practice to litigation armor.
- CMI stripping is radioactive: scrubbing copyright notices, acknowledgements, or attribution from training corpora creates DMCA 1202 exposure independent of fair use – scripts built for it are exhibits, not tooling.
- Acquisition method is liability: torrenting means distribution; corporate networks distributing pirated caches converts a copying theory into a supply theory with sharper edges.
- Internal candor is discoverable: Slack threads and memos calling data “pirated” will be read aloud in courtrooms; compliance language discipline starts before the lawsuit, not after.
- Licensing late is expensive: deals struck after ingestion read as consciousness of guilt, not mitigation – the sequencing of rights acquisition is now part of the record.
FAQ
What exactly is LibGen, and why does it matter legally?
Library Genesis is a shadow library – a pirate repository hosting millions of books and papers, operating through shifting domains and mirrors since the late 2000s, alongside Anna’s Archive (a meta-search layer over LibGen, Z-Library, and others) and Bibliotik. Legally, shadow libraries are unambiguous: their caches are unauthorized copies. That clarity is precisely why they matter in Kadrey – unlike the murkier Books3 dataset (scraped web text where rights status varies file-by-file), a LibGen download has no colorable innocent-provenance story. If Meta’s models trained on it knowingly, the factual predicate for infringement is easy; the only escapes are fair use (contested) and the DMCA-era defenses (now directly attacked by the CMI claims). Shadow libraries thus function as the case’s ballast: they turn an abstraction – “trained on the internet” – into a specific, seizable, stainable artifact.
Could Meta still win on fair use?
Yes – that is the honest state of play as of late December 2024. Fair use asks whether training on copyrighted books is transformative, whether it supplants the market for the originals, how much was taken, and whether the works were published. Meta will argue model training is a transformative statistical process that does not reproduce books to readers; plaintiffs will argue Llama outputs and market displacement – especially for licensed training-corpus competitors – cut the other way. Judge Chhabria has signaled seriousness about market-effect evidence. But fair use cannot rescue the torrenting distribution allegation or the CMI-stripping claim – those are separate wrongs, and the December amendment’s strategic genius is making the case about them, where fair use never reaches.
What happens next in the case?
Sequence to expect: the court rules on the motion for leave to amend (a low bar legally, but watched closely given the discovery-conduct overlay); if granted, DMCA and CDAFA claims enter the pleading stage with Meta’s motion to dismiss directed at them; parallel discovery disputes over the withheld documents continue, with sanctions briefing; and somewhere beyond, summary judgment on the original copyright claims, where the first real fair-use rulings for LLM training could land. Watch also for coordination effects with the New York Times v. OpenAI and Anthropic music-label cases – courts read each other, and 2025 is positioned as the year US judges start writing the AI-copyright canon. For enterprises, the practical takeaway is already available: assume every training-data sourcing decision will one day be an exhibit.
Legacy: the compliance era begins where the folklore ended
The AI industry’s founding folklore held that the internet was a commons, that ingesting it was inevitable, and that legal questions were legacy friction around an unstoppable future. Kadrey’s December 2024 filings are that folklore’s audit. They depict, in the industry’s own words, decision-makers who knew the data was pirated, weighed the risk, tiled the warnings into memos – and proceeded, allegedly with CEO sign-off, while scrubbing the evidentiary traces. Whatever Judge Chhabria decides about fair use, that portrait has already reset industry behavior: licensing desks expanded, provenance tooling funded, torrent-shaped shortcuts retired. The larger legacy is a phase change in how AI data is governed – from folklore to formation of duties, document by document. The era of “training data happens” is ending; the era of “training data is accounted for” has its first great exhibit, and it is a Fortune 500 company’s own email, reading exactly like a Fortune 500 company’s email should never read.
