Sarah vs. the Machines
The copyright war is entering its second act: training, acquisition and substitution are becoming three different legal battles.
Sarah vs.
the Machines
The copyright war is entering its second act. The question is no longer simply whether AI can learn from copyrighted work. It is who controls the corpus, who can prove economic harm—and who gets paid.
AI companies are winning the argument that machines can learn from copyrighted work. They are not winning the argument that they can steal the library first.
The caseboard
The industry is not fighting one copyright question. It is fighting three—and the cases are splitting along those lines.
Copyright is becoming two businesses
Liability
- Piracy and unauthorized copying
- Litigation and statutory damages
- Regurgitation and substitution
- Unpriced legacy exposure
Asset
- Clean, licensable corpora
- Rights and ownership metadata
- Authenticated access
- Freshness, scarcity and indemnity
The battlefield moved
Sarah Silverman sued OpenAI after alleging that her memoir, The Bedwetter, had been copied without authorization into training datasets. Three years later, that question almost looks quaint. Her case now sits inside a consolidated New York fight involving authors, publishers and The New York Times. What began as a claim about one writer’s book has become a forensic examination of how the modern AI corpus was built.
Two federal rulings supplied the industry’s emerging doctrine. In Bartz v. Anthropic, Judge William Alsup treated model training as transformative fair use on the record before him—but separated that use from the creation of a permanent library assembled with pirated books. In Kadrey v. Meta, Judge Vince Chhabria ruled for Meta on a thin record while emphasizing the plaintiffs’ failure to show meaningful market dilution.
Dispatch analysis: Training, acquisition and output are becoming distinct legal products. A developer may have a strong defense for what a model learned and still face enormous exposure for how the source material entered the pipeline.
The $1.5 billion warning
Anthropic’s settlement turns provenance from an abstract compliance concern into a line item. The lesson is not that every training use requires a royalty. It is that a fair-use destination does not necessarily cleanse an infringing route. The acquisition record—purchase, license, scrape or pirate archive—may determine the bill.
Where the money is moving
The market is already rewarding clean access. Shutterstock’s Data, Distribution and Services revenue reached $203.3 million in 2025, up from $137.3 million in 2023. Reddit reported $140 million of “other revenue” in 2025, a category that includes content licensing but is not broken out separately. These are not standardized royalties. They are early evidence of a wholesale market for legally usable human knowledge.
Dispatch analysis: The premium asset is becoming copyright plus access plus provenance plus machine-readable rights. Large catalogs and data intermediaries gain leverage; individual creators risk retaining rights without negotiating scale.
The five-second liability test
The path matters as much as the destination.
Documented access and rights
Permanent-library exposure
Fair-use argument strengthens
Market-harm argument strengthens
That path may determine liability.
Who’s winning?
Training doctrine is moving their direction—on specific records.
Clean catalogs are becoming strategic, defensible inputs.
Rights verification and corpus cleansing become infrastructure.
Rights remain; negotiating scale and proof of harm lag.
The increasingly dangerous link in the training-data chain.
What Dispatch is watching
Does another court clearly separate training from acquisition?
Does a court finally quantify AI-driven market substitution?
Does music establish a different licensing precedent because its rights systems are already institutionalized?
Do clean training datasets begin appearing as material, separately disclosed corporate revenue?
Does provenance become standard diligence for model developers, investors and insurers?
Dispatch AI Rights Ledger
A continuously maintained record of the legal, licensing and provenance exposure embedded in major AI models.
| Model / company | Corpus | Source status | Rights status | Lead matter | Exposure |
|---|---|---|---|---|---|
| OpenAI | Books1 / Books2; news | Disputed | Mixed / unknown | S.D.N.Y. MDL | ●●●○ |
| Anthropic | Books | Purchased + pirate libraries | Mixed | Bartz | ●●●● |
| Meta | Books | Disputed | Disputed | Kadrey | ●●○○ |
| Stability AI | Images | Web-scraped / alleged | Contested | Getty | ●●●○ |
Illustrative framework, not an investment or legal-risk rating. Proposed inputs: corpus source confidence, rights coverage, active claims, adverse orders, output similarity, market-substitution evidence, indemnity and licensed-data share.
Primary record
Bartz v. Anthropic: June 2025 fair-use order.
Kadrey v. Meta: June 2025 summary-judgment order.
U.S. Copyright Office: Copyright and Artificial Intelligence, Part 3.
Shutterstock: 2025 results.
Reddit: 2025 Form 10-K.
Getty v. Stability AI: U.S. complaint.
OpenAI litigation: OpenAI filing index and case statements.