Search Beyond News…

Unsealed court filings reveal a Microsoft executive privately called AI training "the largest theft of labor in history" while the company and OpenAI scraped paywalled New York Times content, undermining their public fair-use defense in a landmark copyright case

Executive summary: Newly unsealed court filings in the New York Times copyright lawsuit against OpenAI and Microsoft reveal a Microsoft executive privately called AI training "the largest theft of labor in human history" while both companies scraped paywalled Times content and internally warned the practice would gut publishers. The admission undermines the defendants' fair-use defense, increases exposure to statutory damages, and could force industry-wide licensing agreements for AI training data, raising costs for model developers and creating revenue streams for publishers.

Who is involved: Microsoft, OpenAI, The New York Times, U.S. District Court for the Southern District of New York.

Likely next: Further unsealing of depositions and internal communications; summary-judgment briefing expected in Q4 2026; potential settlement talks if discovery worsens defendants' position.

Newly unsealed documents in the New York Times lawsuit against OpenAI and Microsoft show a Microsoft manager describing the firms' data-collection practices as historic labor theft. The revelation contradicts the companies' public stance that training on copyrighted works is fair use and may strengthen the Times' claim for willful infringement. The case is shaping up as a pivotal test of whether generative AI developers must license training data at scale.

What's next — scenarios

Base: Courts require licensing for copyrighted training data; Microsoft/OpenAI settle with publishers (50%)

New recurring revenue for publishers; AI model training costs rise; compliance infrastructure becomes a competitive moat.

Upside: Fair use upheld; defendants win on summary judgment (25%)

Status quo continues; AI training proceeds without mandatory licensing; legal uncertainty lifts for model providers.

Downside: Statutory damages awarded; injunction forces model retraining on licensed data only (25%)

Billions in damages; months-long delay in AI product releases; scramble to secure licensed datasets.

What to watch

Timeline

Analysis — what this means

Likely next events

Sectors affected

Regulatory implications

Historical parallels

Key entities

Sources

Related cases

Browse the full archive →