Modelwire
Subscribe

Microsoft internally flagged AI scraping as massive labor theft while harvesting paywalled content

Illustration accompanying: Microsoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings reveal

Unsealed litigation documents expose internal Microsoft communications characterizing large-scale web scraping as systemic labor appropriation, while both Microsoft and OpenAI simultaneously harvested paywalled New York Times content to train models. The filings reveal that insiders predicted this data acquisition would economically devastate publishers. This disclosure crystallizes a core tension in foundation model development: the reliance on copyrighted material without consent or compensation, now documented in corporate correspondence. The case signals that data sourcing practices face mounting legal and reputational risk, forcing labs to reckon with the sustainability of their training pipelines.

Modelwire context

Analyst take

The more significant detail isn't that scraping happened, it's that a Microsoft executive named it as theft in writing, which transforms what was previously a legal theory advanced by plaintiffs into a characterization that originated inside the defendant's own organization. That internal framing is the kind of document that tends to follow a case through appeals.

Modelwire has no prior coverage directly connected to this litigation, so this story sits largely on its own in our archive. It belongs to a broader thread running through the AI industry around data provenance and compensation, a thread that includes publisher licensing negotiations, the emergence of data marketplace startups, and ongoing copyright suits from authors and visual artists. The New York Times case is the highest-profile instance of that pattern, and these filings materially strengthen the plaintiff's position by introducing the defendant's own cost-benefit framing as evidence.

Watch whether Microsoft or OpenAI move to settle the Times case within the next six months, because a settlement before trial would suggest they've assessed the internal communications as too damaging to let a jury see. If the case proceeds to trial with these filings intact, expect other pending publisher suits to accelerate their discovery requests using the same document strategy.

This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.

MentionsMicrosoft · OpenAI · New York Times

MW

Modelwire Editorial

This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.

Modelwire summarizes, we don’t republish. TechCrunch - AI originally reported this story as Microsoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings reveal”. The full content lives on techcrunch.com. If you’re a publisher and want a different summarization policy for your work, see our takedown page.

Microsoft internally flagged AI scraping as massive labor theft while harvesting paywalled content · Modelwire