Newly unredacted filings reveal Microsoft executive Brent Hecht described the company's mass scraping of journalism as "an astonishing theft" and "the largest theft of labor in human history." Internal documents and OpenAI comments warn models are "largely substitutive," creating an alleged "doom loop" that drove up to 93% fewer referrals to The New York Times. Plaintiffs allege paywall circumvention, stripped copyright notices, and a Project Mango dataset with at least 160,903 unique news works. The filings push toward licensing solutions rather than only legal defenses.
Unredacted Filings: Microsoft Exec Called AI Training “The Largest Theft Of Labor” — Paywall Hacks, a “Doom Loop” and a 93% Drop in Referrals

Newly unredacted court filings in The New York Times' lawsuit against OpenAI and Microsoft reveal blunt internal language from both companies about how they built large language models — language that company spokespeople would never use in public.
Key Revelations From The Filings
“An astonishing theft of unprecedented proportions,” wrote Brent Hecht, Microsoft’s Director of Applied Science, in a January 2023 internal memo, later adding that the mass scraping of journalism amounted to the “largest theft of labor in human history.”
Those words are not plaintiff rhetoric alone; they appear in Microsoft documents unredacted in filings made public on September 17, 2026, in a suit joined by the New York Daily News and the Center for Investigative Reporting.
Microsoft internal presentations and analytics foresaw a feedback loop they labeled a “doom loop”: AI answer engines reduce referrals to news sites, weakening newsroom revenue and capacity, which in turn degrades the high-quality journalism the models rely on.
OpenAI’s Internal Assessment
OpenAI managers and researchers also used stark language in private. Nick Turley, head of ChatGPT, warned publishers faced an “existential threat” from products he called “largely substitutive.” Greg Brockman, OpenAI’s president, is quoted as saying the models are “excellent at news.” A researcher allegedly told Brockman about a “hack to get around [the] nytimes paywall,” to which Brockman replied, “ah nice.”
Data, Paywalls, And Allegations
Plaintiffs allege OpenAI bypassed paywalls to obtain content and that copyright notices were stripped from training datasets to reduce the chance models would reproduce source attributions. Microsoft internal data cited in the filings reportedly showed Copilot’s answer engine delivered up to 93% fewer referrals to The New York Times domain compared with traditional Bing search.
Project Mango, an alleged collaboration between Microsoft and OpenAI, is said to have assembled a training dataset containing at least 160,903 unique news works from the publishers at the center of the litigation.
Public Defense Versus Private Language
Publicly, both companies defend their training practices as protected fair use and transformative. Privately, executives used terms like "theft" and acknowledged the substitutive effect on news consumption. Microsoft CEO Satya Nadella has testified that chatbot conversations can substitute for visiting original publisher sites and that paywalled content used for training or grounding should be licensed.
A federal government brief recently backed OpenAI’s fair-use position, adding a political dimension to what courts will decide. Microsoft has tried to distance corporate policy from individual memos, calling Hecht’s remarks personal views.
What This Means
The filings, as presented by plaintiffs, point toward licensing frameworks and new business models as the practical remedies rather than purely legal defenses. When a company’s own internal record describes its practices as theft and predicts substitutive harms, plaintiffs argue, the enduring solution will likely be commercial agreements that compensate content creators and publishers.
Disclosure: The reporting summarized here is based on unredacted court filings and plaintiffs’ exhibits; some materials remain sealed and full context for certain quotes or exhibits may not yet be publicly available.
Help us improve.



























