LmCast :: Stay tuned in

Microsoft exec called AI scraping 'the largest theft of labor in human history'

Recorded: Sept. 18, 2026, 10 a.m.

Original Summarized

Microsoft exec called AI scraping ‘the largest theft of labor in human history,' new unredacted filings reveal | TechCrunch

Disrupt 2026: OpenAI, Anthropic, Replit, and more take over 6 industry stages. 25% off tickets now

Back by popular demand: Save up to $300 on Disrupt

Close

TechCrunch Desktop Logo

TechCrunch Mobile Logo

LatestStartupsVentureAppleSecurityAIAppsDisrupt 2026

EventsPodcastsNewsletters

SearchSubmit

Site Search Toggle

Mega Menu Toggle

Topics

Latest

AI

Amazon

Apps

Biotech & Health

Climate

Cloud Computing

Commerce

Crypto

Enterprise

EVs

Fintech

Fundraising

Gadgets

Gaming

Google

Government & Policy

Hardware

Instagram

Layoffs

Media & Entertainment

Meta

Microsoft

Privacy

Robotics

Security

Social

Space

Startups

TikTok

Transportation

Venture

More from TechCrunch

Staff

Events

Startup Battlefield

StrictlyVC

Newsletters

Podcasts

Videos

Partner Content

TechCrunch Brand Studio

Contact Us

Image Credits:Justin Sullivan / Getty Images

AI

Microsoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings reveal

Rebecca Bellan

12:46 PM PDT · September 17, 2026

New unredacted information in the copyright lawsuit The New York Times brought against OpenAI and Microsoft three years ago reveals an admission that AI scraping was tantamount to theft, and that AI products pose a major threat to publications.
Per the lawsuit, a top Microsoft executive privately described the companies’ AI training practices as “theft,” and OpenAI’s own leadership said its AI models posed an “existential threat” to the publishers and journalists whose work trained them. 

The unsealed material also details how the companies allegedly obtained and used that content by bypassing paywalls undetected, building training datasets via mass scraping, and deliberately stripping copyright notices from training data. 
It’s worth noting that much of the new information comes from The Times’ own brief, not the underlying exhibits, which remain sealed. The quotes below are presented without their original context.
The unredacted filing is the latest escalation in the three-year-old lawsuit, in which The New York Times initially alleged the firms violated copyright law by training generative AI models on its content. 
The question of whether AI firms can legally use copyrighted material to train AI has no clear answer, but judges have been largely favorable to AI companies’ arguments that training constitutes “fair use.” This legal rule lets people use copyrighted work without permission in certain cases, like parody, news reporting, or criticism. Earlier this month, the Trump administration contributed a brief in defense of OpenAI’s unlicensed use of copyrighted material to train its LLMs. 
Several of the new admissions, however, run counter to OpenAI’s fair use defense, particularly the rule’s requirement that use doesn’t substitute or harm the market for the original work.

For example, Microsoft’s own data shows its Copilot “answer engine” caused click-through rates for The New York Times’ domain to drop as much as 93% compared to traditional Bing search. An internal Microsoft presentation written by Microsoft’s director of Applied Science, Brent Hecht, in January 2024 describes the decline as a “doom loop” that would “hurt the performance of our models and the entire web at the same time.”
“It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain,’” reads the Microsoft document, as quoted in the filing. 
Microsoft CEO Satya Nadella also testified in a deposition earlier this year that “anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training,” and made clear that, if he “had been made aware that OpenAI had scraped and trained on information that was behind a paywall,” he would have “invoked [Microsoft’s right to] require OpenAI to retrain its models.” 

Other admissions cut against different pillars of the fair-use test: OpenAI’s head of ChatGPT, Nick Turley, wrote in internal communication that publishers face an “existential threat” from products like the chatbot, which are “largely substitutive” and “will get more and more substitutive as they get better.”
OpenAI President Greg Brockman described the models as “excellent at news.” Nadella agreed under oath earlier this year that conversing with chatbots “has substituted … giving you the information right there on the website on the AI platform versus needing to go to the underlying source.”
That kind of language speaks to how the technology could directly compete with, rather than transform, the original work. 
A Microsoft document states that there is a “real risk” that generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained.” 
The sheer scale of the copying is striking. The documents reveal for the first time that OpenAI’s mid-training datasets alone contain more than 91,692 copies of works published by the NYT, Daily News, and Center for Investigative Reporting. A Common Crawl-derived dataset included more than 2 million documents from nytimes.com alone. 
In a January 2023 internal memo, Hecht called it “an astonishing theft of unprecedented proportions” and “the largest theft of labor in human history.”
The filing lays out in new detail how OpenAI and Microsoft went about acquiring the plaintiffs’ content, including scraping it from the Bing Index. 
“OpenAI delivered the entire GPT-3 training dataset to Microsoft, which Microsoft used to evaluate how to implement OpenAI’s models within its own commercial products,” the filing reads. “Microsoft similarly provided training data to OpenAI through initiatives called Project Taxi and Project Mango.”

The companies allegedly assembled the Project Mango data into a training dataset that contains copies of at least 160,903 unique works from the news publishers. 
In order to get the most out of their scraping, OpenAI employees allegedly came up with a plan to circumvent paywalls without detection. The filings show that when OpenAI researcher Nick Ryder told Brockman about a “hack to get around nytimes paywall,” Brockman replied: “ah nice.” 
OpenAI employees also allegedly built training datasets like WebText and WebText2 that disproportionately relied on scraped news content. They also allegedly pulled millions of articles from Common Crawl, a free, open repository of web crawl data. The findings also describe deliberate efforts to strip copyright notices from training data before it reached the model, since researchers “wouldn’t want model outputting” “copyright notices” to users.
“The evidence revealed here for the first time shows that OpenAI and Microsoft knew that what they were doing was wrong,” Steven Lieberman, counsel for the New York Daily News, said in a statement shared with TechCrunch.
OpenAI and Microsoft did not return requests for comment.

Topics

AI, copyright, Government & Policy, Microsoft, new york times, OpenAI

When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

Rebecca Bellan

Senior Reporter

Rebecca Bellan is a senior reporter at TechCrunch where she covers the business, policy, and emerging trends shaping artificial intelligence. Her work has also appeared in Forbes, Bloomberg, The Atlantic, The Daily Beast, and other publications.
You can contact or verify outreach from Rebecca by emailing rebecca.bellan@techcrunch.com or via encrypted message at rebeccabellan.491 on Signal.

View Bio

October 13 – 15
San Francisco

Last day to book an exhibit table is September 18. Don’t miss out on high-impact leads, investor access, and a brand spotlight in Disrupt’s Expo Hall.

BOOK NOW

Most Popular

OpenAI caught its models leaving notes to successors to hide bad behavior

Rebecca Bellan

Clean tech startup Fluxnium found a way to tap 50,000 years’ worth of nuclear fuel

Tim De Chant

Salesforce and Nvidia’s new reasoning model is everything the AI labs should fear

Julie Bort

Jensen Huang took a call from Trump, and showed off something else, too

Connie Loizos

The 9 buzziest startups from Y Combinator’s latest Demo Day, according to VCs

Marina Temkin
Dominic-Madori Davis

Tesla says it will finally unveil the second-generation Roadster on October 1

Anthony Ha

Revolut confirms customer data breach through fake government requests

Jagmeet Singh

Loading the next article

Error loading the next article

X
LinkedIn
Facebook
Instagram
youTube
Mastodon
Threads
Bluesky

TechCrunchStaffContact UsAdvertiseSite Map
Terms of ServicePrivacy PolicyRSS Terms of UseCode of Conduct
OpenAIHugging FaceFlockStartup BattlefieldDisrupt 2026Tech LayoffsChatGPT

© 2026 TechCrunch Media LLC.

New unredacted filings stemming from the copyright lawsuit brought by The New York Times against OpenAI and Microsoft reveal significant admissions regarding the process of training artificial intelligence models using copyrighted material. A top Microsoft executive privately characterized these AI training practices as "theft," and leadership at OpenAI acknowledged that the resulting AI models pose an "existential threat" to the publishers and journalists whose work was used for training. The documents detail how the companies allegedly obtained and utilized this content by circumventing paywalls without detection, constructing training datasets through mass scraping, and deliberately omitting copyright notices from the training data.

These admissions challenge the prevailing legal argument for fair use, which often permits the use of copyrighted work without permission in certain contexts like criticism or parody. However, the new information suggests that the actions taken by the AI firms may run counter to the requirement that such use must not substitute or harm the market for the original work. For instance, Microsoft's internal data indicated that its Copilot answer engine caused drop in click-through rates for The New York Times domain to as much as ninety-three percent compared to traditional Bing search, leading to an internal assessment of a "doom loop" that would negatively impact model performance and the broader web.

Microsoft CEO Satya Nadella also testified that any content behind a paywall should be licensed for use in training or grounding models, stating that he would have compelled OpenAI to retrain its models had he been aware that information was being scraped from paywalled sources. Further admissions undermined the fair use defense, as OpenAI personnel indicated that the technology directly competes with, rather than merely transforming, the original works. This was supported by a Microsoft document predicting a significant disruption to the employment of the individuals who generated the data used to train the foundation models.

The sheer scale of the alleged copying is detailed in the filings, revealing that OpenAI's mid-training datasets reportedly contained over ninety-one thousand six hundred ninety-two copies of works published by the NYT, Daily News, and the Center for Investigative Reporting. Furthermore, a Common Crawl-derived dataset utilized by the companies included more than two million documents from nytimes.com alone. The text further outlines the methods used, noting that employees developed strategies to bypass paywalls undetected and constructed training sets such as WebText and WebText2 based on scraped news content. There was also an alleged deliberate effort to remove copyright notices from the training data to prevent the model outputting them to users. One internal memo referenced this activity as "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history." The filings ultimately present evidence that OpenAI and Microsoft were aware that their acquisition methods were ethically questionable.