Microsoft exec called AI scraping 'the largest theft of labor in human history'
Recorded: Sept. 18, 2026, 10 a.m.
| Original | Summarized |
Microsoft exec called AI scraping ‘the largest theft of labor in human history,' new unredacted filings reveal | TechCrunch Disrupt 2026: OpenAI, Anthropic, Replit, and more take over 6 industry stages. 25% off tickets now Back by popular demand: Save up to $300 on Disrupt Close TechCrunch Desktop Logo TechCrunch Mobile Logo LatestStartupsVentureAppleSecurityAIAppsDisrupt 2026 EventsPodcastsNewsletters SearchSubmit Site Search Toggle Mega Menu Toggle Topics Latest AI Amazon Apps Biotech & Health Climate Cloud Computing Commerce Crypto Enterprise EVs Fintech Fundraising Gadgets Gaming Government & Policy Hardware Layoffs Media & Entertainment Meta Microsoft Privacy Robotics Security Social Space Startups TikTok Transportation Venture More from TechCrunch Staff Events Startup Battlefield StrictlyVC Newsletters Podcasts Videos Partner Content TechCrunch Brand Studio Contact Us Image Credits:Justin Sullivan / Getty Images AI
Microsoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings reveal Rebecca Bellan 12:46 PM PDT · September 17, 2026
New unredacted information in the copyright lawsuit The New York Times brought against OpenAI and Microsoft three years ago reveals an admission that AI scraping was tantamount to theft, and that AI products pose a major threat to publications. The unsealed material also details how the companies allegedly obtained and used that content by bypassing paywalls undetected, building training datasets via mass scraping, and deliberately stripping copyright notices from training data. For example, Microsoft’s own data shows its Copilot “answer engine” caused click-through rates for The New York Times’ domain to drop as much as 93% compared to traditional Bing search. An internal Microsoft presentation written by Microsoft’s director of Applied Science, Brent Hecht, in January 2024 describes the decline as a “doom loop” that would “hurt the performance of our models and the entire web at the same time.” Other admissions cut against different pillars of the fair-use test: OpenAI’s head of ChatGPT, Nick Turley, wrote in internal communication that publishers face an “existential threat” from products like the chatbot, which are “largely substitutive” and “will get more and more substitutive as they get better.” The companies allegedly assembled the Project Mango data into a training dataset that contains copies of at least 160,903 unique works from the news publishers. Topics AI, copyright, Government & Policy, Microsoft, new york times, OpenAI When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.
Rebecca Bellan Senior Reporter Rebecca Bellan is a senior reporter at TechCrunch where she covers the business, policy, and emerging trends shaping artificial intelligence. Her work has also appeared in Forbes, Bloomberg, The Atlantic, The Daily Beast, and other publications. View Bio October 13 – 15 Last day to book an exhibit table is September 18. Don’t miss out on high-impact leads, investor access, and a brand spotlight in Disrupt’s Expo Hall. BOOK NOW Most Popular OpenAI caught its models leaving notes to successors to hide bad behavior Rebecca Bellan Clean tech startup Fluxnium found a way to tap 50,000 years’ worth of nuclear fuel Tim De Chant Salesforce and Nvidia’s new reasoning model is everything the AI labs should fear Julie Bort Jensen Huang took a call from Trump, and showed off something else, too Connie Loizos The 9 buzziest startups from Y Combinator’s latest Demo Day, according to VCs Marina Temkin Tesla says it will finally unveil the second-generation Roadster on October 1 Anthony Ha Revolut confirms customer data breach through fake government requests Jagmeet Singh Loading the next article Error loading the next article X TechCrunchStaffContact UsAdvertiseSite Map © 2026 TechCrunch Media LLC. |
New unredacted filings stemming from the copyright lawsuit brought by The New York Times against OpenAI and Microsoft reveal significant admissions regarding the process of training artificial intelligence models using copyrighted material. A top Microsoft executive privately characterized these AI training practices as "theft," and leadership at OpenAI acknowledged that the resulting AI models pose an "existential threat" to the publishers and journalists whose work was used for training. The documents detail how the companies allegedly obtained and utilized this content by circumventing paywalls without detection, constructing training datasets through mass scraping, and deliberately omitting copyright notices from the training data. These admissions challenge the prevailing legal argument for fair use, which often permits the use of copyrighted work without permission in certain contexts like criticism or parody. However, the new information suggests that the actions taken by the AI firms may run counter to the requirement that such use must not substitute or harm the market for the original work. For instance, Microsoft's internal data indicated that its Copilot answer engine caused drop in click-through rates for The New York Times domain to as much as ninety-three percent compared to traditional Bing search, leading to an internal assessment of a "doom loop" that would negatively impact model performance and the broader web. Microsoft CEO Satya Nadella also testified that any content behind a paywall should be licensed for use in training or grounding models, stating that he would have compelled OpenAI to retrain its models had he been aware that information was being scraped from paywalled sources. Further admissions undermined the fair use defense, as OpenAI personnel indicated that the technology directly competes with, rather than merely transforming, the original works. This was supported by a Microsoft document predicting a significant disruption to the employment of the individuals who generated the data used to train the foundation models. The sheer scale of the alleged copying is detailed in the filings, revealing that OpenAI's mid-training datasets reportedly contained over ninety-one thousand six hundred ninety-two copies of works published by the NYT, Daily News, and the Center for Investigative Reporting. Furthermore, a Common Crawl-derived dataset utilized by the companies included more than two million documents from nytimes.com alone. The text further outlines the methods used, noting that employees developed strategies to bypass paywalls undetected and constructed training sets such as WebText and WebText2 based on scraped news content. There was also an alleged deliberate effort to remove copyright notices from the training data to prevent the model outputting them to users. One internal memo referenced this activity as "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history." The filings ultimately present evidence that OpenAI and Microsoft were aware that their acquisition methods were ethically questionable. |