Tech News

Microsoft executive called AI training on copyrighted content 'largest theft of labor'

Unredacted filings in The New York Times copyright lawsuit reveal Microsoft and OpenAI executives privately likened AI scraping to theft.

Ceeve Team · 2026-09-17 · 2 min

A neutral summary of Microsoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings reveal, published by TechCrunch. Summarised for the Ceeve reading list; the original reporting is TechCrunch's.

Unredacted court filings in The New York Times' copyright lawsuit against OpenAI and Microsoft reveal that a top Microsoft executive privately described the companies' AI training practices as "theft" and "the largest theft of labor in human history." OpenAI's own leadership stated internally that AI models posed an "existential threat" to publishers and journalists whose work trained them. The documents also detail how the companies allegedly bypassed paywalls undetected, built training datasets through mass scraping, and deliberately removed copyright notices from training data.

A Microsoft document from January 2024 written by director of Applied Science Brent Hecht describes the declining click-through rates from news sources as a "doom loop," noting it was "highly unusual" for an end-product to threaten the economic foundations of its essential suppliers. Internal communications from OpenAI's head of ChatGPT described publishers as facing an "existential threat" from products that are "largely substitutive." The filings reveal OpenAI's training datasets alone contained over 91,692 copies of works from The New York Times, Daily News, and the Center for Investigative Reporting, with one Common Crawl dataset containing over 2 million documents from nytimes.com.

The documents describe how OpenAI employees allegedly circumvented paywalls without detection and deliberately stripped copyright notices from training data. The scale involved Project Mango data containing at least 160,903 unique works from news publishers, and researchers allegedly built datasets like WebText2 that disproportionately relied on scraped news content. Those seeking roles in AI governance or publisher-focused tech should track how these admissions affect fair-use arguments in ongoing litigation and industry standards through resources like Ceeve.

Read the original on TechCrunch


Ceeve tailors your CV to a specific role, writes the cover letter that goes with it, and prepares you for the interview. If a story here is about a company you would like to work for, Ceeve can research it, match your experience against the posting, and rehearse the questions with you before you walk in.