Newly public documents in the copyright lawsuit that The New York Times filed against OpenAI and Microsoft three years ago have revealed the companies’ internal assessments of methods for collecting content to train artificial intelligence and their impact on publishers. A significant portion of the documents is based on The Times’ filing, while the supporting exhibits remain sealed.
According to the filing, Microsoft Applied Sciences Director Brent Hecht described data-collection practices for artificial intelligence in January 2023 as “staggering labor theft at an unprecedented scale” and “the greatest labor theft in human history.” Nick Turley, the OpenAI executive responsible for ChatGPT, wrote that products such as chatbots pose an “existential threat” to publishers, are largely substitutive, and will become even more substitutive as they improve.
Microsoft data showed that Copilot’s “answer engine” feature reduced click-through rates to The New York Times domain by as much as 93% compared with traditional Bing search. In a presentation dated January 2024, Hecht described this as a “doom loop” that could damage both model performance and the internet at the same time. In his testimony this year, Satya Nadella also said that content behind a paywall must be licensed if it is to be used for training or grounding.
The documents state that OpenAI’s intermediate training datasets contained more than 91,692 copies of works published by NYT, Daily News and the Center for Investigative Reporting. A dataset sourced from Common Crawl included more than 2 million documents from nytimes.com alone. The dataset created under Project Mango was alleged to contain copies of at least 160,903 unique works belonging to news publishers.
The court filing alleged that the companies collected content through the Bing Index, that OpenAI employees developed plans to bypass paywalls without being detected, and that copyright notices were removed from training data to prevent them from appearing in model outputs. OpenAI and Microsoft did not respond to requests for comment.
Why it matters
The documents show that the lawsuit is not limited to whether copyrighted materials were used without permission; the impact of AI products on publishers’ reader traffic and revenue model is also at the center of the dispute. In particular, internal assessments suggesting that systems providing direct answers instead of search results could reduce the need to direct users to news websites may strengthen the economic rationale behind publishers’ licensing demands. However, many of the claims in the filing are presented in The New York Times’ petition, and the fact that the supporting exhibits are confidential currently limits independent examination of the scope of the data and the methods used to collect it. The companies’ lack of comment also means that, at this stage, the narrative largely relies on documents submitted by one of the parties to the lawsuit.
Background
Microsoft is not a new name in the FikirPilot archive: we published 15 stories mentioning the name in the past 90 days; the most recent is dated 19 September 2026.