Newly unsealed court documents reveal that executives and employees at Microsoft and OpenAI privately raised concerns about using millions of news articles to train artificial intelligence systems, including ChatGPT. One Microsoft executive described the practice as the “largest theft of labour in human history”, while an OpenAI executive warned that AI products posed an “existential threat” to publishers.
The disclosures emerged from an ongoing copyright lawsuit brought by The New York Times against Microsoft and OpenAI. The case, filed in 2023, centres on allegations that the companies used copyrighted news content without permission to develop and train large language models. The companies have argued that the we legal doctrine of fair use protects their use of the material.
Internal concerns over AI training and the publishing industry
The newly released material includes comments from Brent Hecht, Microsoft’s director of applied science. In internal documents, Hecht reportedly expressed concern about the scale at which AI companies were collecting online content. He described the practice as “the largest theft of labour in human history” and warned that it could create a “doom loop” that would damage both publishers and the AI systems relying on their work.
Another Microsoft document from 2023 warned that millions of people could eventually view large AI models “hoovering up” their work as “an astonishing theft of unprecedented proportions”. Hecht also described large AI models as “a product that destroys its supply chain”, referring to AI systems’ dependence on content produced by publishers, journalists and other creators.
The documents also show concerns within OpenAI about ChatGPT’s potential impact on traditional journalism. Nick Turley, who led the ChatGPT team, wrote that AI represented an “existential threat” to publishers. In another internal communication, he said AI products were “largely substitutive” and would become increasingly so as the technology improved.
Other internal messages raised similar concerns about whether AI systems were directing users towards the sources of information. An OpenAI software engineer wrote in 2023 that, “no matter how prominently we show the links, users won’t click.” The comments are significant because Microsoft and OpenAI have argued that their AI products can provide transformed information while still supporting the broader publishing ecosystem.
Court documents detail how content was obtained
The unsealed filings also allege how OpenAI and Microsoft obtained material for AI training. According to the documents, the companies and their employees used large collections of online content and discussed methods for accessing material that was protected by publisher paywalls. The filings further allege that copyright notices were removed from some training data.
One exchange involving OpenAI president Greg Brockman has attracted particular attention. According to the court documents, an employee told Brockman about developing a “hack to get around the NYTimes paywall”. Brockman responded, “ah nice.” The publishers have cited the exchange as evidence that the companies were aware of questions about access to protected material.
The filings also describe the scale of the datasets involved. TechCrunch reported that the documents identify more than 91,000 copies of works from publishers involved in the litigation within OpenAI’s mid-training datasets. In contrast, a Common Crawl-derived dataset contained more than 2 million documents from nytimes.com alone. The filing also describes projects through which Microsoft and OpenAI exchanged or assembled training data.
Microsoft chief executive Satya Nadella has separately stated in a deposition that “anything that is paywalled should be licensed by anyone who wants to use it”. He also said that if he had known OpenAI had trained on paywalled information, he would have used Microsoft’s contractual rights to require the company to retrain its models. Microsoft has maintained that using copyrighted material for AI training can qualify as transformative use under copyright law.
Microsoft and OpenAI maintain their legal positions
Microsoft has sought to distance the company from some of the internal comments disclosed in the documents. A Microsoft spokesperson, Alex Haurek, said the statements in Hecht’s internal documents did not represent the company’s position in the litigation.
“Microsoft’s position is set out in its court filings, which explain why these transformative uses are consistent with copyright law and why Copilot is not a substitute for publishers’ journalism,” Haurek said.
The distinction matters because the lawsuit ultimately turns on whether using copyrighted material to train AI systems is lawful under US copyright rules. Microsoft and OpenAI argue that the training process transforms the material and that AI-generated responses are not simply replacements for the original articles. The publishers, meanwhile, argue that the systems were built using their work without permission or compensation and can compete directly with the sources that produced the material.
The case is being considered alongside broader legal disputes over using copyrighted material to train generative AI systems. The outcome could influence how technology companies obtain and use large collections of written material, as well as how publishers approach licensing agreements with AI developers. The newly unsealed documents provide additional evidence of internal discussions, but they do not, by themselves, determine whether the companies’ conduct violated copyright law.
The legal dispute also highlights a wider tension in the AI industry. AI developers require enormous amounts of information to train increasingly capable models, while publishers rely on their original reporting to generate traffic, advertising and subscription revenue. As AI tools increasingly answer questions directly rather than simply directing users to external websites, the relationship between AI companies and content creators is becoming a central issue in the technology and media industries.




