New Delhi: A Microsoft executive described the mass collection of online content to train artificial intelligence models as an “astonishing theft” and potentially the “largest theft of labour in human history”, according to recently unsealed court filings in the copyright case involving The New York Times, Microsoft and OpenAI.
The remarks were made by Brent Hecht, Microsoft’s director of applied science, in an internal 2023 memo. The comments have come to light as part of a broader legal battle over whether AI companies can use copyrighted books, news articles and other online material to train large language models without obtaining permission or paying the creators.
The filings also reveal internal discussions at Microsoft and OpenAI about the possibility that AI chatbots could reduce traffic to news publishers by providing information directly to users, rather than sending them to the websites where the original articles were published.
Microsoft executive raises concerns over AI scraping
In the memo cited in the court filings, Hecht warned about the scale of online content being collected for AI training.
He wrote that millions of people could view large AI models “hoovering up” their work as an “astonishing theft of unprecedented proportions”. He also noted that much of the content had not been created with the expectation that it would be used to train AI systems and that creators were generally not being compensated for such use.
Hecht’s comments were internal observations and do not represent a judicial finding that AI training constitutes theft. The underlying legal question remains contested in the US courts.
The comments are significant because Microsoft is one of OpenAI’s major commercial partners and has invested heavily in the company and its AI products.
Copyright dispute remains at the centre
The dispute centres on the use of copyrighted material in training AI models.
The New York Times sued OpenAI and Microsoft in December 2023, alleging that millions of its articles were copied and used to develop AI systems without permission. Other news organisations, including the New York Daily News, The Intercept and the Center for Investigative Reporting, have also brought claims against the companies.
Microsoft and OpenAI have defended their practices under the US fair-use doctrine. Their position is that training AI models on copyrighted material can constitute a transformative use rather than simply reproducing the original works.
The publishers dispute that argument. They contend that AI products can provide information that substitutes for the original reporting, potentially depriving news organisations of readers, advertising revenue and subscriptions.
The question of whether particular uses of copyrighted material qualify as fair use remains for the courts to determine.
AI chatbots could replace visits to news websites
The newly unsealed documents also contain internal discussions about how consumers access news.
Microsoft CEO Satya Nadella acknowledged during a deposition that conversations with chatbots had substituted for users receiving information directly from websites. Instead of clicking on a news article, users can ask an AI system a question and receive a summary or answer within the chatbot.
The news organisations argue that this creates a difficult economic cycle.
Publishers invest money in journalists and produce the articles that provide much of the information used to train AI systems. At the same time, AI-powered search and chatbot products can potentially reduce the number of people who visit those publishers’ websites.
According to figures cited in the court filings, click-through rates for The New York Times and Daily News domains were between 83 per cent and 93 per cent lower on Microsoft’s Copilot “answer engine” than on traditional Bing Search.
The figures are part of the plaintiffs’ court submissions and are being used by the news organisations to support their argument that AI products can compete directly with publishers.
OpenAI executive described products as substitutive
The documents also contain comments from OpenAI executives about the potential impact of AI products on publishers.
Nick Turley, OpenAI’s head of ChatGPT, described the company’s products as “largely substitutive” in internal communication, according to the unsealed filings. He also said they would become increasingly substitutive as the technology improved.
OpenAI co-founder Greg Brockman separately described the company’s models as particularly capable at predicting the text of news articles and performing news-related tasks.
The comments are being highlighted by the publishers as evidence that the companies understood the potential commercial impact of AI systems on the news industry.
Microsoft and OpenAI have maintained their legal defence, arguing that their use of training material is protected by fair use.
Publishers face a changing digital landscape
The dispute goes beyond the question of whether copyrighted articles can be used to train AI models.
Traditional search engines generally direct users to the websites that originally published information. AI chatbots can instead provide an answer directly, potentially reducing the need for users to click through to the original source.
That difference has become increasingly important for news organisations, which rely on website traffic to generate advertising, subscriptions and other revenue.
The newly disclosed documents show that executives and employees at both companies were discussing this potential disruption internally even as AI products were being developed.
Microsoft documents cited in the litigation also warned of a possible “doom loop” in which AI systems could undermine the economic foundations of the websites whose content helps supply information for future AI models.
Court battle could shape AI training rules
The case between The New York Times, Microsoft and OpenAI is part of a much wider legal debate over the use of copyrighted material to develop generative AI.
Authors, publishers, artists and other creators have filed lawsuits against several AI companies, arguing that their work was used without consent or payment. AI companies have generally argued that training models involves transformative processing and that existing copyright law permits such use in appropriate circumstances.
The newly unsealed documents do not by themselves settle those legal questions. Instead, they provide evidence that the court will consider alongside the companies’ formal legal arguments, technical evidence and the specific circumstances surrounding the use of copyrighted material.
For publishers, the case also raises a broader question about whether the internet’s existing content-based economic model can survive if users increasingly obtain information from AI systems rather than visiting the original websites.
As the litigation proceeds, the court’s eventual decisions could have implications for how AI companies obtain training data, how creators are compensated and how publishers distribute their work in an increasingly AI-driven internet.
