OpenAI staff knew the ‘existential threat’ AI posed to publishers, New York Times claims
Unlock the Editor’s Digest for free
Roula Khalaf, Editor of the FT, selects her favourite stories in this weekly newsletter.
OpenAI copied millions of copyrighted articles in an effort to build technology potentially worth “gazillions”, despite recognising the “existential threat” it posed to publishers, according to a new filing in a lawsuit brought by The New York Times.
The newspaper and other media outlets are seeking billions of dollars in damages in a case which hinges on whether OpenAI and co-defendant Microsoft infringed copyright law by training their models on the publishers’ content.
Lawyers for The New York Times said OpenAI was driven by commercial incentives to harvest its material. OpenAI’s co-founder Greg Brockman wrote in around 2017 that he was “deeply motivated by the gazillions” he hoped to gain by commercialising OpenAI’s technology, they claimed.
The New York Times’ filings on Thursday introduce new evidence in one of the most prominent cases among dozens of lawsuits filed by copyright holders against AI groups, accusing them of pirating materials to train models that threaten the jobs of human content creators.
OpenAI has argued that using publicly available content to train its models is “fair use” under copyright law, echoing the legal justification employed by search engines for accessing content. The company has said its large language models transform the original content, rather than replicating it, and that it does not provide a substitute to The New York Times or other publications.
But lawyers for the newspaper claim the “defendants repeatedly copied millions of . . . copyrighted articles in their entirety without permission to produce substitutive commercial AI products”.
The lawyers added OpenAI did so despite knowing their AI tools could replace the underlying material. The AI lab’s head of ChatGPT wrote that publishers faced an “existential threat” from AI products that “are largely substitutive, period [and] will get more and more substitutive as they get better”, according to Thursday’s filing.
The lawsuit also quotes Microsoft’s director of applied science, Brent Hecht, who allegedly wrote that training LLMs on copyrighted content was “an astonishing theft of unprecedented proportions” and that winning the case would “make a complete mockery of the idea of ‘fair use’”.
Microsoft found click-through rates to underlying content were 83 per cent to 93 per cent lower for users of its AI answer tools compared with its traditional search engine.
“These comments reflect one employee’s individual perspective, are not a legal analysis,” Microsoft said in a statement.
The filing also claims OpenAI employed “transgressive and deceptive” techniques to acquire paywalled content, and that Brockman, now OpenAI’s president and de facto number two executive, had been aware of the practice.
When he was informed by an employee about “a hack to get around NY Times paywall”, Brockman responded “ah nice”, according to the filing.
In his testimony for the case, Satya Nadella, Microsoft’s chief executive, said that if he “had been made aware that OpenAI had scraped and trained on information that was behind a paywall”, he would have invoked Microsoft’s rights to require OpenAI “to retrain its models”, under the two companies’ agreement at the time.
Microsoft on Thursday said Nadella’s testimony “spoke to broad principles and changes under way in how people find and consume information” and was not a “conclusion about copyright questions” at the centre of the case.
OpenAI did not immediately respond to requests for comment. The AI lab has previously said: “This case is not about human authorship versus AI. It’s about The New York Times looking for an undeserved payday at the expense of progress that benefits everyone.”
The FT in 2024 signed a licensing deal and partnership with OpenAI.
