OpenAI has been accused by several US news organizations of concealing and destroying evidence related to how ChatGPT was trained on copyrighted news content, as the legal battle over artificial intelligence and copyright intensifies in federal court and litigation costs for one of the plaintiffs exceed $28 million.
The allegations were filed on Thursday in the US District Court in Manhattan by media organizations including The New York Times and the New York Daily News. The plaintiffs are asking a federal judge to impose sanctions on OpenAI, claiming the company obstructed the discovery process by withholding evidence that could be central to determining how its artificial intelligence systems were developed using millions of news articles.
According to the court filing, the plaintiffs allege that OpenAI “chose obstruction” instead of producing datasets and ChatGPT logs that could reveal how copyrighted news content was incorporated into the company’s AI systems. The filing also claims that testimony provided during a recent deposition by an OpenAI employee contradicts the company’s previous statements regarding its ability to search training datasets and system logs for copyrighted material.
Steven Lieberman, an attorney representing the New York Daily News and seven affiliated newspapers, alleged that OpenAI had made misrepresentations over the past two years concerning its ability to identify copyrighted content within its training data. Lieberman said the motion asks the court to penalize OpenAI for “hiding and destroying evidence” related to what he described as journalism used to train ChatGPT.
The case forms part of a broader legal dispute over whether artificial intelligence companies can lawfully use copyrighted material to develop AI systems without authorization. At the center of the litigation is the question of whether AI chatbots compete directly with news publishers by providing information generated from journalistic content while reducing traffic to the original news sources.
The New York Times filed its lawsuit against OpenAI and Microsoft in late 2023, approximately one year after the public launch of ChatGPT. Since then, additional media organizations have joined the case, including MediaNews Group, the parent company of the Chicago Tribune, digital publisher Ziff Davis, and the nonprofit Center for Investigative Reporting.
According to the source text, concerns among news publishers increased further in 2024 after Google introduced AI-generated summaries at the top of search results, reducing the number of readers clicking through to original news websites and affecting advertising revenue.
OpenAI and other artificial intelligence companies have argued that using digitized books, online articles, and other publicly available material to train AI models is protected under the fair use doctrine of US copyright law. That legal argument is currently being tested in numerous lawsuits brought by authors, artists, music companies, and other copyright holders against AI developers.
The source text notes that OpenAI competitor Anthropic recently agreed to pay $1.5 billion to book authors in what it describes as the largest copyright settlement involving AI training to date. The lawsuit brought by The New York Times differs from claims made by authors, focusing instead on allegations of unfair competition by companies that profit from journalism without obtaining permission or compensating publishers.
According to regulatory filings cited in the source text, The New York Times has spent more than $28 million pursuing legal action against AI companies, including a separate lawsuit filed last year against Perplexity. As part of the latest motion, the plaintiffs are seeking sanctions that include reimbursement of attorney fees associated with obtaining evidence they allege was improperly withheld during discovery.
The dispute comes as an increasing number of media organizations have entered into licensing agreements with OpenAI and other technology companies, including Google and Meta, allowing artificial intelligence developers to train models using licensed news archives and content in exchange for payment.

