Meta Faces US Authors in Landmark AI Copyright Battle

As Meta aggressively invests billions to establish itself as a global AI leader, the case could shape not only its future but also the rules governing AI development across the industry.

2 mins read
File Photo of Mark Zuckerberg

In a major legal showdown that could reshape the future of artificial intelligence and copyright law, Meta will face a group of prominent US authors in court on Thursday over allegations that it used pirated books to train its AI models.

The case, brought by a coalition of authors including Ta-Nehisi Coates and Richard Kadrey, challenges Meta’s alleged use of content from LibGen—a shadow library hosting millions of books, academic texts, and comics without permission from rights holders—to train its Llama family of AI models. At stake is whether the $1.4 trillion tech giant’s reliance on such data constitutes “fair use,” or a violation of copyright.

As Financial Times reports, this case marks one of the earliest and most closely watched legal tests of generative AI’s data-hungry development practices, with wide-reaching implications for the future of creative rights and technological innovation.

The authors, supported by the Authors Guild, argue that their work was used without consent or compensation, and that Meta deliberately exploited the availability of pirated content to avoid paying license fees. Mary Rasenberger, CEO of the Authors Guild, criticized the tech industry’s disregard for creators’ rights: “AI models have been trained on hundreds of thousands if not millions of books, downloaded from well-known pirated sites—this was not accidental.”

Court filings reveal that Meta initially held discussions with publishers about licensing but allegedly abandoned those talks after identifying LibGen as a free source. An internal Meta email uncovered during discovery reads, “If we license one single book, we won’t be able to lean into the fair use strategy.”

Other emails show internal concern within Meta over the legality of their approach. In one exchange, Joelle Pineau, then head of Meta’s FAIR AI lab, recommended using LibGen data. Another email from director of product Sony Theakanath, flagged with the subtitle “legal risk,” explicitly stated, “In no case would we disclose publicly that we had trained on LibGen.” Portions discussing “copyright and IP” were redacted.

Meta maintains its position that training large language models (LLMs) on copyrighted content falls under fair use, especially when the technology being developed is transformative in nature. “We disagree with the plaintiffs’ assertions, and the full record tells a different story,” the company said in a statement. “We will continue to vigorously defend ourselves and to protect the development of GenAI for the benefit of all.”

The court case comes amid a flurry of similar lawsuits against tech firms like OpenAI, Microsoft, and Anthropic, all of which face legal scrutiny over the data sources used to build AI tools like ChatGPT and Claude.

Legal experts say the outcome could set precedent. “There is a tremendous amount of uncertainty right now,” said Chris Mammen, partner at Womble Bond Dickinson. “It is extremely important to get these things resolved. Things are going to continue happening in the world at the breakneck pace that technology and our economy are developing.”

The plaintiffs also allege Meta used torrenting—a peer-to-peer file sharing method—to obtain the LibGen database. While Meta is said to have attempted to restrict redistribution, evidence shows that some outbound data may have been shared inadvertently, with portions of related logs now deleted.

Sri Lanka Guardian

The Sri Lanka Guardian is an online web portal founded in August 2007 by a group of concerned Sri Lankan citizens including journalists, activists, academics and retired civil servants. We are independent and non-profit. Email: editor@slguardian.org

Leave a Reply

Your email address will not be published.

Latest from Blog