Meta is under legal fire for its artificial intelligence (AI) training practices after a U.S. court unsealed documents revealing the company’s use of a controversial piracy database, Library Genesis (LibGen). This revelation comes as part of a copyright infringement lawsuit brought by a group of authors, including Richard Kadrey and Sarah Silverman. The authors allege that Meta trained its generative AI models on their copyrighted works without authorization. The case’s outcome could set a precedent for how tech companies use creative content in AI development.
Judge Slams Meta’s Redactions in Court
U.S. District Court Judge Vince Chhabria heavily criticized Meta’s initial efforts to redact information from the court documents, calling the attempt “preposterous.” The judge noted that Meta appeared more concerned with avoiding bad press than protecting legitimate business interests. As a result, the unsealed documents provide new insights into internal communications where Meta employees expressed concerns about using LibGen’s pirated database for training AI models.
Copyright Infringement Allegations
The authors claim Meta violated copyright laws by training its AI language models with unauthorized access to their works. Meta argues that its use of publicly available data falls under the “fair use” doctrine, which permits limited use of copyrighted material under certain conditions. However, the high-profile nature of the authors involved, including well-known names like Silverman, puts additional pressure on Meta’s defense strategy.
Meta’s Link to LibGen
Previously, Meta admitted to using a dataset called Books3—comprising legally sourced online books—for training its AI models. However, the newly disclosed court documents suggest Meta also obtained data from LibGen, a database infamous for hosting pirated content. Internal emails revealed that some employees questioned the ethics and legal risks of using LibGen, but the company proceeded with the practice.
Broader Implications for AI and Copyright
This lawsuit highlights a larger debate about the legality of using unlicensed content in AI training. As the case unfolds, its resolution could reshape industry standards and legal frameworks for AI development. Judge Chhabria’s criticism of Meta suggests the company will need to exercise greater diligence in sourcing data for future AI projects to avoid additional legal challenges.
The case underscores the growing tension between the tech industry’s demand for vast datasets and the rights of content creators, setting the stage for potentially landmark legal decisions.



