Torrenting from Corporate Laptops: Meta Emails Unsealed
- Teh controversy surrounding Meta's AI training practices deepens as allegations of using pirated materials surface.
- Authors assert that meta was aware its AI models were trained using pirated books.
- According to the authors, emails discussing torrenting activities demonstrate that Meta was conscious of the "illegal" nature of their actions.
Meta‘s AI Training Under Scrutiny: Piracy Allegations Intensify
Table of Contents
Teh controversy surrounding Meta’s AI training practices deepens as allegations of using pirated materials surface. Authors in a lawsuit claim Meta knowingly utilized the notorious piracy database, LibGen, to train its artificial intelligence models. This dataset allegedly contains millions of pirated works, distributed through peer-to-peer torrents.
Authors assert that meta was aware its AI models were trained using pirated books. New evidence reportedly indicates that Meta employed LibGen, a dataset rife with copyright-infringing material, for AI training. This data was allegedly accessed and distributed via torrents.
According to the authors, emails discussing torrenting activities demonstrate that Meta was conscious of the “illegal” nature of their actions. Warnings were seemingly ignored, with evidence suggesting Meta actively concealed its torrenting activities while downloading and seeding terabytes of data from various shadow libraries as recently as April 2024.
Meta Allegedly Concealed Seeding Activities
To obscure their activities, Meta purportedly avoided using Facebook servers for downloading the dataset.This was allegedly done to “avoid” the “risk” of anyone “tracing back the seeder/downloader” from Facebook servers. An internal message from meta researcher Frank Zhang described the operation as being in “stealth mode.”
Furthermore, meta allegedly adjusted settings to minimize seeding. Michael Clark, a Meta executive in charge of project management, stated in a deposition that settings were modified “so that the smallest amount of seeding possible could occur.”
Key Allegations Against Meta
- Use of LibGen, a known piracy database, for AI training.
- Conscious effort to conceal torrenting activities.
- Minimizing seeding to avoid detection.
With the emergence of new information, authors are demanding further depositions from Meta staff involved in the decision to torrent LibGen. They argue that these new facts “contradict prior deposition testimony.”
For instance, mark Zuckerberg reportedly claimed no involvement in decisions regarding the use of LibGen for AI model training. However, unredacted messages allegedly reveal that the “decision to use LibGen occurred” after “a prior escalation to MZ,” according to the authors.
Meta has not yet responded to requests for comment. Throughout the litigation, the company has maintained that AI training on LibGen constitutes “fair use.”
In a motion to dismiss filed last month, Meta addressed its torrenting activities, stating that “plaintiffs do not plead a single instance in which any part of any book was, in fact, downloaded by a third party from Meta via torrent, much less that Plaintiffs’ books were somehow distributed by Meta.”
Despite Meta’s confidence in its legal strategy, the new torrenting revelations have seemingly complicated its case. Authors can now expand their distribution theory, which is crucial for winning a direct copyright infringement claim, beyond merely asserting that Meta’s AI outputs unlawfully distributed their works.
As limited revelation on Meta’s seeding proceeds, the company is not currently contesting the seeding aspect of the direct copyright infringement claim. Meta intends to “set… the record straight and debunk… this meritless allegation on summary judgment.”
“Plaintiffs do not plead a single instance in which any part of any book was, in fact, downloaded by a third party from Meta via torrent, much less that Plaintiffs’ books were somehow distributed by meta.”
The Core of the Dispute: Copyright Infringement
The central issue revolves around whether Meta’s use of copyrighted material from LibGen constitutes copyright infringement. Authors argue that Meta’s actions go beyond fair use, while Meta maintains its training practices are legally sound.
The lawsuit highlights the complex legal and ethical questions surrounding AI training and the use of copyrighted data. The outcome could have significant implications for the future of AI progress and copyright law.
Key Dates
- April 2024: Alleged continued downloading and seeding of data from shadow libraries.
- Last Month (February 2025): Meta files a motion to dismiss.
- January 9, 2025: Reuters reports on Meta’s knowledge of using pirated books.
- March 13, 2025: Current Date
Meta’s AI Training Under Scrutiny: Q&A on the LibGen Controversy
The use of copyrighted material to train AI models is a hot-button issue. Meta, one of the world’s largest technology companies, is facing intense scrutiny over its AI training practices. Allegations have surfaced that Meta knowingly used pirated materials from the LibGen database. Let’s delve into the details of this controversy through a Q&A format.
Key Questions About the Meta AI Training controversy
What is the central allegation against Meta?
The core allegation is that Meta used the LibGen database, a known source of pirated books, to train its AI models. This has led to accusations of copyright infringement by authors who claim their works were used without permission.
What is LibGen?
libgen,short for Library Genesis,is a file-sharing website that provides access to copyrighted content,mainly books and articles,without the permission of the copyright holders. It’s often described as a shadow library or piracy database. According to [3], LibGen is basically like PirateBay.
Did Meta know it was using pirated books for AI training?
Authors in the lawsuit allege that Meta was aware its AI models were trained using pirated books from LibGen.They point to internal emails discussing torrenting activities. these emails allegedly demonstrate that Meta was conscious of the “illegal” nature of their actions. Warnings were seemingly ignored, with evidence suggesting Meta actively concealed its torrenting activities while downloading and seeding terabytes of data.
How did Meta allegedly conceal its torrenting activities?
To obscure their activities, Meta purportedly avoided using Facebook servers for downloading the dataset. This was allegedly done to “avoid” the “risk” of anyone “tracing back the seeder/downloader” from Facebook servers. An internal message from a Meta researcher described the operation as being in “stealth mode.” Furthermore, Meta allegedly adjusted settings to minimize seeding, further concealing their activity.
What is “seeding” in the context of torrenting?
In torrenting, “seeding” refers to uploading downloaded files to other users in the network. by minimizing seeding, Meta allegedly attempted to reduce the risk of detection and limit the distribution of copyrighted material from their servers.
What is Meta’s defense against these allegations?
Meta’s primary defense is that its use of copyrighted material from LibGen falls under “fair use.” They argue that training AI models is transformative and does not infringe on the rights of copyright holders. In a motion to dismiss
