Why Would Meta Download So Much Porn?

Wait 5 sec.

Editor’s note: This work is part of AI Watchdog, The Atlantic’s ongoing investigation into the generative-AI industry.Amid tech companies’ ongoing efforts to vacuum up as much data as possible to train AI models, Meta has taken public posts from its own platforms, scraped massive amounts of content from the rest of the internet, and pirated millions of books to train its AI models. Now the company is being accused of downloading much more illicit material: huge volumes of porn, nonconsensual celebrity nudes, blueprints for 3-D-printable handguns, and millions of passwords acquired by hackers.Strike 3, the parent company of Vixen Media Group, a prolific producer of adult films, is suing Meta for allegedly downloading 2,973 of its copyrighted videos. But in its efforts to collect evidence about these purported downloads, Strike 3 also captured information relating to many other files that Meta may have acquired over the course of two years: They include images from “Celebgate,” a 2014 hack that resulted in the leak of private photos belonging to Jennifer Lawrence, Kirsten Dunst, and others, plus several collections of deepfake celebrity porn featuring the faces of Natalie Portman, Scarlett Johansson, Elizabeth Olsen, Gal Gadot, and others.Strike 3’s file-transfer lists include ordinary movies and TV shows, alongside dozens of videos from GirlsDoPorn, which was shut down after six people associated with the site were charged with sex trafficking. Meta allegedly downloaded several individual episodes from the website and two “GirlsDoPorn MegaPack” archives. Strike 3 also presents evidence that Meta made some content publicly available in addition to downloading it.When I reached out to Meta, a spokesperson told me via email that “these claims are bogus.” Meta also noted in a court filing that Strike 3, which regularly files lawsuits against people who download its work illegally, “has been labeled by some as a ‘copyright troll.’”This is a complicated case. Piracy can be hard to track with precision, and both parties are clearly conflicted: Strike 3, in seeking damages, wants Meta to look bad, and Meta would prefer not to be associated with this kind of content. Yet if Strike 3’s data are accurate, they would show that participating in mass piracy is a by-product of modern tech development—whether or not any of the material collected is explicitly used to engineer AI models or other programs.[Read: The hypocrisy at the heart of the AI industry]Although Meta acknowledged the possibility that employees, contractors, or even visitors to Meta’s offices may have used the company’s networks to download the porn in question, its defense in court was that the material was accessed for “personal consumption” rather than as part of a company project. But neither the sheer volume nor the pattern of downloads looks like personal consumption. The file-transfer logs produced by Strike 3 show sequences of files that are unrelated except for a certain keyword or concept. For example, two files that were transferred consecutively on June 2, 2023—“Bajillion Dollar Properties S01E07” and “First Time Home Buyer Anal Fantasy”—are both loosely themed around real estate, but the similarities likely end there. Judge Eumi Lee, who denied Meta’s motion to dismiss the case, cited other such juxtapositions, such as consecutive downloads of “Teenage Mutant Ninja Turtles (1987-1996)” and a video labeled “Teen Sex Sessions 2 (2012).”Strike 3 has argued that the content could be used for AI training. A Meta spokesperson told me, “We don’t want this type of content, and we take deliberate steps to avoid training on this kind of material.” In its legal defense, Meta pointed out that its terms of service prohibit users from using its AI products to generate images containing pornography. Even if Meta isn’t building a porn-generating bot, however, tech companies can use “unsafe” content to test guardrails for their systems—effectively telling the software what not to generate. (Lee noted that Meta’s terms of service are irrelevant to the question of whether it wants to acquire porn.)It’s also possible that Meta is compiling an archive of material for no precise purpose, just as Anthropic acquired millions of books that it says it had no intention of using for AI training but wanted to keep for its “research library.” As AI-training techniques advance, there is no telling what kind of content a company might find useful in the future.Nineteen of the allegedly downloaded files contain models of functional handguns that can be produced with 3-D printers. A file labeled “5.7 million passwords list” purportedly contains 5,718,107 passwords from accounts that were hacked from 2015 to 2019. The downloads also include password-cracking tools that can be used by hackers to break into people’s accounts. Other files include the movies BlacKkKlansman, The Banshees of Inisherin, and everything made by Studio Ghibli from 1979 to 2020; music by Dua Lipa, Elton John, and Paul McCartney; radio shows and podcasts such as The Howard Stern Show and 1619; and software including Microsoft Windows, Adobe Photoshop, Ableton Live, and NBA 2K23.The files were downloaded through BitTorrent, a system that allows people to share files among themselves. When you download files via BitTorrent, you typically share those files with other people at the same time; if nobody is sharing the files, then nobody can download them. According to Strike 3’s data, which lawyers from the company told me it regularly gathers in an effort to track the illegal distribution of its intellectual property, devices on Meta’s networks also made some of Strike 3’s videos available to others—a practice known as “seeding”—for as long as three months after fully downloading them.Meta additionally used BitTorrent to acquire large quantities of copyrighted books around the same time, as shown in testimony from another lawsuit. Some Meta employees were uncomfortable with this activity. According to court documents, Jelmer van der Linde, a former Meta employee who was asked to download books to train the company’s AI models in 2024, told his supervisor in a work chat message that it seemed like a “very very dark grey area. Legally, ethically …” He added that torrenting “copyrighted material is not very legal either.” He was transferred to a different project, and another developer downloaded the books instead. He has since left Meta and now works at the Ellison Institute of Technology in Oxford.[Read: The unbelievable scale of AI’s pirated-books problem]To verify Strike 3’s data, I reached out to Tom Chothia, a professor of cybersecurity at the University of Birmingham who has studied BitTorrent. He reviewed a technical description of Strike 3’s data-collection methods and told me that, assuming the records weren’t tampered with, “they have established that people at Meta were uploading” Strike 3’s videos.One thing that complicates the case is that most of the torrenting did not occur over Meta’s corporate networks. Strike 3’s data show sequences of related file transfers across several networks, including Meta’s. Strike 3 argues that the other networks were used remotely in a coordinated way, perhaps to make the activity harder to trace. Meta has called the patterns a coincidence. But Lee found them compelling, writing that Meta’s coincidence theory “strains belief” and that the activity is likely indicative of “algorithmically coordinated behavior.”Nearly 300 of the transfers were traced to a private residence in Mountain View, California. Strike 3 claims that this was the home of a Meta contractor’s father, and that the file transfers stopped when the contractor stopped working for Meta. In response, Meta argued that these downloads were “plainly indicative of personal consumption.” According to Strike 3’s logs, the files downloaded to the residence include, in addition to porn, Russian-language books, Chinese movies, and a wide variety of cracked software, among other things.Unfortunately, in the AI era, illegal mass-downloading by tech companies is not uncommon. Other lawsuits have revealed that OpenAI and Anthropic have also used BitTorrent to download copyrighted books. Whether they have uploaded those books as well, or what other content they may have downloaded through BitTorrent, is not yet clear.BitTorrent is a strange world in which to find respectable megacorporations. It is dominated by media piracy, and very little legitimate file-sharing occurs. But if AI models are intended to replace human creators, then they have to be trained on as much human-created content as possible. BitTorrent is perhaps the fastest and cheapest way to download high-quality work in large quantities. AI companies have claimed that their models can create original work, but so far, the development of artificial intelligence has involved a lot of theft of other people’s intelligence.