World's largest open library calls for volunteers to scan and preserve physical books as AI companies buy, scan, and destroy them — Anna's Archive says ‘time is running out’ as ‘knowledge is permanently monopolized on private servers’

Wait 5 sec.

Members of the world's largest 'truly open library in human history' have issued a worldwide call to volunteers, begging them to scan and upload books online to prevent AI companies from obtaining and destroying huge quantities of books. The Anna's Archive appeal follows multiple reports that large AI companies are buying up books to train AI models, often resulting in the destruction of the works. According to the blog post, the problems started when Anthropic was hit with a $1.5 billion penalty as it settled a copyright infringement lawsuit from 2024. This settlement is the largest amount ever in a copyright case, with the AI tech firm paying $200 per title based on its collection of 7 million pirated books stored in a central library. However, that same ruling affirmed that the use of existing works to train AI models is fair use. This opened the floodgates for AI firms, which started buying up books online to scan and destroy them.Independent bookstores all across Europe are seeing a trend like this, where they receive random, relatively large orders that are shipped to local addresses. While orders like these still come through in the digital age, especially from institutional buyers looking to fill out new libraries, they say that most “legitimate” orders often come with coordination and negotiation, not just a straight purchase order. Aside from that, the suspicious orders also include books that no one would buy today, like “Pass Your Driving Test, 2018 Edition” and “The Eddie Hobbs Guide to your SSIA.” An investigation revealed that some of the books are sent to an Amazon processing facility, where workers cut off their spines and feed them into industrial scanning machines.Unfortunately, this method of scanning destroys the books, meaning the printed record is erased in favor of a digital one. There are non-destructive methods of scanning these materials, like using a V-shaped scanner that follows the natural curve of the book binding, but it’s likely that this method is slower and more expensive compared to just stripping out the spine and feeding it into an automatic scanning machine. The blog also says that destroying the original material means that competitors could not use these very same books to train their AI models, giving the original buyer an advantage, especially if the book that is destroyed is a rare one with no other copies in the world, and that it also avoids legal risks.Using these books, especially those published before 2022, which are said to be “untouched by machines,” comes with legal and ethical issues, but the bigger concern of many is that the destruction of this material could lead to a monopoly of knowledge. If the original copy no longer exists, the knowledge stored in it would only be available to the AI company that scanned it, and anyone else who wants to gain access to it would have to pay for the privilege unless the company shares the original for free (although this would come with its own legal issues). Aside from that, it would probably only be available in processed form if access is limited through the AI model, as most LLMs will not produce it verbatim for fear of copyright restrictions, which is why the author said, “Knowledge is permanently monopolized on private servers.”This fear has led to the callout for volunteers to scan books and upload them to Anna’s Archive. “If every person scans a book, and there are 10 million volunteers worldwide, we can obtain 10 million pieces of invaluable wealth,” the blog post said. It also added that uploaders of small scans are often recognized and awarded lifetime membership to the shadow library. Those who want to conduct large-scale scans and upload many titles can reach out to the archive for support as the archive says that it “can help pay for the scanning fees and other rewards.”“This is a race against time,” u said in the Anna's Archive blog post. “Our ideal is to scan and upload all the world’s publications before publishers completely block knowledge, and before AI companies scan and destroy all the world’s books and papers.”