GreekReporter.comTechnologyHow Anthropic's Secret Project Destroyed Millions of Books to Train AI

How Anthropic’s Secret Project Destroyed Millions of Books to Train AI

Getting your Trinity Audio player ready...
Library books in shelves
Court records reveal how Anthropic destroyed millions of books for AI training. Credit: Terry Kearny, Flickr, Public Domain

Anthropic destroyed millions of books after scanning their pages to create digital copies used in the development of its AI models.

Details became public after records were unsealed in a copyright lawsuit filed by authors. The documents offered a rare look into how Anthropic collected the written material used to develop models behind products such as Claude.

After hiring former Google executive Tom Turvey, who had worked on Google Books partnerships, Anthropic significantly expanded the project in early 2024. Documents obtained by The Washington Post revealed that the company spent tens of millions of dollars purchasing physical books, while internal planning records showed it sought to keep the initiative hidden from public view.

Dario Amodei, the CEO and co-founder of Anthropic
Dario Amodei, the CEO and co-founder of Anthropic. Credit: TechCrunch / Wikimedia Commons / CC BY 2.0

Books were especially valuable because they contained professionally edited, human-written content. According to one of Anthropic’s co-founders, this material could help AI systems generate higher-quality writing than models trained primarily on the often lower-quality content available online.

How Anthropic destroyed millions of books for AI training

To assemble its collection, Anthropic purchased used books in large quantities, with some acquisitions involving tens of thousands of titles at a time. The company obtained books from suppliers such as Better World Books and the British retailer World of Books.

Anthropic also explored potential partnerships with libraries and bookstores, including New York’s Strand Book Store. However, Strand later stated that it had not sold any books to the company.

The Strand bookstore in Greenwich Village, NYC
The Strand bookstore in Greenwich Village, NYC. Credit: Ajay Suresh / Wikimedia Commons / CC BY 2.0

After acquisition, the books underwent destructive scanning, a process that involved dismantling the physical copies to extract their contents. One contractor proposed scanning between 500,000 and two million books over a six-month period. Workers used hydraulic cutting machines to remove the spines, allowing the loose pages to be processed through high-speed scanners. The discarded paper was then collected and recycled.

Anthropic’s purchases reflected a broader demand for printed books as training material for AI systems. Companies such as ISBNdb have marketed services aimed at helping AI developers acquire large collections of physical books. Its services include locating books, managing large-scale purchases, and maintaining buyer confidentiality through non-disclosure agreements.

Bulk book demand reaches independent sellers

The surge in demand has also affected independent booksellers. Citing reporting from 404 Media, Decrypt reported that one seller saw weekly sales rise from roughly twenty books to several hundred after buyers seeking material for AI development entered the market.

Although the increased demand provided a financial boost, the seller expressed concerns that rare and out-of-print books could vanish once they were scanned and destroyed, potentially removing valuable works from circulation.

The move toward acquiring physical books came after earlier attempts to access large digital collections. Court documents revealed that Anthropic co-founder Ben Mann downloaded books from LibGen, an online archive known for distributing copyrighted works without authorization.

Mann later shared details about another unauthorized repository, the Pirate Library Mirror. Anthropic stated that material obtained from LibGen was not used to train a commercial model that generated revenue. The company also said that the second collection was not used to train a full AI model.

Courts separate AI training from unauthorized acquisition

The distinction between these methods of obtaining books became central to legal proceedings. In June 2025, US District Judge William Alsup ruled that Anthropic’s use of books to train its AI models could fall under fair use because the training process transformed the original material into a new type of work. Alsup also determined that Anthropic could keep digital copies after destroying millions of books it had legally purchased.

Nonetheless, the ruling did not resolve allegations involving illegally obtained digital copies. Instead of proceeding to a trial over those materials, Anthropic reached a $1.5 billion settlement without admitting liability. Under the agreement, eligible authors and publishers were expected to receive approximately $3,000 per affected title.

Copyright cases expand as preservation concerns grow

Court records indicate that Anthropic was part of a broader effort among AI companies to acquire extensive book collections. Court records indicate that employees at Meta discussed using LibGen as a potential source for training Llama 3, although some raised concerns about the legal and ethical implications.

OpenAI has also acknowledged downloading material from LibGen but stated that the files were removed before the release of ChatGPT. Meanwhile, copyright lawsuits involving companies such as Meta, OpenAI, Microsoft, and Google are continuing through the courts.

ChatGPT Artificial Intelligence
ChatGPT, an OpenAI AI product. Credit: focal5 / CC BY-NC 2.0 / Flickr

The controversy has also sparked broader concerns about the preservation of physical books. The Internet Archive took a different approach by maintaining original printed copies while providing digital versions for lending. Publishers successfully challenged the Internet Archive’s controlled digital lending program, leading the nonprofit to remove more than 500,000 books from its lending service.

Elon Musk has likewise argued that rare books used in AI projects should be preserved and digitized without being dismantled in the scanning process.

See all the latest news from Greece and the world at Greekreporter.com. Contact our newsroom to report an update or send your story, photos and videos. Follow GR on Google News and subscribe here to our daily email!



National Hellenic Museum
Filed Under

More greek news