AI firms targeting London’s rare book shops in ‘dystopian’ hunt for training data
London’s rare book shops appear to have been targeted by AI companies searching for books to train their models, after court documents revealed Anthropic bought, scanned and destroyed millions of physical books as part of a controversial project to feed an insatiable demand for training data.
One London-based seller of rare books told City PM they received two unusually large enquiries for specialist books from anonymous email addresses, with no names or company details attached.
Both buyers asked for additional photographs of the books before placing sizeable orders. When the bookseller asked who they were and where the books were going, neither replied.
The enquiries mirror reports across Europe, where antiquarian booksellers say they have received similar requests for thousands of obscure titles from anonymous buyers.
404 Media reported this week that second-hand booksellers had seen a sharp rise in bulk purchases they suspect are being orchestrated by AI companies, with the buyers’ identities typically concealed.
Social media users have dubbed the practice ‘dystopian.’
Slicing off spines and scanning the pages
Recent court documents showed Anthropic, the Claude developer, launched an internal programme known as ‘Project Panama’ after concluding books were “essential” for training advanced AI models.
The AI giant said it taught systems “how to write well”, rather than relying on what one co-founder described as “low-quality internet speak.”
To build that dataset, Anthropic bought millions of second-hand books, sliced off their spines using industrial cutting machines, scanned every page and recycled the rest.
According to Futurism, internal planning documents also showed executives wanted to keep the programme out of the public eye, stating: “We don’t want it to be known that we are pursuing this project.”
A US federal judge later ruled the scanning process was protected by fair use because Anthropic legally bought each book, converted it into a digital copy and destroyed the original, instead of creating additional copies.
The company separately agreed a $1.5bn (£1.13bn) settlement over claims relating to its earlier use of pirated digital books.
‘AI company destroys two million books’
The ruling appears to have accelerated a new market for printed books as AI companies look beyond the internet for training material. As 404 Media first reported, suppliers are now sourcing between 1,000 and one million books at a time for AI developers, with older printed works particularly prized because they pre-date the flood of AI-generated content online.
ISBNdb, one company now offering bulk book-buying services to AI labs, dubbed printed books as “the world’s best AI training data”.
It said books published before 2022 are “structurally guaranteed” to be free of AI-generated text and calls for strict non-disclosure agreements to protect the identity of buyers.
It said books published before 2022 are structurally guaranteed,
The company also acknowledges the reputational risk surrounding the practice, stating on its website: “‘AI company destroys two million books’ is not a headline that generates sympathy.”
