AI companies are reportedly shredding books after using them to train AI models
AI companies are purchasing millions of secondhand books through intermediaries to source high-quality training data, then destroying the books after scanning. The practice has raised ethical concerns about the loss of rare books and the privatization of knowledge.
14
Join the conversation
Follow us
Add us as a preferred source on Google
Having contributed to the growing shortage of memory and storage, AI companies seemingly have a new target in their sights: humanity's literary history. A recent investigative report from 404 Media reveals that these companies are reportedly purchasing millions of secondhand books through intermediaries to source high-quality training data for their AI models, avoiding public backlash.
Go deeper with TH Premium: AI shortages
(Image credit: Nvidia)
AI data centers are swallowing the world's memory and storage supply
Chip scarcity assaults auto industry amid the worsening Nexperia and DRAM crisis
Samsung and SK hynix shorten memory contracts as pricing power shifts back to suppliers
Memory makers are set to earn $551 billion from the AI boom
AI relies on vast amounts of data to advance, but not just any data. It has to be high-quality data. The problem is that mediocre AI-generated content, commonly referred to as "AI slop," has proliferated across the Internet. This type of content contaminates the data pool and is counterproductive for AI to train on. As a result, leading AI companies have turned to human-authored sources for knowledge, specifically print sources that predate 2022 and are more likely to contain original, uncontaminated content.
There is precedent for AI companies turning to physical books for training AI. For instance, Anthropic, one of the leading AI companies involved in a lawsuit, reportedly invested millions of dollars in extracting information from countless printed books to build its Claude AI models and then destroying them. The company bought books from Better World Books. Although the court decision affirmed that using books for AI training falls under fair use in copyright law, Anthropic faced a staggering $1.5 billion fine for maintaining a repository of seven million pirated books that infringed the copyrights of authors and publishers. Similarly, a coalition of publishers recently filed a lawsuit against Google, accusing the tech giant of allegedly and illegally using millions of copyrighted books to develop its Gemini AI models.
Latest Videos FromTom's Hardware
Watch full video here:
ISBNdb, an online database that reportedly has over 111 million cataloged books, has been a long-favorite platform for booksellers, libraries, and distributors to sell books. With the explosion of the AI industry, ISBNdb has pivoted its business to offer specialized services to bulk-purchase books for AI companies. According to 404 Media, the orders range from 1,000 copies to as many as one million books in a single transaction.
One professional bookseller, who wanted to remain anonymous, purportedly spoke to 404 Media about the unprecedented surge in book sales, which began in April of this year. The seller previously moved around 20 books in a good week, but in recent months, weekly sales have skyrocketed to several hundred books. It represents a fivefold increase over the normal volume. Other booksellers on platforms such as Alibris and Biblio have reported similar spikes in bulk purchases.
While there is no concrete proof that ISBNdb or some other AI company is making the purchase, there are some red flags. Notably, the large-scale purchases only included books with an International Standard Book Number (ISBN), the unique 13-digit code used globally to identify books. There were no patterns in terms of subject, genre, or author. It also appeared that the purchasers disregarded the pricing for the books and snapped up titles at any cost, even if they were overpriced.
During the Anthropic lawsuit, Tom Harvey, who previously participated in the creation of Google Books before leading Anthropic's "Project Panama" digitalization project, confirmed that the AI firm hired several document scanning companies. Datamation Information Services, which offers high-volume, non-destructive, and destructive book scanning services, was one of them. The former method employs different tools, like overhead scanners, flatbed scanners, or V-shaped imaging systems. The latter method, on the other hand, would have personnel gut the books and feed the individual pages into a high-speed industrial scanner. Logically, AI companies opt for the destructive route since it is more efficient and lower-cost. The result is the destruction of millions of books.
Obviously, printed books represent a treasure trove of information for AI. However, many debate the ethics of removing books from circulation since it is uncertain whether AI companies filter the rare or even out-of-print books from the common titles during digitalization. The other major issue is that scanned books go directly into a private database to train AI, which the general public does not have access to. True, we will have smarter AI, but at the cost of the information not being available to future generations.
Follow Tom's Hardware on Google News, or add us as a preferred source, to get our latest news, analysis, & reviews in your feeds.
Reply
Do we really want Stephen King as the basis for training AI without developing context into the AI?
:ROFLMAO:
Reply
Reply
Reply
Reply
Reply
I don't understand why they'd destroy the books. :/
This situation reminds me of a t-shirt I saw recently: "Make Orwell fiction again."
Inventory space for unused assets. That way, they don't even need a room to store them, and it gets cheaper. Always money.
Well, sellers could create a different type of business: lend the book, then do a reverse-logistics to receive it back. Charge more for it, as it seems they don't care about costs. Book sellers keep inventories while turning a profit, preserve hard to replace books, AI companies don't look so evil, and everybody wins. I guess it's too late for that, but still, would solve many problems.
Reply
Do we really want Stephen King as the basis for training AI without developing context into the AI?
Actually: yes. AI's have to know what "popular culture" is to comment on it at all.
What is REALLY funny is ask an AI about most any a Stephen King novel and it will agree it's pure trash, even if entertaining, or great literature. It all depends on how you ask the question.
Reply
I don't understand why they'd destroy the books. :/
I see nothing wrong with this.
You're allowed to buy a book and keep it for reference without any violation of copyright laws. What you're not allowed to do is make copies except for personal use. An AI does the same thing, but it has to have the 'book' digitized in a database...that's what its 'memory' is after all... in order to reference it and make fair-use quotes for servicing 'clients'.
That digitized copy is probably a copyright violation that will get them in trouble. What I don't know is whether or not destroying the original copy of copyrighted material they legally purchased relieves them of the copyright violation for making digital copies (I am not a lawyer). Especially to use as AI's do.
Also, look at the recent Anthropic settlement so see what's at play here. There were at least two parts: one is they got pirated material (never good) but the part that got the 1.5B dollars is they kept copies (digitized in the AI's memory/database) and used them.
But at any rate it's no different from me reading the latest NY Times best seller on summer vaca then ripping pages out to start a fire in the cabin stove this winter. My library will not take donated books, they'd just toss them in a dumpster; they only want books they obtain and are in line with their collection management policies.
Reply
Do we really want Stephen King as the basis for training AI without developing context into the AI?
At least they're using the books and not the movies. I would hate to see what a Stephen King-inspired meatball recipe looks like!
Reply
Show more comments