Originally published bySlashdot
An anonymous reader quotes a report from 404 Media: As AI companies search for more training data to improve their models, one company is offering old, printed books as an ideal source because they are guaranteed to be free of the very AI slop AI companies are producing. "The world's best AI training data is sitting on a shelf," ISBNdb, a company that produces what it claims is "the world's largest book database," and that offers high-volume book acquisition services for AI companies, says on its site. "Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative."
In one article on its site, ISBNdb explains that printed books published before 2022 are ideal for AI training data because they don't include AI generated text. As the article correctly notes, much of the data that AI companies can scrape from the internet today is likely to include AI generated text, which could result in "model collapse," a process by which AI models that are trained on AI generated data results in worse models that are more prone to errors. The article also notes that book authors who object to their writing being scraped for training purposes can now easily poison AI models by producing writing designed to manipulate and sabotage the resulting AI models.
"Print books from the pre-LLM era are structurally guaranteed to be free of this contamination. That alone is a significant advantage [...] "Physical books published before this date [pre-2022] are structurally clean of modern poisoning tools." [...] ISBNdb advertises that it can keep the identity of AI companies secret. "Strict NDA [non-disclosure agreement] on every engagement," ISBNdb's site says. "Every project begins with a legally binding non-disclosure agreement. Your identity, strategy, and acquisition targets are never disclosed." ISBNdb notes that AI companies may not want to be caught destroying printed books during the scanning process. "The optics problem is real," ISBNdb's site says. "'AI company destroys two million books' is not a headline that generates sympathy."
Read more of this story at Slashdot.
πΊπΈ
More news from United StatesUnited States
NORTH AMERICA
Related News
My Local AI Assistant Got Worse When I Remembered Too Much
16h ago
RedHook malware turns on your phone's Wireless Debugging to stream your screen β and it never touches the consent dialog
16h ago
Naming Things Without Pain
17h ago
Running Docker on Proxmox: LXC vs VMs and the Firewall Rules That Actually Matter
15h ago
How to Query Databricks from Salesforce Apex (Without Copying a Billion Rows)
16h ago