AI Firms Reportedly Buy and Destroy Print Books for Training Data
In one line: A report examines why AI companies may be buying print books in bulk, tearing off their bindings to scan them, then destroying the copies to build training datasets.
Key points
- Local outlet The Clarion-Ledger looks at why AI firms would purchase physical books only to scan and discard them.
- The books are reportedly cut apart into loose pages, then run through high-speed scanners to produce digital text.
- The apparent goal is to secure clean, high-quality source material with lower copyright-litigation risk than scraped web data.
Why it matters
The provenance and legality of data used to train large language models is at the center of recent copyright litigation. Scanning lawfully purchased books may offer a stronger legal footing than unauthorized web crawling, yet destroying the physical copies raises separate questions about waste and ethics.
Read more
- Why AI companies may be buying books from MS stores and destroying them — The Clarion-Ledger