Why Anthropic's Book-Scanning Practice Draws Scrutiny
In one line: The Guardian looks at Anthropic's practice of scanning physical books and then discarding them to build training data, unpacking the context and the controversy.
Key points
- As AI firms need vast amounts of text to train large language models, digitizing physical books by cutting and scanning them has come under scrutiny.
- Scanning reportedly requires slicing off each book's binding, effectively destroying the original — a detail that has fueled symbolic backlash.
- The practice is tied to ongoing debates over copyright and fair use around digitizing lawfully purchased books.
Why it matters
With the provenance and acquisition methods of AI training data now central to copyright litigation and regulation, physically discarding books highlights where technical justification collides with cultural sentiment. Transparency in how training data is sourced is likely to remain contested.
Read more
- Why is Anthropic destroying books? — The Guardian
How this story unfolded
- OWASP Publishes 2026 Top 10 Security Risks for LLM Applications
- Alibaba Debuts Its 'Most Powerful' AI Model to Date
- Author Gets $6,200 From Anthropic Over Book Use — Yet Says Years of Litigation Left Him Poorer
- Midjourney Demands Disney, Universal, Warner Bros. Disclose Their Own AI Use in Copyright Fight
- LLM
- — Large Language Model의 약자로, '거대 언어 모델'이라고 해요. ChatGPT, Claude 같은 AI가 바로 LLM이에요. 엄청나게 많은 텍스트를 학습해서 사람처럼 글을 쓰고 대화할 수 있어요.