본문으로 건너뛰기
All news

Anthropic Bought Millions of Books, Scanned Them, Then Destroyed Them

In one line: Anthropic reportedly purchased millions of physical books, cut and scanned them for AI training text, then discarded the originals.

Key points

  • Anthropic is reported to have bought used books in bulk, sliced off their bindings, and digitally scanned the pages to build training text.
  • The scanned paper originals were reportedly discarded rather than reused.
  • The practice surfaced through copyright litigation, raising the question of how physically digitizing lawfully acquired books differs from unauthorized web scraping.

Why it matters

The provenance and legality of AI training data is among the industry's biggest risks. A "buy-then-scan" approach may carry a stronger legal defense than unauthorized crawling, but destroying the books themselves raises its own copyright and resource concerns — and could shape how labs source training data going forward.

Read more

How this story unfolded

  1. Microsoft Ramps Up Open Rivalry With OpenAI and Anthropic
  2. Microsoft Steps Up Direct Rivalry With OpenAI and Anthropic
  3. Author Gets $6,200 From Anthropic Over Book Use — Yet Says Years of Litigation Left Him Poorer
  4. Anthropic Launches Claude Opus 5: Near-Frontier Performance at Half the Price
LLM
— Large Language Model의 약자로, '거대 언어 모델'이라고 해요. ChatGPT, Claude 같은 AI가 바로 LLM이에요. 엄청나게 많은 텍스트를 학습해서 사람처럼 글을 쓰고 대화할 수 있어요.

뉴스레터 구독

무료 뉴스레터

매주 핵심 AI 소식, 한 번에 받기

쏟아지는 AI·LLM 뉴스 중 꼭 알아야 할 것만 골라 메일로 보내드려요. 뉴스레터 발송이 시작되면 구독자분들께 가장 먼저 보내드립니다.