Amazon has been purchasing and physically destroying rare books to scan their pages for artificial intelligence training data. Investigative reports revealed that the company strips the spines off out-of-print volumes at specialized facilities, feeding its large language models with exclusive offline text.
This aggressive data-gathering tactic highlights the severe shortage of fresh training material facing tech giants. As internet sources dwindle and face risks of AI contamination, pre-digital literature provides pristine text necessary to maintain model accuracy and prevent quality degradation.
- Amazon cuts up and scans rare books for AI training
- Offline texts provide crucial data beyond internet limits
- Older literature helps prevent AI model collapse
- The operation was uncovered using a tracking device
Sources:
