Amazon destroys rare books to feed AI training pipelines

Amazon's systematic acquisition and destruction of printed books for AI training data represents a scaling strategy that raises questions about resource consumption and cultural preservation in the foundation model era. The practice, tracked via AirTag logistics, reveals how major labs source training material at industrial volume, destroying physical artifacts in the process. This touches on both the data sourcing practices underpinning large language models and the broader tension between AI infrastructure needs and the depletion of rare intellectual resources.
Modelwire context
Analyst takeThe story's real news isn't that Amazon uses books for training (known) but that the scale and method are now observable via logistics tracking. This means the data sourcing playbook for major labs is becoming transparent in ways it wasn't before, and the destruction of rare materials suggests labs are optimizing for volume over curation.
This is largely disconnected from recent activity in the space we've covered. It belongs instead to the emerging category of AI infrastructure transparency stories, where operational practices become visible through supply chain leaks or forensic tracking. As labs scale foundation models, the sourcing decisions they make (what gets bought, what gets destroyed, what gets preserved) reveal their actual cost-benefit calculations. This matters because it shows labs treating rare intellectual resources as consumable inputs rather than irreplaceable assets, a trade-off that will shape what training data looks like for the next generation of models.
If other major labs (OpenAI, Google, Anthropic) issue public statements about their book sourcing practices or data retention policies within the next 60 days, that signals the practice is becoming a reputational liability. If Amazon's book destruction accelerates rather than slows after this reporting, that confirms the company views the PR cost as lower than the operational benefit of the current approach.
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsAmazon · AirTag · The Decoder
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The Decoder originally reported this story as “AirTag reveals how Amazon destroys rare books for AI training”. The full content lives on the-decoder.com. If you’re a publisher and want a different summarization policy for your work, see our takedown page.