Amazon is destroying rare books to train its artificial intelligence models, a practice first reported by Ars Technica. This revelation, also covered by TechCrunch, shows a company founded on selling books now disposing of them to fuel its AI ambitions. The books are especially valuable for training large language models (LLMs) because these texts contain unique data, distinct from the content already widely available online.

Internally, Amazon’s AI team uses a striking logo: a T. rex preparing to devour a book. This imagery offers a blunt symbol of the company’s aggressive approach to data acquisition for AI development, an approach that has become increasingly visible across its various services.

A Pattern of Broad Data Harvesting

The destruction of rare books for AI training is not an isolated incident. It fits into a broader pattern of Amazon seeking extensive data to improve its AI capabilities. Earlier, Wired reported that Amazon was using content from its streaming platform, Twitch, to train its AI models. Twitch users discovered their content was being absorbed unless they actively opted out. This policy sparked widespread questions from thousands of users about why their streams were used for AI training without explicit, opt-in consent.

Both instances highlight Amazon’s strategy of leveraging its vast holdings—from physical books to user-generated digital content—as raw material for AI development. The company appears to prioritize data ingestion for its AI over traditional content preservation or explicit user permission, requiring users to take action to protect their data.

Coincidentally, Amazon has also reinstated a user agreement clause that may hinder class-action lawsuits against it, according to Livemint. This clause requires customers to resolve disputes outside of court, shifting potential legal battles to arbitration or individual claims. The timing of this reinstatement, amid revelations of aggressive data collection for AI, suggests a preemptive move to manage potential legal challenges arising from its AI training practices or the outputs of those models.

Such a clause could reduce the likelihood of large-scale litigation related to data privacy, content usage, or algorithmic bias, all of which are growing concerns as AI models become more pervasive and powerful. By directing disputes away from class actions, Amazon minimizes its legal exposure at a time when its AI data demands are expanding.

Shifting AI Focus in Key Markets

While Amazon collects diverse data globally, its AI application strategy varies by market. In Indian ecommerce, for example, Amazon’s AI focus has shifted significantly. For much of the past two years, the AI race among companies like Amazon, Flipkart, and Meesho revolved around consumer-facing applications. Now, as Inc42 reported, the focus is increasingly on seller-side stacks. This means developing AI tools to assist sellers with inventory management, pricing optimization, logistics, and other operational tasks, rather than primarily on customer recommendations or chatbots.

This shift indicates that the vast quantities of data Amazon collects, whether from rare books or Twitch streams, are likely feeding a multifaceted AI strategy. The company is building models for consumer interactions and for optimizing its complex operational backbone, impacting sellers and internal processes.

The destruction of rare books, the use of Twitch content, and the reintroduction of the “no class-action” clause collectively illustrate Amazon’s aggressive pursuit of data for AI development, alongside its efforts to control potential legal fallout. The company’s internal T. rex logo for its AI team seems to accurately reflect its approach to sourcing the unique data needed to train advanced models, even if it means discarding historical texts.

What is Amazon doing with rare books?

Amazon is destroying rare books to train its artificial intelligence models, as these texts offer valuable, unique data beyond what is available online.

What is Amazon’s policy on using Twitch content for AI?

Amazon is using content from Twitch streamers to train its AI models unless those streamers actively opt out of the program.

Compiled by Launch91 Desk from the sources linked above. More about Launch91.