back
Physical Ownership Is Not a Training License
seikouAI

Physical Ownership Is Not a Training License

About the Author

Sources

Amazon’s scanning operation turns an old copyright distinction into a live AI question about what buying one copy actually permits.

Las Vegas, VGT3

A rare-book seller had reason to wonder where an unusually large order was going. Instead of guessing, the seller and 404 Media put a tracking device inside one of the books.

The shipment traveled across the United States and eventually arrived at an Amazon warehouse in Las Vegas. Employees at the facility, identified as VGT3, told 404 Media that their work consists of processing large quantities of printed books. The bindings are cut away so the pages can be scanned quickly. The physical copies do not survive the process.

Amazon did not deny that it was buying the books. A company spokesperson said that Amazon “purchases books through commercial channels to help develop and improve the products and services our customers use.”

What Amazon did not publicly identify was equally interesting. It did not say which models were being developed with the scans, how the resulting files were being used, how long they would be retained, or what legal theory Amazon relies on when copyrighted books enter the operation.

That leaves an unusually clean copyright question sitting underneath a visually provocative story about books being sliced apart. Amazon paid for the books. What, exactly, did it buy?

What the Sale Transfers

Copyright law has answered the first part of that question for decades. Buying a book gives the purchaser ownership of that physical copy. It does not transfer the copyright in the work contained inside it. Section 202 of the Copyright Act makes the separation explicit: ownership of a copyright is distinct from ownership of the material object in which the work is embodied.

The familiar first-sale doctrine gives the owner of a lawfully made copy considerable freedom over that object. A reader can resell the book, donate it, lend it, tear pages out of it, or throw it away. Amazon can buy a used book and destroy it without asking the author whether the author approves of what happens to that particular copy.

Scanning introduces another right. Section 106 gives the copyright holder the exclusive right to reproduce the protected work, subject to the limitations elsewhere in copyright law. A digital scan creates another copy. The fact that the paper original is subsequently destroyed does not somehow cause the digital reproduction to disappear from the legal analysis.

The U.S. Copyright Office made the point directly in its report on generative AI training. Creating a training dataset from copyrighted material involves acts of copying.

Downloading, transferring, converting, filtering, and preparing material for training may implicate the reproduction right unless the developer has permission or can rely on another defense.

So the purchase receipt does useful legal work. It establishes lawful acquisition of the physical object. It does not, by itself, function as an AI-training license. That is where the simple answer ends.

Fair Use Enters the Warehouse

An unauthorized reproduction is not automatically an infringing reproduction. Fair use can permit copying without the copyright holder’s consent, and American courts have long allowed substantial copying in some technological contexts.

Google famously scanned millions of books without obtaining permission from every copyright holder. The Second Circuit upheld the Google Books project because its search functionality and limited display of snippets served a sufficiently different purpose and did not provide the public with substitutes for the books themselves.

The Internet Archive later discovered how sensitive that reasoning is to the use made of the resulting digital copy. It also scanned physical books. Its controlled digital lending program made full digital versions available to readers. In Hachette Book Group v. Internet Archive, the Second Circuit rejected the fair-use defense, emphasizing that the digital copies served the same basic purpose as publisher-authorized ebooks and affected an established licensing market.

The scanner was not decisive in either case. The downstream use was.

Generative AI presents an even harder version of the problem because the digital copy may never be shown directly to the public. It can instead become an intermediate asset in a training pipeline. Courts then have to evaluate why the copy was made, how it was used, what the resulting system does, and whether that use damages markets copyright law recognizes.

Amazon’s silence about the specific destination of its scans therefore leaves out information that would be central to any serious fair-use analysis.

Bartz Is the Hard Case for the Copyright Argument

Anyone writing about Amazon’s operation now has to deal with Bartz v. Anthropic. Ignoring it produces an easier story and a much weaker one.

Anthropic had assembled an enormous collection of books for developing Claude. Its sourcing practices included pirated digital libraries, but it also bought physical books, cut off their bindings, scanned them, and retained digital replacements. The resemblance to the Amazon operation reported by 404 Media is obvious.

In June 2025, Judge William Alsup ruled that Anthropic’s use of books to train its large language models was fair use. He also concluded that the company’s one-to-one conversion of lawfully purchased print books into digital copies for its internal library was fair use when the physical copies were destroyed, and the digital versions replaced them for storage and search.

The same opinion treated Anthropic’s acquisition of millions of books from pirate libraries very differently. Building a permanent central library from unlawfully obtained copies was not excused merely because some of those books were later useful for training.

That part of the dispute eventually resulted in the $1.5 billion class settlement, which received final approval in July 2026. The settlement did not overturn the earlier ruling that the training use and the one-to-one scanning of purchased books were fair use. In fact, the settlement proceedings expressly recited those findings while resolving claims arising from Anthropic’s pirated acquisitions.

Amazon therefore has a serious legal argument available if its operation resembles the purchased-book side of Bartz.

It also has only a district-court decision. Bartz did not create a nationwide rule that buying one physical copy grants every AI developer the right to scan it. Fair use remains a defense applied to particular facts, and the ultimate use of Amazon’s scans is still unknown.

Kadrey Refuses a Comfortable Rule

Two days after the Bartz decision, another federal judge in the same district reached a Meta-friendly result while explaining the law quite differently.

In Kadrey v. Meta, Judge Vince Chhabria granted Meta summary judgment against thirteen authors whose books had been used to train Llama. The authors had not produced sufficient evidence that Meta’s particular use damaged the market for their works. On that record, Meta won.

Chhabria nevertheless rejected the idea that the decision settled AI training generally. His opinion focused heavily on the possibility that generative systems could flood markets with works competing against human-created books. Better evidence of that effect, he suggested, could change the fair-use calculation in another case.

The result is less contradictory than it first appears.

Bartz demonstrates that a court can regard AI training and destructive scanning of purchased books as fair use. Kadrey demonstrates that even a court ruling for an AI developer can consider unauthorized training potentially unlawful on a different evidentiary record.

The Copyright Office has taken much the same position. Some training uses are likely transformative, but the answer depends on the works involved, their source, the purpose of the model, and its effect on the relevant market.

There is no general American AI-training exception to copyright law. There is also no general rule requiring a training license whenever copyrighted material is copied into a model development process.

Amazon has arrived directly inside that unresolved space.

Rare Books Create a Different Problem

Copyright is only part of the discomfort surrounding destructive scanning. A common used paperback and a scarce first edition may contain the same copyrighted text. From a copyright perspective, their rarity may change very little. From a preservation perspective, the difference can be enormous.

The Guardian reported days before the Amazon investigation that booksellers in the United Kingdom and Ireland had been seeing unusual bulk orders for eclectic collections of older books. Some sellers suspected AI-related acquisition, although the buyers and ultimate uses were not established. The requested works could jump from obscure historical material to unusual editions and translations without any collecting logic familiar to booksellers.

Anthropic had already exposed the industrial logic behind such acquisition. Internal documents concerning its Project Panama described an ambition to “destructively scan all the books in the world.”

Books can provide edited, pre-generative-AI text that does not exist in convenient online collections. Once that text becomes scarce as training material, the dusty shelf begins to look like a data center.

The preservation problem does not depend on whether a particular work remains under copyright. A public-domain book can be culturally scarce. A copyrighted bestseller can be physically abundant.

Current copyright doctrine is poorly designed to decide whether a company should destroy a scarce physical object after buying it. That is a stewardship question, and it deserves to be treated separately rather than smuggled into fair-use analysis.

Disclosure Before New Copyright Law

The obvious legislative response would be to require developers to obtain licenses before scanning copyrighted books for AI training. The current case law does not support such a clean rule, and Congress would immediately have to confront research uses, public-domain mixtures, collective licensing, and the practical scale of modern training datasets.

I would start somewhere more basic: provenance.

Companies conducting industrial-scale acquisition of books for AI development should maintain records showing what they acquired, where it came from, whether the source was lawful, what was scanned, what happened to the physical copy, and where the digital reproduction went. If a model was trained on the resulting material, the relevant model or development program should be traceable to that acquisition record.

The same system could distinguish ordinary commercial books from material identified as rare, antiquarian, or otherwise preservation-sensitive. Destructive scanning should not be the default for a copy whose physical survival has independent cultural value. Non-destructive scanning, resale, donation to an archive, or transfer to a preservation institution are available alternatives once rarity is recognized before the binding hits the cutter.

Licensing can then develop where copyright law actually requires it. The Copyright Office has so far resisted imposing a compulsory federal licensing system and has recommended allowing voluntary markets to continue developing. That approach makes more sense if developers disclose enough about acquisition and use for copyright owners to know what market they are negotiating in.

Commercial Channels Do Not Answer the Copyright Question

Amazon’s statement is carefully framed. The books were purchased through commercial channels. That addresses one of the ugliest problems exposed by earlier AI copyright litigation: the use of pirate libraries when lawful copies were available.

It does not resolve everything that happens afterward. A commercial purchase establishes lawful possession of a physical copy. Scanning creates a reproduction. Training may introduce another use. Fair use can protect some or all of those acts, as Bartz demonstrates, while different facts can push the analysis elsewhere.

The Copyright Act already contains the conceptual pieces. What remains unsettled is how courts will apply them to industrial AI development at this scale.

The next case may therefore turn on evidence more mundane than the image of machines cutting books apart. A court may want to know which copy Amazon bought, whether one digital copy replaced it, which additional copies were created during processing, how long they were retained, which model received the material, and what that model was built to do.

Those records either exist or they do not. A receipt for the book will be only the beginning.