Anthropic spent tens of millions of dollars buying old books, cutting off spines, scanning page by page, and finally mashing the paper into pulp. The most emotionally charged part of this incident is “book destruction,” but what really deserves discussion is not just whether the book was destroyed.
The more important question is: why would a company training large models rather build an expensive physical-book processing pipeline than buy ebooks directly or sign licensing agreements with publishers? This approach, known as “Project Panama,” brings together the conflicts between AI training data, traditional publishing rights, digital-content licensing, and model distillation.
It may seem absurd, but it follows a concrete business and legal rationale.
What did Project Panama do?
According to leaked project materials, the project began in 2024. Anthropic bought used books in bulk from secondhand booksellers and handed them to a professional scanning company for processing.
Once a book enters the assembly line, it generally goes through four steps:
- Use a hydraulic paper cutter to cut off the spine.
- Separate the book into loose sheets that can pass through an automatic feeder.
- Digitize the pages with a high-speed duplex scanner.
- After scanning, the paper is sent for recycling and turned into pulp.
High-speed scanners can process dozens or even hundreds of pages per minute. When the target range of operations expands from hundreds of books to hundreds of thousands or even millions, cutting off the spine is almost the only way to maintain throughput.
The materials say that the project processed an extremely large number of printed books in roughly six months. This was not a temporary operation in which a few people flipped pages and took photographs, but an industrial workflow linking procurement, transport, disassembly, scanning, optical character recognition, storage, and disposal.
Why must the spine be cut off?
Many people’s first reaction is: scanning is understandable, so why destroy the book? The answer comes first from efficiency.
A fully bound book cannot pass through a scanner’s automatic document feeder. If the binding is preserved, an operator must turn every page and capture it with a flatbed scanner, camera, or specialized device.
This approach mainly has four issues:
- Page turning speed is slow, and labor costs are high.
- If the pages are bent, the positions near the binding line are prone to deformation.
- Thick books are hard to flatten, and the accuracy of text recognition decreases.
- Prolonged exposure to scanning light sources and repeatedly flipping pages increases the physical burden on operators.
After the spine is removed, the book becomes a stack of uniformly sized sheets. The pages can be fed continuously and scanned on both sides at high speed.
For someone scanning an occasional book, preserving the original is not difficult. But for a project processing millions of books in a matter of months, a few extra minutes per volume can become prohibitively expensive.
Cutting books is therefore not done for dramatic effect; it is a choice driven by the speed and reliability required at industrial scale.
Can a cut book be restored?
Removing the spine does not necessarily mean the book must be discarded. In some library digitization projects, the pages are rebound after scanning.
Restored books are usually one or two millimeters narrower than original, and hardcover covers may need to be remade, but the main text can still be used. “Scan and destroy it” is therefore not a decision made by the equipment, but rather a follow-up disposal method actively chosen by the project team.
This is also the starting point of the controversy. If books can be rebound, why not preserve them, donate them, or put them in the library?
The answer involves not only storage costs but also legal risks regarding whether digital copies and physical originals exist simultaneously.
Why is it difficult to use book scanning without dismantling books?
Scanning techniques that preserve original books have always existed. Common methods include:
- Press the book onto a flatbed scanner and scan it page by page.
- A book scanning platform with a sloped surface is used to reduce pressure on the spine.
- Use the camera above to take photos of the two unfolded pages.
- Automatically lifts and scans pages using a suction page-turning device.
These devices are suitable for archives, historic volumes, rare books, and collections that must not be damaged. Their problem is not feasibility, but speed and cost.
Manual page-by-page turning significantly reduces processing volume. Automatic page-turning devices have complex structures and are sensitive to paper thickness, static electricity, damage, and adhesion.
For books that are old, have fragile paper, or have special layouts, manual intervention is still required. If the goal is to protect a small amount of precious materials, these costs are entirely reasonable.
If the goal is to prepare millions of ordinary publications for model training, these methods struggle to compete with a cut-and-feed workflow.
Google Books offers another point of comparison
Large-scale book scanning is not a new phenomenon after generative AI emerged. Google partnered early with libraries and publishers to digitize massive collections.
Those projects also face issues such as scanning speed, binding structure, text recognition, and copyright boundaries. The difference is that traditional book search projects usually design rights boundaries around “retrieval” and “display fragments.”
Public-domain books can be displayed and used more fully. Books still protected by copyright are generally limited to short excerpts, with users directed to purchase or borrowing options.
This model is easy to explain to publishers: digital services help readers discover books and do not replace selling entire books.
Model training is different. The training process doesn’t simply present a book as a search result, but instead absorbs the language, facts, structure, and expression patterns from the book into the model’s parameters.
Publishers find it difficult to measure this usage by the traditional “number of copies sold.”
The controversy isn’t just about “whether you bought the book.”
In training data disputes, the source of the physical book is very important. Downloading an e-book directly from a pirated website is not the same as buying a physical book on the second-hand market.
The former first faces the issue of whether the source is legitimate. The latter completes at least one genuine physical goods transaction.
This is also why “buying old books” is the first step in a complete solution. AI companies can claim that they obtain a legally purchased physical book and scan and process the items they own.
But this does not automatically resolve all disputes. Whether purchasing a physical book grants the right to train the model, whether training constitutes transformative use, and whether the model will reproduce protected content can all be independent issues.
Buying the book is therefore not a universal get-out-of-liability card. It simply makes lawful provenance easier to establish than when electronic copies of unknown origin are used directly.
Why destroying originals after scanning is crucial
The most counterintuitive step in this approach is precisely destroying the original book. From the perspective of traditional copyright logic, rights holders are most sensitive to unauthorized copying and dissemination.
If a paper book is scanned into a digital copy while the original book still exists, the apparent situation changes from “one” to “two.” If the original is subsequently destroyed, the project team can emphasize:
- The original book was purchased legally.
- Digitization serves internal analysis or model training.
- The physical originals will no longer circulate.
- The complete digital copy was not resold as an e-book.
This arrangement attempts to describe the act as a medium conversion, rather than creating an additional independent circulating copy. Whether it constitutes fair use still depends on the specific case, usage, market impact, and output limitations.
Different jurisdictions may also reach different conclusions. So a more accurate statement is: destroying originals helps build legal arguments, but does not mean all similar operations are inherently legal.
Why not buy training rights directly from the publisher?
Since scanning paper books is so troublesome, the most natural solution seems to be to buy the publisher’s electronic documents directly. The main practical obstacle is not technology, but pricing.
Publishers are familiar with selling copies. After a copy is sold, the publisher can pay the author royalties under the relevant contract.
But model training is neither a conventional one-time reading nor the sale of one copy to one reader. Publishers must answer a series of questions that remain unresolved:
- Should one training use be priced as one copy, hundreds of copies, or thousands?
- How much potential sales will the model reduce after mastering the knowledge in the book?
- If the same book is used for pre-training, fine-tuning, and evaluation, should separate fees be charged?
- Do model updates and repeated training require reauthorization?
- How do authors, publishers, and digital distribution platforms distribute revenue?
- When the model only learns universal language rules, where are the boundaries of rights?
AI companies are also reluctant to accept a very high fixed standard too early. Once a certain price becomes industry norm, the cost of training millions of books can rapidly balloon.
When neither side accepts the other’s pricing formula, buying physical used books and digitizing them in-house becomes an alternative with predictable costs and a controllable process.
Why old books specifically?
Buying old books isn’t just because the price is low. For training data, the publication date itself is also valuable.
Books published before 2022 were largely created before generative-AI content entered the internet at scale. These texts were written, edited, proofread, and formally published, while their provenance and dates can usually be confirmed through ISBN records and other metadata.
Compared to web texts with mixed sources, formal publications usually have several characteristics:
- Their long-form structure is complete.
- The language has been edited.
- Professional topics are more focused.
- Their publication dates are traceable.
- They are less likely to contain AI-generated material.
This makes old books a high-density corpus with clear boundaries. The importance of ISBN databases lies in their ability to help buyers organize bibliographies by author, subject, edition, and publication year.
For companies that need to systematically purchase training data, this is much more efficient than searching for books individually on second-hand platforms.
What does “AI-generated content pollution” mean?
If the model is continuously trained on text generated by other models, the data distribution may gradually narrow. Errors, clichés, and statistical biases may also be repeatedly amplified.
In research discussions, “model collapse” is often used to describe such risks, but it does not mean “as long as AI text is mixed in, the model will definitely fail.” The actual impact depends on:
- The proportion of AI-generated content in the training set.
- Whether the data has been filtered and deduplicated.
- Whether there are still sufficiently diverse samples of human creation.
- How to design training objectives and data ratios.
- Whether the synthetic data has been verified and corrected.
The value of old books therefore does not come from any mysterious quality. They provide a body of clearly dated human writing that was less influenced by generative AI.
As AI content on the public web increases, such traceable corpora will become increasingly scarce.
From printed books to ebooks, the rights obtained through purchase have changed
The rules for printed books are relatively straightforward. After buying one, readers can read, collect, resell, give it away, or even use it to prop up a table.
Restrictions often focus on unauthorized copying and public dissemination. E-books have changed this relationship.
After paying fees, users often receive not full ownership of the document, but a reading license restricted by platform terms. E-books may restrict the number of devices, borrowing periods, copy scope, and file formats.
In certain cases, platforms may even withdraw the content. On the surface, it’s all about “buying books,” but the actual control gained by print book buyers and e-book users is not the same.
This also explains a seemingly strange phenomenon: AI companies may not have more flexible access to electronic documents directly than buying physical old books.
Electronic documents are more convenient, but often come with clearer contract restrictions.
Large model companies are also restricting others from “learning from themselves”
There is also a clear symmetry in this controversy. AI companies want broad access to human-created works for training, yet they usually prohibit other companies from calling their models in bulk for distillation.
Ordinary users pay pay-as-you-go to call models and generally do not encounter problems. But if someone registers a large number of accounts, continuously asks a flood of questions, and trains competing models with answers, the platform usually considers them to violate the terms of service.
The underlying business logic is very similar to the concerns of publishers:
- Publishers worry that models will absorb books and reduce their sales.
- Model companies worry that competitors will distill their capabilities and reduce their API revenue.
- Both sides are trying to prevent paying users from turning a single purchase into a new product that can replace themselves.
The difference is that traditional copyright systems are built around copies of works, while model capabilities are difficult to map to any specific copy. What the AI era lacks is a mature framework for describing how much a machine learned, what it displaced, and what should be paid.
Why projects need to be kept confidential
From a business perspective, large-scale procurement of bibliographies, scanning capabilities, data recipes, and costs all fall under competitive information. But trade secrets alone are not enough to explain the sensitivity of this matter.
Destroyed books carry strong cultural symbolism. An ordinary old book may be worth only a few dollars on the market, yet people still believe it carries the author’s labor, personal memory, and public knowledge.
When “buy, cut, scan, pulp” is compressed into a single sentence, a company can easily be portrayed as systematically destroying human culture to train machines. That perception may be incomplete, but it is highly transmissible.
Even if the project team believes the process has legal basis, they realize it is difficult to gain intuitive public approval. This leads to a situation where “it can be done, but not suitable to be publicly discussed.”
Confidentiality itself further deepens suspicion, as it makes outsiders believe that companies are aware of ethical issues.
Ordinary old books and rare books should not be confused
When discussing “book destruction,” it is also necessary to distinguish the actual condition of the books. Ordinary old books that circulate extensively, have multiple collections, and digital copies are not the same as rare editions, out-of-print books, signed editions, and local documents.
Once the former is dismantled, cultural loss is usually limited. Once the latter is destroyed, it may result in irrecoverable information loss.
A responsible procurement and scanning process should at least establish the following screenings:
- Whether there are multiple accessible collections.
- Whether it is an out-of-print or small-batch publication.
- Whether it contains unique information such as signatures, annotations, or bookplates.
- Whether it has local history, family history, or archival value.
- Whether it can be rebound after scanning or transferred to a collecting institution.
“Legal purchase” only indicates the source of ownership and cannot replace the judgment of scarce cultural materials. If rare books were also indiscriminately sent into shredding and pulping processes, the controversy would no longer be just an emotional issue.
What does this controversy really reveal?
The “Panama Plan” is not just a simple corporate curiosity story. It exposes three gaps in the content industry in the AI era.
First, traditional copyright focuses on the number of copies, while model training focuses on information absorption. Training a model on a book does not create another volume in its original form, but it may still affect the work’s market value.
Second, publishers’ licensing systems are still centered on sales and reading. They have yet to establish a training license pricing method jointly accepted by authors, publishers, and AI companies.
Third, AI companies take different positions toward external content and their own models. They emphasize the transformative nature of training while using their terms of service to prevent others from distilling their outputs.
These three types of contradictions will not automatically disappear with a lawsuit or a settlement. As long as high-quality human text remains a key ingredient for models, similar disputes over procurement, authorization, and data sources will continue to arise.
Summary
Anthropic buys old books, cuts them, scans, and then destroys them, which seems like an uncomfortable technical pipeline. But it’s not just about saving scanning time.
Buying physical books helps establish lawful provenance; trimming the spines increases industrial throughput; destroying the originals strengthens the argument that this is a change of medium rather than the creation of extra copies; and choosing old books balances cost against demand for data with less AI-generated contamination. The entire approach shows how poorly rules formed in the print-publishing era map onto model training.
Paper book trading focuses on “who owns this book,” e-book licensing focuses on “how users can read it,” and model training questions “can machines learn, and what does it replace after learning?” Until new authorization and allocation mechanisms mature, AI companies will continue to seek viable paths within existing rules.
What truly needs to be addressed is not simply labeling “book destruction” as legal or illegal, but establishing a set of rules that can simultaneously measure author rights, publishing markets, technological innovation, and cultural preservation.