The closest they got was admitting that the rare books weren’t anything that someone might care about in the sense that people assume when we hear “rare books”
> As the bookseller who sold them told me, there are not many people in the world who would care about them in the same way people might care about the first edition of Oliver Twist, but that doesn’t mean they’re not valuable.
Okay? But then why exactly where they considered valuable enough to warrant an entire article about them going to a book scanning facility? Without revealing anything about these books I have no idea if they were classic literary works that were underappreciated, or if this was some old guide about How to Use Microsoft Office 97.
There was a more balanced take on Twitter (which I’m unable to find again, because Twitter) from a book seller who said it was more of the latter type: Books that were rare because they were no longer in demand and most everyone had thrown their copies away. Some parts of the media are doing backflips to try to imply that these are cherished literary classics being fed into the shredder to deprive humanity of something valuable, but the book seller seemed happy to be making sales for useless old books that no human was interested in buying.
These LLM training runs have already ingested essentially the whole public internet. What marginal value is to be gained from scanning and destroying obscure books?
For me the biggest functional issues with LLMs don't seem to have any connection with "I wish they had read this obscure community cookbook from 1946". Is that going to get Claude to stop saying "honestly" to me? Is it going to get LLMs to stop making up sources that don't exist? What is in it for Amazon or any LLM company to chase more obscure data like this.
In any case, all the indignation about destructive digitization misses the point that rare books takings space in a warehouse for years without being bought will eventually be destroyed anyway.
Not even the title of one of those rare books?
It'll take a while before publishing collapses due to the availability of the same information via an LLM, but once this happens new books will get a lot more expensive for AI companies and a lot cheaper for everyone else.
I always struggle with their articles because it feels like they built it for rage bait on a topic and leave the other interesting topics out of it. It is an absolute shame for books to be destroyed but what does it really mean to be rare here? I know they kind of tried to differentiate but it sounds like this could be John Doe’s self help book that never sold well. If you ever are connected to a library you will start to realize how many books simply get thrown out or sold for nothing because nobody wants them.
To me the problem is partly the copyright law. I think it’s. Hard problem but I always lean more towards books having short copyright shelf lives and making it legal for digital copies to be shared after which I think would eliminate I good part of this problem. Not to mention 99% of the books published are probably garbage but that is highly subjective.
So it is bit of a meta rant but I think there a couple holes to go down that could be extremely interesting but they always write these informationally light articles. Like scrolling through a NYT visualization for just some shipping datapoints. Don’t really dig deep on anything and then end with a trust me bro these are rare books that Amazon is destroying for AI when I cannot be that upset with the amazons of the world. There are a lot of reasons for a business to digitize books, most books are worthless and it makes sense to cut the bindings for scanning. I would rather talk about how could you fix copyright to make this less an issue but is it even an issue with how many books get thrown out?
My writing got better and better, and I got better publishers who actually hired editors, but the books sold fewer and fewer copies. Even crappy self-serving poorly written stackoverflow posts are often good enough. Then LLMs killed stackoverflow. Is that real ironic or Alanis ironic?
But I'm not holding my breath waiting for a book deal from OpenAI.
I was going to say a very similar thing, and this is something I strongly disagree with when it comes to HN's moderation, and it goes like this:
The product is the outrage.
And HN should know better and mods should actively discourage, warn and prevent accounts (who karma farm, among other things) from even being able to post rage-bait articles. These aren't "hacker curiosities", they're just insipid bullshit.
It would seem the redaction of the rare titles is a way to avoid de-anonymization and subsequent harm to the business of the seller who agreed to place a tracker in one of the books. That being said, maybe they could have chosen a better methodology which would have allowed the disclosure of the title, although ultimately I’m not sure the title matters too much outside of their claim they were “rare”.
Like Amazon consuming and presumably destroying rare books should be enraging to everyone, regardless of political persuasion.
The difference is huge between 404 and Fox. Fox is out here trying to tell people there's a trans agenda, and that Biden was a lunatic leftist. They are just making up stories and publishing them because they know their audience engages. 404 definitely make editorial choices about which stories to pursue but I've largely found them to be grounded in real depictions of stuff that is happening.
This situation has absolutely no relation to that. This is the exact opposite, and the only reason this information isn't available publicly and is in risk of getting lost is copyright law.
TLDR: Amazon isn't the nazis in this story, copyright law is.
Anyway ... this case is:
Bartz v. Anthropic PBC, No. 3:24-cv-05417-WHA, U.S. District Court for the Northern District of California, decided by Judge William Alsup
“the purchased print copy was destroyed and its digital replacement not redistributed, this was a fair use.”
https://copyrightalliance.org/wp-content/uploads/2025/06/Bar...
Obviously this violates precedent, I had an internal LLM (probably a frontend for Claude or ChatGPT) find them (just like the convictions for file sharing in the 2000s required counter-to-the-law reasoning by judges, fair use was almost never accepted as a valid excuse, even when it obviously was, but of course Sony was a billion dollar company and needed to be in the right. In fact that this had to happen was explicitly given as a reason to create the DMCA)
Anyway, some precedents:
Hotaling v. Church of Jesus Christ of Latter-Day Saints, 118 F.3d 199 (4th Cir. 1997)
“Although the Church acknowledges that its sole remaining copy is not the one it originally acquired … it maintains that the remaining copy does not infringe Hotaling's copyright because it is a replacement copy…”
(this reasoning was rejected by the court)
https://law.justia.com/cases/federal/appellate-courts/F3/118...
Atari, Inc. v. JS & A Group, Inc., 597 F. Supp. 5 (N.D. Ill. 1983)
... defendant sold a device for making backup copies of copyrighted Atari cartridges and argued that §117 permitted replacement/archival copying. The court rejected the broad replacement theory.
https://law.justia.com/cases/federal/district-courts/FSupp/5...
This very court has clearly declared that making a copy of a copyrighted work for replacement purposes is illegal, on multiple occasions.
I would like to point out that this isn't Anthropic's only extreme-WTF law violation. When the original judgement against them was made against them, they were forced to admit that using books to train models was illegal if acquired illegally AND THEN WERE ALLOWED TO KEEP DOING IT. That's not how this works. In my opinion Anthropic and OpenAI and everyone else need to at minimum take training material that was acquired in violation of copyright out of their training data unless and until they have a separate licensing agreement with the copyright holders.
Because of the copyright-filesharing court wars of the 2000s, which were also handled dishonestly by courts (whether we're talking US or EU courts), and the absurd copyright extensions, they had to now make some new excuse, and settled on this very sad, very destructive option. It's not even defensible legally, imho, but of course the biggest wallet must win. I don't understand. It's such a sad joke at this point, and it's not like the courts even still had credibility after the file sharing cases.
I wonder which sad excuse will be forthcoming from the courts once and if we do have someone release a movie made by an AI model that is obviously a direct ripoff from some high-budget movie. We all know that that's coming, and probably sooner than most people predict.
What if a person with small children and an elderly, incontinent pet with a penchant for peeing on books wants to buy it - can I sell this book to such a dangerous purchaser who might destroy it?
I love book as much as the next person but the hyperbole about "rare books" is absurd. Nobody is buying the Gutenburg Bible and destroying it for AI. The books in question are certainly not rare enough to be in museums - without titles there's no proof these are anything of real value.
This is the citation needed that is missing from every report so far, including this one which deliberately refuses to reveal anything about these books.
You’d think if there were examples of actually valuable, rare books being shredded that the journalists would at least be able to name one such example. Instead it’s always vague posting about the destruction without ever naming any examples.
I think it’s because if they named some example titles, everyone would see that they don’t care about these books being shredded.
I even imagine that the market price for these “rare” books is helping filter out anything truly valuable and rare. It just reads as a rage bait tmz article. The quantity of used books including “rare” books that get thrown into the dump is astronomical.
My neighbor self-published a book, printed I think 100 copies at his own expense. It's literally a rare book. I can't imagine he nor anyone would care if an AI company bought a copy, no matter what they did with it.
Maybe these "rare books" are first edition Mark Twains, and. maybe they're unwanted books that would otherwise have gone to be pulped. The distinction is important and by not giving any evidence or even a qualitative claim about the types of books, 404 media is being pretty weak here.
The rare qualifier is used precisely because no reasonable person thinks this is stealing.
Nothing good comes from hoarding whether it being toilet paper, money or knowledge.
No reasonable person thinks buying 1 of something is "hoarding".
Advertisement
•
We placed a tracking device in a shipment of rare books to see which AI company was buying it, and found an Amazon facility where Amazon scans and destroys books.

The symbol for Amazon's VGT3, the Las Vegas facility where it scans book for AI training data.
Amazon is buying massive quantities of books, scanning them for AI training data, and destroying them in the process.
A 404 Media investigation was able to reveal Amazon’s book buying operation, which hasn’t been previously reported, by placing a tracking device in a rare book we suspected would be acquired by an AI company for training data, and following it around the country to its final destination.
That final destination was an Amazon warehouse in Las Vegas, Nevada. Amazon employees who work at this location say all they do is receive massive shipments of printed books which they then cut the bindings off in order to scan the books more quickly. The printed book is destroyed in the process. The logo of the Amazon team that works at this warehouse, called VGT3, is a dinosaur, brandishing its teeth and with a book in its hands.
“Amazon purchases books through commercial channels to help develop and improve the products and services our customers use,” an Amazon spokesperson told me in a statement.
The world’s AI companies are constantly looking for, and spending extreme resources to locate, more material to train their AI models. With books, that sometimes means destroying them in the process, something that large parts of the public have spoken up against, and which we can now confirm Amazon is doing.
In July, I published a story about booksellers who reported a historical spike in sales starting in the past year. They suspected this spike in sales was due to AI companies acquiring any books they can in search of new training data. Printed books are valuable as training data because a lot of the text they contain is not readily available on the internet, which AI companies have already scraped. The data is also conveniently organized and, if the book was printed before 2022, is guaranteed to be free of AI-generated text, which can make any AI model that is trained on it worse via a recursive process called “model collapse.”
📖
Do you know work at a facility where you scan books? I would love to hear from you. Using a non-work device, you can message me securely on Signal at @emanuel.404. Otherwise, send me an email at emanuel@404media.co.
These booksellers suspected AI companies were behind these large bulk purchases because of the high number of books they were buying, the seemingly random choice of books, and the fact that these buyers, unlike libraries and universities, did not seem price sensitive at all. But booksellers couldn’t say for certain who was behind the large purchases because the marketplaces where they sell their books keep the buyers anonymous. When an order comes in, a bookseller ships the sold books to a warehouse operated by the marketplaces, where books are sorted and then sent to the buyer.
In July, one bookseller told me they received a very large order of around 1,000 books on Biblio, one of these marketplaces. The seller agreed to put an Apple AirTag provided by 404 Media in one of the books included in this order so we could see where the book was going. And by extension, which company, AI or otherwise, was behind this massive order.
404 Media granted the bookseller anonymity because they worried sharing this information would harm their business. Biblio did not respond to a request for comment.
The bookseller sent the shipment to Biblio, and it arrived at a California airport. It then traveled by plane to Milwaukee International Airport in Wisconsin. Later that day, the book traveled to a warehouse belonging to a specialized shipping and distribution company called Trifinity, right outside of Kenosha Regional Airport, about 30 miles south. The book remained there for about two weeks, at which point it began traveling west by what appeared to be a truck. I could see the book travel via the highway and spend a night over at a trucking travel center around Grand Junction, Colorado. The next day, the book arrived at its final destination, an Amazon warehouse called LAS8 in Las Vegas.
LAS8 is one of several large Amazon warehouses in the area, each with a different specialty. LAS7, right across the street, for example, is a fulfilment center, while LAS8 appears to mostly operate as one of Amazon’s “print on demand” operations, which will print and ship books to Amazon shoppers as they are buying them. Initially, I was confused about why the book would arrive there, but Amazon employees who work at this location and who discuss working conditions there with other Amazon employees online explain that the the north end of the LAS8 warehouse, where I saw the book arrived, housed a different Amazon operation with a different code: VGT3.
Many Amazon warehouses have unique symbols to represent that specific site. VGT3’s symbol, painted on the entrance to the site and inside, is of a Tyrannosaurus rex, the massive carnivorous dinosaur, with its mouth open, holding an open book.

The entrance to Amazon's VGT3 facility.
“I work at VGT3 here in Vegas, and all we do is scan books,” one Amazon employee wrote on a forum for Amazon workers. “Some are assigned to cut books, and others go to receive where they get books and scan the bar codes. We didn't have rates, but now we do, but it's not stressful.”
“Working at VGT3 is nice all we do is scan books,” another Amazon employee wrote. “It's so cool.”
I saw VGT3 employees online talk about how working in this operation is a good but boring job because workers do the same, easy, repetitive tasks all day. I also saw some discussion indicating that VGT3 jobs are desirable for this reason, and because the warehouse sometimes offered night shifts. Earlier this year, some employees expressed concerns that Amazon would shut down the site because they had worked through the books they had and weren’t getting enough shipments of new books, though the site is still operational today.
That employees are scanning the barcodes or ISBNs on books — a unique serial number given to every published book — before scanning their content gives further credence to another theory put forth by booksellers: AI companies are trying to methodically scan every printed book in the world by working through the list of ISBNs. One bookseller told me they suspected this was the case because the very large orders they were getting never included very rare books that do not have ISBNs.
Elsewhere online in 2024, booksellers who sell their books on Amazon said they received a spike in orders to be shipped directly to the VGT3 location for a customer named “Amazon FC.” This was so unusual given its size they suspected the orders were some kind of scam. An Amazon representative chimed in to say they checked and that the orders were legitimate.
Like other major tech companies, Amazon is developing its own large language models (LLMs). Amazon considers its family of models branded Nova to be “frontier” models, meaning it considers them to be competitive with other cutting edge LLMs from Google, OpenAI, and Anthropic. These models require massive amounts of training data in the form of human-written text. Internally, Amazon also used an AI coding agent called Kiro to develop its own software.
We first learned that AI companies wanted to scan millions of books for training data because of a lawsuit from book authors against Anthropic, which revealed Anthropic’s “Project Panama.” The goal of the project was to acquire books from commercial bookselling marketplaces, cut the spines off the books, and scan them. It’s possible to scan books without destroying them, but cutting the spine makes it cheaper and faster. Additionally, the judge in the lawsuit ruled it was fair use and not a copyright violation for Anthropic to scan a book for training data in part because it destroyed the original, printed copy. Essentially, it’s the customer’s right to take physical media and store it digitally, and destroying the original copy means that copy isn’t duplicated and resold, and isn’t cutting into the publisher’s business.
We’re not revealing the titles of the books included in the shipment we tracked, but they are rare, meaning there are not many copies of them in circulation. Sometimes that’s because not many copies of them were ever printed, and sometimes because they are in a foreign language not many people speak. As the bookseller who sold them told me, there are not many people in the world who would care about them in the same way people might care about the first edition of Oliver Twist, but that doesn’t mean they’re not valuable.
“There are different types of value,” the bookseller said. “There's monetary value, obviously, but there are a lot of other types of value. There's historical value, intellectual value, sentimental value. All sorts of things, and all of those the AI companies don't care about. They just want the content as a bunch of words strung together.”
Amazon provided its statement above but did not respond to questions about why it cut the books it scanned, how many facilities like this it has across the world, and their response to the backlash from people upset that AI companies are destroying books.