The symbol for Amazon’s VGT3, the Las Vegas facility where it scans book for AI training data.

Amazon is buying massive quantities of books, scanning them for AI training data, and destroying them in the process.

A 404 Media investigation was able to reveal Amazon’s book buying operation, which hasn’t been previously reported, by placing a tracking device in a rare book we suspected would be acquired by an AI company for training data, and following it around the country to its final destination.

That final destination was an Amazon warehouse in Las Vegas, Nevada. Amazon employees who work at this location say all they do is receive massive shipments of printed books which they then cut the bindings off in order to scan the books more quickly. The printed book is destroyed in the process. The logo of the Amazon team that works at this warehouse, called VGT3, is a dinosaur, brandishing its teeth and with a book in its hands.

“Amazon purchases books through commercial channels to help develop and improve the products and services our customers use,” an Amazon spokesperson told me in a statement.

The world’s AI companies are constantly looking for, and spending extreme resources to locate, more material to train their AI models. With books, that sometimes means destroying them in the process, something that large parts of the public have spoken up against, and which we can now confirm Amazon is doing.

In July, I published a story about booksellers who reported a historical spike in sales starting in the past year. They suspected this spike in sales was due to AI companies acquiring any books they can in search of new training data. Printed books are valuable as training data because a lot of the text they contain is not readily available on the internet, which AI companies have already scraped. The data is also conveniently organized and, if the book was printed before 2022, is guaranteed to be free of AI-generated text, which can make any AI model that is trained on it worse via a recursive process called “model collapse.”

📖 Do you know work at a facility where you scan books? I would love to hear from you. Using a non-work device, you can message me securely on Signal at @emanuel.404. Otherwise, send me an email at emanuel@404media.co.

These booksellers suspected AI companies were behind these large bulk purchases because of the high number of books they were buying, the seemingly random choice of books, and the fact that these buyers, unlike libraries and universities, did not seem price sensitive at all. But booksellers couldn’t say for certain who was behind the large purchases because the marketplaces where they sell their books keep the buyers anonymous. When an order comes in, a bookseller ships the sold books to a warehouse operated by the marketplaces, where books are sorted and then sent to the buyer.

In July, one bookseller told me they received a very large order of around 1,000 books on Biblio, one of these marketplaces. The seller agreed to put an Apple AirTag provided by 404 Media in one of the books included in this order so we could see where the book was going. And by extension, which company, AI or otherwise, was behind this massive order. 404 Media granted the bookseller anonymity because they worried sharing this information would harm their business. Biblio did not respond to a request for comment.

  • BilSabab@lemmy.world
    link
    fedilink
    English
    arrow-up
    2
    ·
    9 hours ago

    local AI companies are the reason Ukrainian book publishing industry experienced an uptick in sales lately. Bros stock up entire catalog at once and no one complains because its a hefty paycheck.

    • Tollana1234567@lemmy.today
      link
      fedilink
      English
      arrow-up
      4
      arrow-down
      1
      ·
      10 hours ago

      it should be labeled as “limited amount of books in circulation”, not rare. rare means 1 or only handful of compies.

  • hard_zero1@discuss.tchncs.de
    link
    fedilink
    English
    arrow-up
    11
    ·
    1 day ago

    If this was done by trustworthy organizations as an effort to digitize and preserve all the books, I would appreciate it. Maybe an association of libraries should do that and, to get the cost back, sell the digital versions / ebooks to the AI companies for an additional price). Then, at least, we would not loose the contents of those rare books to the AI conpanies and prevent them from obtaining a monopoly on the data. And each book would only get destroyed once.

    But libraries/bookshops are probably not allowed to sell digital versions, and AI companies are not allowed to use borrowed ebooks.

    • Duamerthrax@lemmy.world
      link
      fedilink
      English
      arrow-up
      10
      ·
      22 hours ago

      That exists already. Archive.org has a digitization service, but then archive.org would put a public copy up and the AIbros wouldn’t be the sole owners of that training data.

      https://digitization.archive.org/

      There’s also methods to digitize books without destroying the book and for the purpose of AI training that should be more then sufficient. AIbros are just so arrogant to think that other people haven’t solved the problem already or that their time is too valuable to be slowed by proper methods.

      https://www.diybookscanner.org/

      • ggtdbz@lemmy.dbzer0.com
        link
        fedilink
        English
        arrow-up
        1
        ·
        8 hours ago

        I remember looking into this, the number of books I have that I’d like to scan are not enough to justify building one of these unless they can be folded up and passed on.

        Might be best to just get a plexiglass slab with some kind of anti glare covering and try to get flat pages

    • frongt@lemmy.zip
      link
      fedilink
      English
      arrow-up
      3
      ·
      1 day ago

      Yeah. Libraries got in trouble during COVID for relaxing their ebook borrowing rules. AI companies don’t give a shit and are happy to settle lawsuits for a fraction of the money they take in from investors.

  • queermunist she/her@lemmy.ml
    link
    fedilink
    English
    arrow-up
    4
    ·
    1 day ago

    I don’t understand the concept of a “rare” book. Every book should be infinitely reproducible, the fact that they aren’t is a crime.

      • queermunist she/her@lemmy.ml
        link
        fedilink
        English
        arrow-up
        2
        arrow-down
        2
        ·
        12 hours ago

        The text is what makes it valuable though? I guess there’s some value in, like, the actual literal physical copy too, but that has nothing to do with rarity. A bible can have historical value because of who owned it, but that’d be strange to describe as “rare” I think.

    • Ebby@lemmy.ssba.com
      link
      fedilink
      English
      arrow-up
      10
      ·
      edit-2
      14 hours ago

      There are many ways a book can be rare. First editions, signed copies, and books with little demand. Can’t fire up the printers and make those again.

      In the case of little demand, I have a book “Two thousand leagues under the seas” not the more common “Twenty thousand leagues under the sea”. It’s rare in that it’s very difficult to find more literal translation than the version we are accustomed with. I suspect search engines simply think I’ve made a typo. Either way, why print it if there is little demand or can’t find the book?

      • zaphod@sopuli.xyz
        link
        fedilink
        English
        arrow-up
        1
        ·
        3 hours ago

        The title doesn’t make sense, or is it an extremely shortened version and it’s a pun on the fact that it shortened the journey? Do you happen to know who the translator was?

      • Leomas@lemmy.world
        link
        fedilink
        English
        arrow-up
        3
        ·
        12 hours ago

        I agree with your point overall, but why is the book two thousand leagues under the sea more literal, when the original is called “Vingt Mille Lieues sous les mers”, is a league 10 times as much, as a Lieue? Genuinely curious (and hard to google)

        • Ebby@lemmy.ssba.com
          link
          fedilink
          English
          arrow-up
          1
          ·
          10 hours ago

          Huh, good question. I remember a forward in the book about the difference translations, but I will admit that was back in middle school. I still have the book. Perhaps I’ll try to dig it up tomorrow.

          I did however find a new translation online that sounds interesting. I remember all the scientific names suuuucked back then and even my copy said something along the lines of “abridged for sake of brevity” on one of his journal entries. Even my translation noped out of that.

    • Solrac@lemmy.world
      link
      fedilink
      English
      arrow-up
      3
      ·
      11 hours ago

      You don’t hate them enough. The only response is to do onto them, as they do to us, as they do to these books

    • Test_Tickles@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      ·
      1 day ago

      Look up Google and Project Ocean. Google already did this 2 decades ago. They even went to great lengths to build machines that would very slowly and gently turn pages and non-destructively scan books.
      They already have a digital library of 25 million books just sitting there, ready to be instantly and non-destructively copied infinitely.

      • Mirshe@lemmy.world
        link
        fedilink
        English
        arrow-up
        2
        ·
        1 day ago

        No no, that’s too slow and slow is expensive. Time is money and we have to move faster and break more things in order to disrupt the market. /s

  • ominous ocelot@leminal.space
    link
    fedilink
    English
    arrow-up
    1
    ·
    edit-2
    1 day ago

    Are we talking about 14th century handcrafted masterpieces or 2000s university textbooks? Do they destroy cultural heritage items or dusty low value sold-by-weight books?

    I need to know if I have to feel angry and agitated or indifferent.

    • kbobabob@lemmy.dbzer0.com
      link
      fedilink
      English
      arrow-up
      1
      ·
      7 hours ago

      We will never know because they didn’t give any information. No info about the book(books?) that were tracked or even how they tracked them.

    • frongt@lemmy.zip
      link
      fedilink
      English
      arrow-up
      6
      ·
      1 day ago

      The latter. These are books that have been sitting on the shelf in a warehouse for years. They’re not sold by weight, but they are stuff like random technical manuals for stuff very few people care about.

      My only hope is that an eventual lawsuit forces the AI companies to release the digitized versions to an archive or library.

    • Test_Tickles@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      ·
      1 day ago

      Textbooks are not “rare”, nor are they something that universities and libraries would be price sensitive about. The books are being bought from booksellers who make a living buying and selling rare books. So, while I doubt that they are all going to “masterpieces”, they are going to be books valuable enough to support an industry of people and expensive enough that universities and libraries would be price sensitive about them.

    • skisnow@lemmy.ca
      link
      fedilink
      English
      arrow-up
      1
      ·
      1 day ago

      We’re almost certainly talking about literally everything a bot scraping catalogues decides it doesn’t have. I highly doubt it’ll be discriminating.

    • SnailMagnitude@mander.xyz
      link
      fedilink
      English
      arrow-up
      0
      ·
      1 day ago

      I don’t think they are devouring the Garima Gospels, Cuneiform tablets and deleting hieroglyphs from granite but the worry is if we don’t stop them now it’s a slippery slope and the VHS tapes will be next.

      • ominous ocelot@leminal.space
        link
        fedilink
        English
        arrow-up
        1
        ·
        edit-2
        1 day ago

        Without details it feels, like a fear, uncertainty, doubt strategy just for the clicks, to me. I’m missing crucial information for an informed decision. The article is not suitable for much more than igniting anger.

        LLM and AI companies rise a lot of valid criticism, besides its undeniable pros. I don’t like a witch hunt just because it is AI and someone tells me to be angry. I need more facts before I go fetch my pitchfork.

        I.e. I dont get why they destroy the books. Google Books did fine without unbinding/ destroying books. It would help to know what kind of books are processed. And what their process is.

        I don’t have information if the OCRed data is made available to the public. Which would be nice, since the article talks about data that is not yet digitally available.

        • themachinestops@lemmy.dbzer0.comOP
          link
          fedilink
          English
          arrow-up
          1
          ·
          1 day ago

          This is what I believe they are doing:

          https://finance.biggo.com/news/37a2899b-5f1a-4571-9bb5-1156c0ac4605

          _**The owner of MW Books in Ireland also reports that since May, the shop has been receiving large orders with wildly varied content. “The orders are piecemeal, presumably machine-driven, and there’s very little thematic rhyme or reason to them: everything from 18th-century African agricultural implements to biographies of 1950s racing drivers,” he says. “Personally, I don’t think we have the right to dictate what customers do with books after they buy them, but the selection is certainly intriguing.”

          Another anonymous UK used bookseller reveals that some seemingly random orders, possibly from different buyers, all end up being shipped to the same address. He provided a delivery postcode used by different buyers, which corresponds to a cluster of freight warehouses near London Heathrow Airport. The bookseller says that since January, he has received orders covering 6,000 books from similar buyers, amounting to thousands of pounds. These buyers are willing to pay “top prices” and haven’t asked for discounts even when purchasing large quantities at once. “That doesn’t happen very often. It’s disruptive to the used book trade.”**_

          They are probably buying books that don’t have any digital copies since they are perfect for training AI. As for why they are using destructive scanning, I believe it is because either destructive is cheaper and faster or they don’t want the competition to have access to the training data. Since these are not easy to find books, Amazon sells books, any book they try to obtain is probably extremely hard to find anywhere.

          • dhork@lemmy.world
            link
            fedilink
            English
            arrow-up
            1
            ·
            6 hours ago

            Interesting that the buyers are willing to pay “top prices” and yet unwilling to spend the money to scan these books non-destructively.

  • antonim@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    1 day ago

    So, what is the book that they tracked? What other books were in the order alongside it?

  • altkey (he\him)@lemmy.dbzer0.com
    link
    fedilink
    English
    arrow-up
    0
    ·
    1 day ago

    The symbol for Amazon’s VGT3, the Las Vegas facility where it scans book for AI training data.

    I thought it’s 404media’s obviously satirical preview picture to dab on aibros, but the truth is even more weird

    • Leon@pawb.social
      link
      fedilink
      English
      arrow-up
      0
      ·
      1 day ago

      Hardly surprising that slop bros have no taste. I’m not an artist and it’s not exactly difficult to point out the flaws with that logo. Someone could’ve stopped it, but they didn’t care enough.