• Dave.@aussie.zone
    link
    fedilink
    arrow-up
    33
    ·
    4 days ago

    At this point, I’m just hoping that when we are sifting through the smoking wreckage of the AI crash we’ll find a hundred million books that we can add to Anna’s Archive.

    And some RAM. RAM would be nice.

      • BigPotato@lemmy.world
        link
        fedilink
        arrow-up
        9
        ·
        4 days ago

        They mean the PDFs of them or ePub or whatever they’re scanning them into. Yeah, the original is gone but imagine if the Library of Alexandria went up in flames but every ash contained a complete work for free to everyone to read at their leisure.

        Oh, and Ram.

        • Eldritch@piefed.world
          link
          fedilink
          English
          arrow-up
          5
          ·
          4 days ago

          Storage space is valuable for their AI stuff. Why would they keep around copies like that on valuable storage space if they don’t think they’re going to be using them again once they’ve trained their model? I really hope they are scanning them to a format or something like that that. But I doubt it. These aren’t the most forward looking or intelligent people you’ll find. I wouldn’t be at all surprised to find out that they never did anything more than scan it into train the model and then flush it from the system. It’s 100% on brand for them.

          • Dran@lemmy.world
            link
            fedilink
            arrow-up
            7
            ·
            4 days ago

            You always keep training data because you never know what you need for the next generation. Ironically they might be preserved (privately though unfortunately) pretty well

      • MonkeMischief@lemmy.today
        link
        fedilink
        arrow-up
        2
        ·
        4 days ago

        They are ripping them from their spines and pulping them once the have what they want.

        They treat their books like they treat their employees.

    • Einskjaldi@lemmy.world
      link
      fedilink
      arrow-up
      5
      ·
      4 days ago

      They aren’t never seen before unique rare books, they’re just out of print and not lots of copies. They’re destroying them to sidestep copyright concerns.

  • WatDabney@sopuli.xyz
    link
    fedilink
    arrow-up
    16
    ·
    4 days ago

    It’s not simply that they don’t care about the value of the books.

    The tech oligarchs are trying to convert information into property, with themselves as sole owners

    • cmbabul@slrpnk.net
      link
      fedilink
      arrow-up
      1
      ·
      3 days ago

      And it’s not even just about being a monopoly for profit reasons, this time, if past knowledge becomes completely centralized by the powerful it can no longer be trusted

  • UnderpantsWeevil@lemmy.world
    link
    fedilink
    English
    arrow-up
    19
    arrow-down
    3
    ·
    4 days ago

    Rare Books

    I’m still waiting to see what an example of a “Rare Book” is supposed to be. Are these first edition copies of To Kill A Mockingbird and Ulysses? Or are we just talking about books by new authors that were never widely distributed or reprinted, because their sales numbers were no good.

    This post is for paid members only

    I guess I’ll never know.

    • Jo Miran@lemmy.ml
      link
      fedilink
      arrow-up
      14
      arrow-down
      1
      ·
      edit-2
      4 days ago

      There are so many niche instructional books that are now out of print. If these books were scanned (yes, destroyed in the process) and then added to a free and open book archive for all to access, I wouldn’t have a problem with it. I have a problem with these AI training centers because all that knowledge is just going into a back hole.

      There are music theory books and penmanship books I’m actively hunting for through piracy and ebay because they are out of print and rapidly getting lost to time. If I find them, I plan to scan and share them, not for piracy but because the original authors works do not deserve to disappear just because they no longer generated profit.

      • mushroommunk@lemmy.today
        link
        fedilink
        arrow-up
        4
        ·
        4 days ago

        That’s not even touching in the price. Sometimes they sit there untouched simply because whoever happens to have one of the few copies put a ridiculous price on it.

        I’m not saying it needs to be free (I mean I do but that’s a separate argument). I’m saying many of these sellers want as much money as physically possibly with no real world basis for the prices, which is why they sit there.

        There’s several old Welsh poetry books I’ve gone looking for only to find copies marked for a thousand bucks pricing out any regular person and leaving only companies with stupid money to burn or maybe one day a collector.

        • Jo Miran@lemmy.ml
          link
          fedilink
          arrow-up
          3
          ·
          4 days ago

          One of the books in my hunt list is the one below. It originally sold for $16 but is now out of print. Notice the reseller’s price.

          • mushroommunk@lemmy.today
            link
            fedilink
            arrow-up
            4
            ·
            4 days ago

            Yup, and it’s just gonna sit there and go up in price as no one buys it because it’s “rare” and out of print. I feel ya

            • floofloof@lemmy.ca
              link
              fedilink
              arrow-up
              1
              ·
              4 days ago

              The big tech companies will drive up the price of second-hand books, just as they drove up the price of computer hardware and electricity.

    • frongt@lemmy.zip
      link
      fedilink
      arrow-up
      5
      arrow-down
      1
      ·
      4 days ago

      Rare meaning uncommon, hard to find, not many copies are available.

      But that doesn’t necessarily mean valuable. These have been sitting on shelves in warehouses unsold for a long time. No one else wanted them.

      But let me be clear, I’m against their destruction for private use. The digitized copies should be made freely available, if it all possible.

      • UnderpantsWeevil@lemmy.world
        link
        fedilink
        English
        arrow-up
        1
        ·
        4 days ago

        But that doesn’t necessarily mean valuable.

        Well, the value in the books is that they (a) haven’t been digitized yet and (b) aren’t contaminated with AI autowriting.

        The buyers that would find the most value in these books are, themselves, AI training companies.

        But let me be clear, I’m against their destruction for private use.

        shrug That’s how books are digitized. You break the spine, split out the individual pages, and run them through an industrial scanner. I guess we can go back to reading from scrolls to alleviate this step. Past that, Idk what the problem is.

        The digitized copies should be made freely available, if it all possible.

        I don’t hate this idea. But I might argue that the Library of Congress should be digitizing published works as part of the copywriting process anyway. And, in fairness, the LoC currently hosts 21 petabytes of digitally archived data across 91 million unique works in 470 languages.

        This keeps getting floated as some kind of scandal. I see it compared to “The burning of the Library of Alexandra” over and over again. But it appears to be nothing more than another, more primitive form of data harvesting of documents barely more valuable than Reddit shitposts. Less Alexandra and more the graffiti scribbled across Pompeii.

    • Zaktor@sopuli.xyz
      link
      fedilink
      English
      arrow-up
      1
      ·
      4 days ago

      It would likely be stuff under copyright or stuff out of copyright that wasn’t considered a priority in other book digitization efforts. 1930 is the current copyright-free publish date.

      So probably not a prestige edition of a famous work, but when buying in bulk like this who knows.

  • CannedYeet@lemmy.world
    link
    fedilink
    arrow-up
    5
    ·
    4 days ago

    The real problem here is copyright law. We need a system where registration is required to extend the copyright term beyond an initial relatively short term. That way we can identify orphaned works and make them available to the public.