• Prove_your_argument@piefed.social
    link
    fedilink
    English
    arrow-up
    49
    ·
    vor 3 Tagen

    Google already did this ~15 years ago with the google library project, but they didn’t buy books. They took them from libraries and then as a result of scanning old and rare books, they were generally damaged or destroyed.

    I know. I saw it first hand. There wasn’t an incinerator or anything, but the carts full of books that disintegrated half the time. They had quotas for pages scanned and falling below them was a firable offense. You had to rush through it and there was no process to set aside books that disintegrated when the pages were turned.

    • daannii@lemmy.world
      link
      fedilink
      English
      arrow-up
      3
      ·
      vor 3 Tagen

      Preserving books that are falling apart isn’t really the same as intentionally destroying so that the copyright can’t be traced.

  • AmbitiousProcess (they/them)@piefed.social
    link
    fedilink
    English
    arrow-up
    14
    arrow-down
    1
    ·
    vor 3 Tagen

    Not a fan of the fact they’re doing this (obviously), given that I despise the waste and societal consequences AI companies are bringing upon us, but before anyone assumes they’re just doing this for the sake of being evil or something like that:

    Destructive Scanning is the cheapest and easiest to automate methods of scanning books. This is fairly common for any large-scale digitizing projects that aren’t dealing with old, limited-in-supply books. (in this case, it’s just books before 2022, which still likely have many copies in circulation)

    They’re also doing this to comply with court requirements for fair use. According to Ars Technica, a judge ruled that the destructive scanning qualified as fair use, only because Anthropic had spent the money to buy the books outright, destroyed them afterwards, and not distributed digital copies after.

    The alternative is either Anthropic pirating the books (authors get paid $0), or buying the books, then distributing them back on the market through resellers (authors get paid the cost of the book, but then someone else buys the same copy secondhand and the author misses out on what would have otherwise been a fresh sale)

    • sos242@thelemmy.club
      link
      fedilink
      English
      arrow-up
      3
      ·
      vor 3 Tagen

      Judge should be aware that public libraries exist within the law. Destruction of books should not be a legal requirement, books scanned without destruction could go to a public library without violating laws.

      This seems to be a cost savings method of scanning cheap books in bulk. Not difficult to set up a different workflow for rare books.

      • AmbitiousProcess (they/them)@piefed.social
        link
        fedilink
        English
        arrow-up
        1
        ·
        vor 3 Tagen

        Destruction of books should not be a legal requirement, books scanned without destruction could go to a public library without violating laws.

        Yeah, I wish this was the case. I don’t think that under Anthropic’s prior ruling at least, that it would be allowed though, since after donating to the library (assuming they non-destructively scanned the books) they would then have to discard the entire set of digital copies they made, as that would mean they copied the work, and thus didn’t engage in a legal “transformation” of that work that maintains only one copy, and thus there would be no reason for them to do it in the first place.

        I wish our copyright law could at least consider a donation to a library as equivalent to destruction of a work so companies could keep a scan but not have it be considered a copy. That way at least we’d be getting the giant influx of books going to libraries from AI training bs.

    • skisnow@lemmy.ca
      link
      fedilink
      English
      arrow-up
      5
      arrow-down
      1
      ·
      vor 3 Tagen

      The alternative is…

      I mean there is one you missed out but ok

          • AmbitiousProcess (they/them)@piefed.social
            link
            fedilink
            English
            arrow-up
            2
            ·
            vor 3 Tagen

            I mean… sure, but that’s obviously not something that’s just gonna forcibly happen (at least not soon), nor is it a decision they’re gonna make on their own.

            I was just describing the alternatives from the perspective that all the AI companies are trying to do things like this to compete with one another, so if they’re going to get the data from these books one way or another, there’s a limited set of ways that they can actually do that.

            I’d love to see all these AI companies shut down, and the higher-ups responsible held accountable for intentionally destroying communities and the environment, as well as permanently poisoning online discourse and our ability to perceive what’s real and what’s not… but I just don’t see that as something the system will allow to happen, at least not currently.

    • daannii@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      ·
      edit-2
      vor 3 Tagen

      I heard someone on another social media site say that if the only physical copies are gone then it’s hard to prove copyright infringement.

      So they are intentionally shreading books knowing that they can’t be sued for copyright infringement if there is no copy of the og material to prove it.

      • AmbitiousProcess (they/them)@piefed.social
        link
        fedilink
        English
        arrow-up
        2
        ·
        vor 3 Tagen

        I have no clue if that’s true, but that sounds pretty far fetched imo.

        As already mentioned, the books being purchased are not like, the only sole remaining copies of the work. Many of the authors are probably still alive, and there’s likely thousands or millions of each book distributed all over the country of origin, if not the world, in the hands of individuals, bookshops, libraries, archives, etc.

        A judge said it was fair use when they purchased and solely used the books for AI training without redistributing them or a copy of them afterwards, so they’re just doing that in order to not get sued again. Not much more to it.

        • daannii@lemmy.world
          link
          fedilink
          English
          arrow-up
          2
          ·
          edit-2
          vor 3 Tagen

          Aren’t some of the books considered rare ? I mean that’s why people are mad.

          And yeah I agree they already use stolen works.

          I’m just saying I heard someone say that.

          I’m curious, myself, if it holds true or not. I don’t know enough about copyrights.

          But I do know Google has made quite a lot of effort to scan books. Ive found stuff on Google books that is legit 100 years old in German research on optics. (I study depth perception). And I was really surprised they had a photocopy of it.

          But maybe Google wanted a hefty price for access. Maybe a subscription.

          • AmbitiousProcess (they/them)@piefed.social
            link
            fedilink
            English
            arrow-up
            2
            ·
            vor 3 Tagen

            Aren’t some of the books considered rare ? I mean that’s why people are mad.

            I looked it up further, and while I’ve seen some people claim they’re using rare books, as far as I can tell that’s just based on a single bookseller saying that some of the books he distributed to ISBNdb (which the AI companies are buying through) were “rare or out of print”, but he didn’t provide any details on what those books were or how rare they were exactly, and he’s also the only source I’ve seen for that entire claim, so I’m not really sure how common that would actually be if he’s literally the one single person they could find who sold rare books to them.

            Regardless, I still obviously am not a fan of that happening, I’d much rather that those books could be archived, even if that meant destruction but then digitizing them in, say, the Internet Archive instead, but at the end of the day I just wanted people to know that there wasn’t exactly no reason for them doing it the way they are. It’s not just malice for the hell of it, they got a court order that said how they could do it legally, so they did.

  • humanspiral@lemmy.ca
    link
    fedilink
    English
    arrow-up
    5
    ·
    vor 3 Tagen

    The destruction part is for the scanning. Still a bad look, and copyright issues seem to apply.

  • kreskin@lemmy.world
    link
    fedilink
    English
    arrow-up
    6
    ·
    vor 3 Tagen

    As a sad consolatation prize, In some cases this might be the only way certain old books survive. Some books from when I was young were just OK and they basically cant be found anymore. Try finding a complete collection of the ubiquitous “choose your own adventures” books. There were 184 of those published. There were 49 TSR Dungeons and Dragons Endless Quest books (basically choose your own adventure in d+d) and those are like 20 bucks each now, if you can find them at all. They are 50 years old. Anyone Remember Mountain of Mirrors, Pillars of Pentegarn, or Revenge of the Rainbow Dragons?
    Good Times.

    • verdigris@lemmy.ml
      link
      fedilink
      English
      arrow-up
      6
      arrow-down
      1
      ·
      vor 3 Tagen

      Except they’re not surviving, are they? They’re being harvested for semantic associations and then destroyed.

      Anthropic could easily offset the PR blow by offering a free digital library of all the works they’re destroying, but of course that would require them to actually pay a fair price to the owners.

  • notsure@fedia.ioBanned
    link
    fedilink
    arrow-up
    2
    arrow-down
    1
    ·
    vor 3 Tagen

    Con’t move forward without destroying the lessons of the past. Capitalists, probably…