Google already did this ~15 years ago with the google library project, but they didn’t buy books. They took them from libraries and then as a result of scanning old and rare books, they were generally damaged or destroyed.
I know. I saw it first hand. There wasn’t an incinerator or anything, but the carts full of books that disintegrated half the time. They had quotas for pages scanned and falling below them was a firable offense. You had to rush through it and there was no process to set aside books that disintegrated when the pages were turned.
Not trusting google but IMO digitizing books is completely different from training an LLM on them and subsequently destroying them.
The originals getting damaged is certainly not ideal, but if the contents are being preserved it’s not as bad.
I know google isn’t giving away that stuff for free either.
I promise google used the scanned books with OCR to train their LLMs. This sort of thing isn’t even new from them.
As long as they don’t delete the PDFs afterwards it’s still categorically different in my eyes.
EoD removing the text from existence is the crime.
This is the legal way.
This is what’s basically mandated, by rules about copyright. They’re not allowed to just grab a torrent of existing scans - but they can create a backup of a physical copy. They cannot then sell that copy, or they’d have to delete their backup. Since the most pragmatic way to scan a zillion books is to remove the spine and scan the stack of pages… that concern solves itself.
This is the direct result of people complaining about piracy - as if Lemmy gives a shit about piracy in any other context. And now twice a week we have headlines acting like they’ve uncovered a library firebombing campaign.
It’s one copy of every book. They don’t need to do it for anything older than Steamboat Willie. Their motivation is quantity, so they’re not bidding millions toward illuminated manuscripts, they’re taking obscure nonsense by the shovel-full. If it’s rare it’s because nobody gave a shit.
And it’s probably preserved in torrents anyway.
I hate when comments are 50% right and 50% infuriatingly incorrect.
…
It’s one copy of every book.
Per entity that wants to train their AI. So probably never just one copy of any given book.
…
This is the direct result of people complaining about piracy
People complain about piracy? What alternate timeline am I in right now? Or am I still in the one where it’s primarily Corporations that complain about piracy.
If you mean people complaining about AI’s stealing art/etc. AI’s don’t pirate, they perjure other people’s work as their own. Massive difference. I didn’t download the car and sell it like I built it.
…
They don’t need to do it for anything older than Steamboat Willie. Their motivation is quantity, so they’re not bidding millions toward illuminated manuscripts, they’re taking obscure nonsense by the shovel-full.
Lack of need isn’t a restriction. Indiscriminately seeking quantity implies they’re not checking if it was made before steamboat willy.
…
If it’s rare it’s because nobody gave a shit.
Most incorrect thing you said. Not how physical objects work.
…
And it’s probably preserved in torrents anyway.
Bro I literally collect books made after steam boat willy that never made it to the internet.
And now twice a week we have headlines acting like they’ve uncovered a library firebombing campaign.
Maybe you don’t read books so you don’t care as much as people who do. But in the books I’ve read mass book destruction is called book burning and it’s never contributed to positive outcomes for anyone working class.
How else do you think mass-produced things become rare? Either they printed a zillion copies and nearly all of them were thrown away, or the print run was tiny for lack of demand. Action Comics #1 is worth a fortune and Puget Sound: Proceedings Of A Seminar Held January 21, 1987 is not.
Per entity that wants to train their AI.
So like twelve copies? There’s no explosion of mom & pop AI startups that would justify the comparisons to Nazi book burning that are in every thread about this. If any work’s last physical copy is threatened, it was probably headed for the trash. Is it better-off being lost entirely, with nobody having a scan?
And yes, people complained when these companies did the obvious thing and grabbed a torrent. Just as loudly as they’re complaining about them doing this on the up-and-up.
If any work’s last physical copy is threatened, it was probably headed for the trash.
The fact that you see no problem with that is why I no longer respect you intellectually enough to consider you improving my lemmy experience.
have a nice life bro
deleted by creator
Preserving books that are falling apart isn’t really the same as intentionally destroying so that the copyright can’t be traced.
Bit extreme but I genuinely think the humans that decided to do this should spend the rest of their lives in prison.
Silly person, living, breathing humans didn’t make the decision! A non-human person in the form of Google made the decision and as such can’t be held liable for any harm done.
Besides, there’s no room in prison what with all the minor drug offenses locking people away for years.
Can you imaging that person probably has kids and keep them in 9-6 day care + after school programs so they can work on their jobs to destroy books?
Not a fan of the fact they’re doing this (obviously), given that I despise the waste and societal consequences AI companies are bringing upon us, but before anyone assumes they’re just doing this for the sake of being evil or something like that:
Destructive Scanning is the cheapest and easiest to automate methods of scanning books. This is fairly common for any large-scale digitizing projects that aren’t dealing with old, limited-in-supply books. (in this case, it’s just books before 2022, which still likely have many copies in circulation)
They’re also doing this to comply with court requirements for fair use. According to Ars Technica, a judge ruled that the destructive scanning qualified as fair use, only because Anthropic had spent the money to buy the books outright, destroyed them afterwards, and not distributed digital copies after.
The alternative is either Anthropic pirating the books (authors get paid $0), or buying the books, then distributing them back on the market through resellers (authors get paid the cost of the book, but then someone else buys the same copy secondhand and the author misses out on what would have otherwise been a fresh sale)
The alternative is…
I mean there is one you missed out but ok
What?
Shut it all down
I mean… sure, but that’s obviously not something that’s just gonna forcibly happen (at least not soon), nor is it a decision they’re gonna make on their own.
I was just describing the alternatives from the perspective that all the AI companies are trying to do things like this to compete with one another, so if they’re going to get the data from these books one way or another, there’s a limited set of ways that they can actually do that.
I’d love to see all these AI companies shut down, and the higher-ups responsible held accountable for intentionally destroying communities and the environment, as well as permanently poisoning online discourse and our ability to perceive what’s real and what’s not… but I just don’t see that as something the system will allow to happen, at least not currently.
Judge should be aware that public libraries exist within the law. Destruction of books should not be a legal requirement, books scanned without destruction could go to a public library without violating laws.
This seems to be a cost savings method of scanning cheap books in bulk. Not difficult to set up a different workflow for rare books.
Destruction of books should not be a legal requirement, books scanned without destruction could go to a public library without violating laws.
Yeah, I wish this was the case. I don’t think that under Anthropic’s prior ruling at least, that it would be allowed though, since after donating to the library (assuming they non-destructively scanned the books) they would then have to discard the entire set of digital copies they made, as that would mean they copied the work, and thus didn’t engage in a legal “transformation” of that work that maintains only one copy, and thus there would be no reason for them to do it in the first place.
I wish our copyright law could at least consider a donation to a library as equivalent to destruction of a work so companies could keep a scan but not have it be considered a copy. That way at least we’d be getting the giant influx of books going to libraries from AI training bs.
I heard someone on another social media site say that if the only physical copies are gone then it’s hard to prove copyright infringement.
So they are intentionally shreading books knowing that they can’t be sued for copyright infringement if there is no copy of the og material to prove it.
I have no clue if that’s true, but that sounds pretty far fetched imo.
As already mentioned, the books being purchased are not like, the only sole remaining copies of the work. Many of the authors are probably still alive, and there’s likely thousands or millions of each book distributed all over the country of origin, if not the world, in the hands of individuals, bookshops, libraries, archives, etc.
A judge said it was fair use when they purchased and solely used the books for AI training without redistributing them or a copy of them afterwards, so they’re just doing that in order to not get sued again. Not much more to it.
Aren’t some of the books considered rare ? I mean that’s why people are mad.
And yeah I agree they already use stolen works.
I’m just saying I heard someone say that.
I’m curious, myself, if it holds true or not. I don’t know enough about copyrights.
But I do know Google has made quite a lot of effort to scan books. Ive found stuff on Google books that is legit 100 years old in German research on optics. (I study depth perception). And I was really surprised they had a photocopy of it.
But maybe Google wanted a hefty price for access. Maybe a subscription.
Aren’t some of the books considered rare ? I mean that’s why people are mad.
I looked it up further, and while I’ve seen some people claim they’re using rare books, as far as I can tell that’s just based on a single bookseller saying that some of the books he distributed to ISBNdb (which the AI companies are buying through) were “rare or out of print”, but he didn’t provide any details on what those books were or how rare they were exactly, and he’s also the only source I’ve seen for that entire claim, so I’m not really sure how common that would actually be if he’s literally the one single person they could find who sold rare books to them.
Regardless, I still obviously am not a fan of that happening, I’d much rather that those books could be archived, even if that meant destruction but then digitizing them in, say, the Internet Archive instead, but at the end of the day I just wanted people to know that there wasn’t exactly no reason for them doing it the way they are. It’s not just malice for the hell of it, they got a court order that said how they could do it legally, so they did.
The destruction part is for the scanning. Still a bad look, and copyright issues seem to apply.
As a sad consolatation prize, In some cases this might be the only way certain old books survive. Some books from when I was young were just OK and they basically cant be found anymore. Try finding a complete collection of the ubiquitous “choose your own adventures” books. There were 184 of those published. There were 49 TSR Dungeons and Dragons Endless Quest books (basically choose your own adventure in d+d) and those are like 20 bucks each now, if you can find them at all. They are 50 years old. Anyone Remember Mountain of Mirrors, Pillars of Pentegarn, or Revenge of the Rainbow Dragons?
Good Times.Except they’re not surviving, are they? They’re being harvested for semantic associations and then destroyed.
Anthropic could easily offset the PR blow by offering a free digital library of all the works they’re destroying, but of course that would require them to actually pay a fair price to the owners.
Space Hawks!
Straight out of Rainbow’s End by Vernor Vinge.
https://en.wikipedia.org/wiki/Rainbows_End_(Vinge_novel)
Straight from the Torment Nexus:
Alex Blechman (@AlexBlechman) tweeted: Sci-Fi Author: In my book I invented the Torment Nexus as a cautionary tale.
Tech Company: At long last, we have created the Torment Nexus from classic sci-fi novel Don’t Create The Torment Nexus. 8 Nov 2021
Con’t move forward without destroying the lessons of the past. Capitalists, probably…




