AI companies are shredding rare books
Posted by anon373839 6 hours ago
Comments
Comment by squidbeak 5 hours ago
It pisses me off to reflect that they can sit on works until copyright expires, keeping them out of print. There's no real need for any of these so-called rare books to be rare while they're under copyright.
And related to this, the books that are in print are mostly only in print in the shittiest way. I often see well-made books from the 17th or 18th centuries which are still in good nick. It's ridiculous that in the 21st century, publication standards have fallen to the point where for most works a disposable format is the only type available - where no amount of money could buy a truly decent hardback copy.
If we have to have copyright laws, I'd like to see two changes to them.
When a publisher has no incentive to keep an edition in print, it should be available to any other publisher to print, without compensation to the original publisher, and with renegotiated royalties for the author.
And if the publisher keeps a book in print - but only in bestseller-grade materials, bogroll paper that furrows in any humidity and perfect binding that molts its pages a couple of dry seasons later - and if it refuses to print a durable hardback copy with signatures, good paper and decent print - something that will still be readable in several generations' time - any other publisher keen to have a crack at it should be able to, again without any compensation for the original publisher, though perhaps in this case, with matching royalties for the author.
Comment by giancarlostoro 3 hours ago
Issue is small streamers have no legal defense.
Edit, adding the channel I am referring to:
Comment by monknomo 3 hours ago
Comment by Ekaros 23 minutes ago
We do not continue to pay for most things once they are created. Unless they are continuous services. Artistic works should not be any different.
Comment by reorder9695 3 hours ago
Comment by monknomo 2 hours ago
Comment by marcosdumay 2 hours ago
Comment by qwytw 1 hour ago
> If it's good enough for patents, I don't see why it isn't good enough for copyrights
Because there are fundamentally differ concepts and serve different purposes?
Comment by altmanaltman 2 hours ago
If I wrote it, I own the copyright on it, why should I or my future family give away something I worked really hard for? Why do only authors must care about public benefits?
Comment by thdr 3 hours ago
That would have a few undesirable consequences... for example, you wouldn't hire a 70 year old writer for your commercial project no matter how brilliant they are.
The complexity of our legal system is in many cases justified. The problems are often the numbers (duration of copyright protection etc.)
Comment by fn-mote 2 hours ago
This makes sense when you’re thinking of a painting or a book.
Who owns the copyright to Windows or MacOS? A corporation. How do you deal with that?
> you wouldn't hire a 70 year old writer for your commercial project no matter how brilliant they are
Commercial projects are works-for-hire and the copyright is not owned by the person who does the work.
The proposal for the limit of copyright needs to be refined.
In the US, the current rule is:
> For … a work made for hire, the copyright endures for a term of 95 years from the year of its first publication or a term of 120 years from the year of its creation
Comment by wiether 3 hours ago
> for example, you wouldn't hire a 70 year old writer for your commercial project no matter how brilliant they are
If you hire them, then you own the work you paid them to do, no?Comment by graeme 2 hours ago
You could do "life of author or X years, whichever is longer". Or include a period after death.
But you see how the complexities come in.
Comment by clort 2 hours ago
It is said that the vast vast majority of works don't earn anything significant after a few years in any case, meaning the only possible reason to have long copyrights is so that a very few people can get stinking rich. But those people already got rich, in the first few years.. society does not benefit from them getting richer.
20 years fixed term is my proposal.
Comment by TeMPOraL 2 hours ago
Comment by this_was_posted 2 hours ago
Comment by Ekaros 21 minutes ago
Comment by delecti 2 hours ago
Not that I'm arguing for 30 per se, just that I don't see what goals of copyright would be advanced more by adding an "or until death" complication.
Comment by qwytw 1 hour ago
Somebody decides to make a movie based on your book? You get nothing at all from it... The movie bit would be problematic even for books that were reasonably popular at the time. e.g the Witcher adaption came out almost exactly 20 years after the last book, for GOT it wasn't that far from being the case as well (at least for the initial volumes). Studios would be incentivized just to wait a couple of years to avoid paying anything.
I think it could be reasonably to have a fixed limit if the rights are held by corporations, though.
Comment by Always42 3 hours ago
Comment by ktm5j 3 hours ago
Comment by weberer 3 hours ago
Comment by AussieWog93 2 hours ago
Comment by duzer65657 1 hour ago
Comment by giancarlostoro 2 hours ago
Comment by fn-mote 2 hours ago
Think this through some more.
Authors and artists are still creating after that age.
Comment by inferniac 3 hours ago
sure, you don't have to pay royalties, but any other publisher can now publish too
Comment by cindyllm 3 hours ago
Comment by Ikatza 1 hour ago
Comment by kasey_junk 5 hours ago
Books in that time were _luxury_ goods. Most people could not afford them. One of the ways that was changed was to introduce cheap, mass produced bindings that were lower quality than the bespoke artisianal bindings done by specialist craftsmen.
You can still get custom bindings done. There exists whole niches on the internet of crafters that will take a production run book and strip its binding and make you extremely high quality and custom bindings and covers.
Comment by ryanmcbride 4 hours ago
With how the quality of things seems to have been degrading over the years (either real or just me getting older and experiencing the impermanence of all things) I've been trying to adopt an attitude of "if this practice existed before the industrial revolution, I can _probably_ do it" and it's been really great to learn how things were made before they had to be mass produced as cheaply as possible.
Comment by DrewADesign 4 hours ago
Comment by kasey_junk 4 hours ago
Comment by DrewADesign 2 hours ago
Comment by pfdietz 4 hours ago
Comment by DrewADesign 2 hours ago
Comment by Lutzb 4 hours ago
There is a market for these type of books, albeit a very small one.
Comment by ascagnel_ 3 hours ago
Comment by m4rtink 2 hours ago
Apparently it is getting harder to find people who can do that as most schools no longer have book binding as a course you can study.
Comment by tookmund 4 hours ago
At the risk of stating the obvious, any poorly made books from then wouldn’t have lasted this long and so you would never see them.
Comment by ortusdux 4 hours ago
Comment by ShinyLeftPad 4 hours ago
Comment by phoghed 4 hours ago
I’m sure there’s ongoing litigation, and better sources than this, but fair use was determined in June 2025 in a sf federal district court https://www.goodwinlaw.com/en/insights/publications/2025/06/...
Similar conclusion vs meta https://www.jw.com/news/insights-kadrey-meta-bartz-anthropic...
And the more recent $1.5B settlement did not overturn it https://www.reuters.com/world/us-judge-approves-anthropics-1...
Comment by EnergyAmy 5 minutes ago
Comment by sofixa 4 hours ago
Comment by bscphil 4 hours ago
Comment by rileymat2 4 hours ago
Comment by hn_acker 4 hours ago
Comment by nativeit 3 hours ago
Comment by shimman 4 hours ago
Comment by TeMPOraL 3 hours ago
In fact, the whole problem of shredding books (destructive format-shifting) was created by copyright laws in the first place, and that in itself is a concession hard won against the IP establishment - and all that way before LLMs became a thing.
Comment by phoghed 4 hours ago
Comment by roarcher 2 hours ago
Comment by rc5150 3 hours ago
Comment by phoghed 3 hours ago
> step aside
Isn’t that what I explicitly just did in the comment you replied to?
Comment by HedonicEscal8r 2 hours ago
Comment by ButlerianJihad 4 hours ago
Fair Use is not an activity that you engage in. Fair Use is not a category with criteria that you meet. Fair Use is not a precedent that paves the way for everything afterwards.
Fair Use is a defense that can be used in court when you’re named in a copyright lawsuit. Fair Use is how you justify your actions before the court finds infringement.
Comment by hn_acker 4 hours ago
> Notwithstanding the provisions of sections 106 and 106A, the fair use of a copyrighted work... is not an infringement of copyright.
[1] https://en.wikipedia.org/wiki/Google_LLC_v._Oracle_America,_....
Comment by phoghed 3 hours ago
Now if you look at how fair use is used colloquially, everyone understood what I meant except the autistic pedants.
Comment by Marsymars 2 hours ago
e.g. Canada doesn't have fair use, but from wiki on fair dealing in Canada, "According to the Supreme Court of Canada, it is more than a simple defence; it is an integral part of the Copyright Act of Canada, providing balance between the rights of owners and users."
Comment by gruez 4 hours ago
>You don’t know what “Fair Use” is.
This isn't the opinion of some armchair HN commenter. Actual judges have affirmed this, as other commenters in this thread has pointed out.
Comment by parineum 4 hours ago
The comment uses the same language as it's parent.
Comment by parineum 4 hours ago
Comment by arduanika 4 hours ago
Comment by DonsDiscountGas 2 hours ago
Comment by MemoryHoleHQ 4 hours ago
Comment by Aurornis 3 hours ago
This is purely a response to market demand. Publishers aren’t going to put in the extra expense of binding high quality versions of every book so it can occupy warehouse space while consumers everywhere buy the cheap paperback.
> And if the publisher keeps a book in print - but only in bestseller-grade materials, bogroll paper that furrows in any humidity and perfect binding that molts its pages a couple of dry seasons later - and if it refuses to print a durable hardback copy with signatures, good paper and decent print - something that will still be readable in several generations' time - any other publisher keen to have a crack at it should be able to, again without any compensation for the original publisher, though perhaps in this case, with matching royalties for the author.
This is all based on the idea that there is hidden demand for something, but publishers are choosing to deprive us all of it for reasons. That if we open up the laws, another company will come along and satisfy this hidden market opportunity and associated profits that publishers are declining to take.
The simpler explanation is that these high quality editions aren’t being published because the publishers have the data about demand for them. They know they won’t sell.
If the goal is preservation, laws forcing publishers to print on slightly nicer paper isn’t going to solve the problem. It needs to be a robust digital archive and it needs to exist somewhere other than in unsold warehouse inventory or some book collector’s shelf. You’re trying to solve a problem with last century’s technology.
Comment by kevin_thibedeau 4 hours ago
Those books predate the development of wood pulp paper. It isn't the publisher's fault they can't economically print on rag paper anymore.
Comment by ctolsen 3 hours ago
Comment by jolmg 3 hours ago
This doesn't affect publishers.
It affects humanity as a whole by having large private parties hoard books to partly destroy them. The covers and spines can have historical value too. There's not even a reason for these companies to release their scans to the public once the copyright expires.
It prevents proper preservation by archivists and preservationists.
Comment by cataphract 1 hour ago
Comment by nyeah 4 hours ago
Comment by trescenzi 2 hours ago
Comment by streetfighter64 4 hours ago
How do you adjudicate that? And wouldn't it just lead to loopholes such as "Ghost Printings" (cf. Ghost Flights https://en.wikipedia.org/wiki/Ghost_flight_(commercial_aviat... ) where the books are technically printed in the required volume but practically unavailable to customers through one method or another. Because, the cost of wastefully printing a few books to warehouse, is less than the potential losses of the IP rights, probably.
Comment by flir 3 hours ago
But there are worse outcomes.
Comment by nobodyandproud 1 hour ago
Comment by eudamoniac 3 hours ago
There is not a shortage of books in print. I don't understand why people get so hung up on a few of them being out of print. It is okay for someone to own something cool and not let anyone see it. The cool thing doesn't suddenly become a societal necessity because it is words written down.
Comment by jolmg 2 hours ago
Because the point isn't so much in reading whatever text as if all text is the same. The point is the spreading of knowledge. A single book can contain knowledge not present in any other.
> But I don't see how that desire results in the laws needing to be changed so that you can read everything you want
In order for your "people" (country, etc.) to do better, you want them to be educated. In order for people to understand one another, you want them to be able to see all the same various perspectives there are. It makes perfect sense for laws to aim for these goals. This is why libraries exist.
> It is okay for someone to own something cool and not let anyone see it. The cool thing doesn't suddenly become a societal necessity because it is words written down.
Books aren't simply trinkets, like an item you bought at a gift-shop.
> I understand you want to read the books.
I think you're looking at this too much as what people want for their own individual selves, when it's more of what people want for everyone. It's about what they believe is best for society as a whole. They don't need to want to read a book themselves.
Comment by eudamoniac 1 hour ago
Comment by andrepd 1 hour ago
Comment by SwtCyber 3 hours ago
Comment by stellamariesays 4 hours ago
Comment by trollbridge 5 hours ago
We use a special guillotine type cutter to cut off the binding and then store the pages in a sealed plastic bag which goes in the archives; they're stored there indefinitely in case the book needs rescanned for some reason. We also keep the original, uncompressed copies of the books on magnetic disks.
We also go out of our way to try to find rare books published in 1931, 1932, etc. so they are ready to go once the copyright expires.
And no, no AI company has ever come to us and asked to run training on all of our scanned copies.
Comment by palmotea 5 hours ago
Who is we?
> then store the pages in a sealed plastic bag which goes in the archives; they're stored there indefinitely
Is that the best thing for archival storage? Like could things like chemical breakdown increase the humidity in the sealed bag or concentrate corrosive chemical vapors? I was under the impression the best environment was an actively climate-controlled environment.
Comment by ComputerPerson 3 hours ago
Most books have some mold by the time they're ~100 years old, it's just not enough to cause a serious problem. Sealed enclosures (wrappings, bags, tubs) are a nightmare situation. Even packing books too tighly on shelves accelerates mold growth to problematic levels.
Comment by starkparker 3 hours ago
Comment by pfdietz 2 hours ago
Comment by qingcharles 1 hour ago
Comment by palmotea 30 minutes ago
> I would think this totally dries the pages out and then they just turn to dust, from experience.
They used to sell Boveda two-way humidity control packs that would keep a bag at a constant low-ish 32% humidity, but last I looked those were discontinued (it looks like they've pivoted pretty heavily to marijuana storage and higher-humidity products).
Comment by ComputerPerson 48 minutes ago
Archivists recommend standard ambient conditions or a little drier for long term storage. As you've said, too dry and the pages fall apart; permanent damage.
I suppose the fancy silica gels that maintain specific humidities would work in bags.
Comment by vander_elst 4 hours ago
Comment by Aurornis 3 hours ago
The value of most very old books for AI training is very low. You don’t really want your AI training data to start biasing toward outdated writing styles. Most of the valuable knowledge has been covered again in modern texts in more depth and detail.
There is interesting value in old texts and it’s important to have them archived. It’s less valuable for stirring into the giant pot of AI training data, though.
Comment by sosodev 2 hours ago
I have a hard time believing that text valuable to humans would not be valuable to AI.
Comment by dgellow 2 hours ago
Comment by htrp 4 hours ago
Yet
Comment by butlike 5 hours ago
Comment by remus 5 hours ago
Comment by ryukoposting 4 hours ago
As for consumer-grade solutions, look for the Fujitsu SV600.
Comment by phasefactor 5 hours ago
I have done it at home for my books since the mid-00s.
Comment by open-paren 4 hours ago
Comment by flir 3 hours ago
You've trashed the book (and its lifespan) but some books are for using, not keeping.
Comment by JKCalhoun 2 hours ago
I have books but would have 5x as many if I could not capture them digitally. (When I go to move or go through a purge, some of the books I scanned do go to a used book store.)
So I scan in part to keep my physical book-footprint smaller, but also my scans all get cleaned up and uploaded to archive.org. Mainly I scan young-adult science books from the 50's and 60's (since they were so influential and have all but disappeared except on eBay and the like).
Comment by yorwba 5 hours ago
Comment by ck2 5 hours ago
Comment by est31 6 hours ago
Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright?
IDK about the shredding, it's not nice, but it's more a problem with copyright law than AI companies.
Scanning books you own should be legal from a copyright point of view, and not require shredding.
Second, one should think about abandoned property provisions for copyright works published more than 50 years ago and in danger of being forgotten: once challenged, either you as the owner have to prove that the work is preserved for future generations (e.g. in various libraries around the world), or you have to authorize further copies, or you give up copyright on the work.
Comment by ACCount37 5 hours ago
What happens to the pages after? No one needs them anymore, so they get mulched and recycled.
That would be the dominant scanning method even if copyright wasn't a thing. But then again - if copyright wasn't a thing, there would be much less need to scan any physical media.
The reason why OpenAI can't just go on Amazon, buy a "digital edition" of a 2018 book and use that is that it would violate the license in ten ways, and then the DMCA laws that forbid breaking DRM on top of it.
Comment by voakbasda 5 hours ago
In the end, digital publishing just isn’t right and will lead to massive gap in our historical records. They require active curation and cannot be preserved simply by resting on a dusty shelf.
Comment by butlike 5 hours ago
When today's algae evolve enough into tomorrow's sentient creatures, they're really only going to need up to the industrial revolution and should probably stop right before that.
Comment by compass_copium 3 hours ago
Comment by WorldPeas 26 minutes ago
Comment by edoloughlin 4 hours ago
Strictly speaking, no one needs the Sistine Chapel or the Pietà etc. It would be a shame if they were mulched and recycled, though.
Comment by doublerabbit 3 hours ago
ChatGPT know them, i'd count that as digitalised why keep the originals?
Comment by TeMPOraL 3 hours ago
Comment by dragonwriter 5 hours ago
Because there’d be much less content created in any media to capture in the first place.
Comment by eru 5 hours ago
Comment by classified 5 hours ago
Comment by graemep 5 hours ago
An 18th century book would be out of copyright so why would it be illegal to keep the original and scan it?
Comment by sethops1 5 hours ago
It's cheaper to scan the books if you do it destructively. Cost. That's why they're shredding irreplaceable texts. Nothing to do with copyright.
https://www.404media.co/ai-companies-are-buying-tons-of-old-...
Comment by cestith 4 hours ago
Comment by dpark 2 hours ago
Comment by breakyerself 2 hours ago
Comment by dpark 1 hour ago
404 Media published a story about this as well and cites a bookseller who notes that all of the books there have sold have had ISBNs (and are thus from 1967 or later and generally would have active copyright).
”very large purchases were of books that had little in common, except for the fact that they all had ISBNs. This seller also sells rare books that do not have ISBNs, and none of those were part of the bulk purchases”
Comment by cestith 5 minutes ago
Comment by mc32 5 hours ago
Comment by soco 5 hours ago
Comment by qingcharles 1 hour ago
It's not just the big players buying up these archives, either. There are a lot of smaller players, especially in the OCR space, who are buying up huge swathes of works in languages which have much smaller digital footprints, e.g. Arabic.
Comment by 9dev 5 hours ago
Comment by qingcharles 1 hour ago
After that you are diving into forum posts etc to see if you can find anyone who has even mentioned owning a copy or having seen a copy.
I don't know what happens to some works. Supposedly thousands, tens of thousands, or sometimes apparently a million or more copies published and yet not a single copy surfaces for years.
Comment by phoghed 4 hours ago
Comment by shimman 4 hours ago
Comment by ctoth 2 hours ago
1. Why are you here?
2. What is the purpose of this comment?
Comment by eru 5 hours ago
Most rare books are rare because no one cared enough about them. Ie most rare books are rubbish.
Comment by croes 5 hours ago
Comment by jfyi 3 hours ago
Comment by lousken 5 hours ago
Comment by kingstnap 5 hours ago
But during covid archive.org decided to just remove the limit and lend unlimited copies concurrently which started the debacle with the publishers.
Comment by cvadict 5 hours ago
IIRC, this was 100% it. Lending one digital version of one physical asset was likely already a violation copyright. Lending UNLIMITED digital versions of one physical copy was DEFINITELY a blatant violation of copyright.
Comment by phasefactor 5 hours ago
Comment by butlike 5 hours ago
Comment by ndiddy 4 hours ago
"IA maintains that it delivers each Work “only to one already entitled to view [it]”―i.e., the one person who would be entitled to check out the physical copy of each Work. But this characterization confuses IA’s practices with traditional library lending of print books. IA does not perform the traditional functions of a library; it prepares derivatives of Publishers’ Works and delivers those derivatives to its users in full. That Section 108 allows libraries to make a small number of copies for preservation and replacement purposes does not mean that IA can prepare and distribute derivative works en masse and assert that it is simply performing the traditional functions of a library. 17 U.S.C. § 108; see also, e.g., ReDigi, 910 F.3d at 658 (“We are not free to disregard the terms of the statute merely because the entity performing an unauthorized reproduction makes efforts to nullify its consequences by the counterbalancing destruction of the preexisting phonorecords.”)."
Comment by Cthulhu_ 4 hours ago
This wasn't a very smart move of them. I get why they did it but they put themselves at a huge legal risk.
Comment by lousken 2 hours ago
Comment by WorldPeas 25 minutes ago
Comment by Incipient 5 hours ago
Comment by the-grump 5 hours ago
Publishers had accepted the prior arrangement before The Archive decided to push it, if not explicitly then implicitly by not suing.
I'm a believer in The Archive's mission, and I wish they had treated the goodwill they'd accumulated as something worth preserving and not a currency to be spent.
It has been stated by many before me: lending books should have been handled by a separate entity, especially when they removed the physical backing requirement.
Comment by kmeisthax 5 hours ago
Furthermore, in the discovery for the Internet Archive case, publishers had already found a case where IA had lent out books despite knowing their partner libraries wasn't actually withdrawing loaned-out copies from circulation. The CDL premise was always just a suggestion, and IA would have still lost their case if they hadn't done the National Emergency Library (NEL) stunt or if they'd been sued in another venue that hadn't had the ReDigi case as precedent.
It's important to note that whenever a company decides to sue for copyright, it is often late, because the company is banking infringements up to the 3-year statute of limitations and because building a meritorious case takes time. The lack of a timely lawsuit proves almost nothing about the intent of a publisher with a valid case against you.
The thing is, I don't even think the whole stunt damaged much of the IA's goodwill? I know of a few people who withheld donations to IA, but that was mainly under the assumption that publishers would be getting a billion-dollar damage award that would immediately bankrupt IA and result in it's archives being sold off to Lexis-Nexis or something. The funny thing is, IA wound up settling for a sum so small they had to promise never to reveal it, and the danger is gone, so the only thing people complain about now is just that the NEL stunt maybe pushed them "above the radar" or something.
It's still insane that shredding books for AI training is legal, but this isn't.
Comment by TeMPOraL 3 hours ago
AI training happens to be one of the fields exercising that option, but since it's the current favorite topic for people to hate on, here we are.
Comment by azan_ 5 hours ago
Comment by lousken 2 hours ago
Comment by infinite_spin 5 hours ago
Comment by redsocksfan45 5 hours ago
Comment by JumpCrisscross 5 hours ago
How are these things remotely related? If anything, Archive.org’s callous, thoughtless approach nuked the hands of legitimate archival efforts.
Comment by ACCount37 6 hours ago
Turns out that buying an old book for $5 and destructively scanning it for $25 is way cheaper than paying extortion fees to the copyright-mongers.
What I don't buy is it being "rare, precious books". First, they're not after ancient texts - they're after the books that there's still copyright on. Second, when it comes to books, "old" doesn't mean "valuable" - plenty of libraries destroy old books because there's no demand for them, and storage costs you. This is how those scanning companies get books for so cheap.
Comment by tencentshill 5 hours ago
Comment by ACCount37 5 hours ago
You know, piracy online is nice and simple - but it's kind of hard to get physical media without paying what the previous owner considers "a fair amount" to part with it.
Comment by chii 5 hours ago
they should pay the marginal value that the next buyer would buy.
Do you also think that a person dying of thirst ought to pay the maximum price they could possibly pay for water?
Comment by qwytw 1 hour ago
Well the argument is that they should pay for the right to produce derivative works not for the physical copy.
Comment by soco 5 hours ago
Comment by Minor49er 4 hours ago
Comment by qwytw 1 hour ago
I suppose an argument might be made for fair use if they released their model weights publicly without financially profiting from it.
Comment by bitwize 3 hours ago
Comment by glaslong 2 hours ago
Comment by TeMPOraL 2 hours ago
Instead of blindly jumping on a manipulated outrage bandwagon, people would do well to maybe read some of those old books, not even the rare ones - they tend to contain plenty of parables and stories explaining basics of morality and civilized conduct. We used to teach that to kids at homes and in primary education...
Comment by qwytw 1 hour ago
Turning it the other way around is it deeply immoral and sociopathic for someone to hoard massive amounts of money/resources if there are people dying or suffering around them (even if not on the spot but e.g. due to poor access to healthcare)?
Comment by Catloafdev 3 hours ago
Comment by Analemma_ 3 hours ago
Comment by SirFatty 5 hours ago
Comment by infinite_spin 5 hours ago
Comment by jerf 5 hours ago
Moreover, while it is emotionally appealing to some people to want to add some sort of "responsibility to society" to people who own the old books, it's a very emotional plea that can't really be manifested in the real world. As already pointed out in other places, "an old book" itself doesn't really mean much in terms of what its value is in any particular dimension. Plus, I am always deeply suspicious of anything that expands to "Other people, who are not me, should expend vast quantities of resources so that in the next five or ten times I think about this issue for the rest of my life I feel slightly better about this issue" which is what this really amounts to. I think that as superficially appealing as that may be, it's really a very hostile and demanding position to take.
Personally, to the extent that I would want to lay a "social responsibility" on the AI companies, I'd like to see something like they are either obligated, or ideally, just do it of their own free will, to make the scans of the books that are out of copyright available for some reasonable fee (ideally, "free because we like the PR", but given the scope demanding it be free is not reasonable), and without them trying to lay any further claims on the public-domain results. Trading "one old book somewhere, inaccessible to the world" for "a scan of the book and an OCR of it" I would judge a net win for society for rather a lot of these old books, which are by no means "worthless" sitting in some old collection somewhere but would be a lot more useful for being available.
[1]: https://en.wikipedia.org/wiki/Decoupage , since I imagine a number of people won't know what that is.
Comment by eru 5 hours ago
They could give a copy of the data once to some third party organisation that then seeds it in bittorrent or something like that. Basically, what I want to say is that this doesn't need to be an ongoing obligation for the scanner to be worthwhile for society.
Comment by master_crab 5 hours ago
Comment by yeoyeo42 5 hours ago
they're paying for the books, no shady things going on there. whether the publishers should deserve more than a single copy's worth is a separate question.
having the law such that its illegal to scan a book and then keep it, but legal to scan it and destroy it - gg no re there, law people retardmaxxed themselves as they tend to do with anything related to digital data.
Comment by margalabargala 5 hours ago
Comment by TeMPOraL 2 hours ago
Comment by silverlimetea 3 hours ago
They change the info (many buyers), destroy the source material, and now the lie is in the LLM.
That's it!
Comment by butlike 5 hours ago
Comment by pfdietz 3 hours ago
Comment by pessimizer 2 hours ago
More like destructively scanning it for $0.25, if you include the wear and tear on the machine, the salary of the guy who unjams it when it chokes, and recycling fees.
Comment by wwweston 5 hours ago
> paying extortion fees to the copyright-mongers
Yes yes, greedy fat cat publishing oligarchs treading on the poor put-upon scrappy AI underdogs. /s
Back in reality, the fraction of people who got into publishing books to get rich collecting rents is… not large. There’s so many other fields that are likely to reward participants with more wealth that it’s absurd — even with all the passion for the work in tech it’s probably relatively less pure.
And whatever the excesses of copyright have been, the whole bargain has always been on more pro-social foundations and stronger intellectual foundations than “extortion” sneers. It recognizes that incentives matter and work that’s valuable should be rewarded and incentivized.
A culture that takes a Robin Hood approach to low marginal cost billing points but fawns over the hypercapitalized distribution King Johns isn’t creating a freer or richer society or fighting the real cartel center, it’s indulging resentment and caricature.
Comment by pessimizer 2 hours ago
That's because the business is buying copyrights in bulk from those people. Copyright-mongers ≠ publishers or writers.
Comment by sherr 6 hours ago
Comment by vessenes 6 hours ago
That said, supporting Anna's archive is one of the best things you could do for humanity long term in my opinion.
Comment by dwohnitmok 4 hours ago
Comment by awinter-py 1 hour ago
vinge wrote his singularity piece in the 80s I believe
Comment by dwohnitmok 57 minutes ago
> Stan Ulam [28] paraphrased John von Neumann as saying:
>> One conversation centered on the ever accelerating progress of technology and changes in the mode of human life, which gives the appearance of approaching some essential singularity in the history of the race beyond which human affairs, as we know them, could not continue.
> Von Neumann even uses the term singularity, though it appears he is thinking of normal progress, not the creation of superhuman intellect. (For me, the superhumanity is the essence of the Singularity. Without that we would get a glut of technical riches, never properly absorbed (see [25]).)
which hews much closer to how people who are fans of the term use it (the Singularity explicitly refers to the rise of superhumanly intelligent systems, not just the progress of technology overall).
Comment by orthoxerox 5 hours ago
Comment by vessenes 5 hours ago
Comment by 59percentmore 3 hours ago
Comment by clickety_clack 5 hours ago
Comment by trollbridge 5 hours ago
Also, holy cow, hard disks (as in the magnetic oxide kind) got a lot more expensive.
Comment by vessenes 5 hours ago
Comment by cyberrock 4 hours ago
Comment by JumpCrisscross 4 hours ago
Let AI companies do this. But require them to make the digital copies public. Maybe with a multi-year delay, to give the original scanner advantage to doing it.
[1] https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....
Comment by sixdimensional 4 hours ago
Let's say one of the books to be digitized and destroyed is the sole remaining copy of a book from 1850, which is now considered public domain.
On one hand, hoarding such a book, stealing its content from the public domain, locking its content behind a for-profit machine, and destroying the only remaining copy is clearly wrong. It's equivalent to stealing a public resource, just like mining minerals or oil on public lands without a permit or mineral rights. Pure extraction.
On the other hand, taking care to digitize the copy and making it available for free in perpetuity, as well as being required through regulation to provide access to that content through, let's say a public utility LLM/AI available for free through libraries and online... and perhaps after fair due diligence being required to preserve physical copies in a public archive of rare books of which there are no known remaining physical copies...
That seems much more reasonable to me at least. I can imagine there are many who would not see it that way though. Do we see it happening or gaining regulatory, moral and/or public support?
Comment by s1artibartfast 4 hours ago
It provides a freedom to circulate, but not access to the material. It is not a public owned resource.
Turning a copy over to the public or state might be an interesting requirement for obtaining a copyright, but instituting that fix for new works now would have a 70 year lag time.
Think of it this way, if I copyright a book and put it in my dresser for 70 years, that doesn't give the public the right to access it or come into my house and scan it after expiry
Comment by Legend2440 3 hours ago
This has been a requirement in the US since 1790. It is called mandatory deposit: https://www.copyright.gov/help/faq/mandatory_deposit.html
They don't keep every book though.
Comment by JumpCrisscross 2 hours ago
It increasingly looks like a grand compromise around copyright and AI is needed. Expanding public domain when it comes to AI companies in this respect seems merited. (The other bits are fair use if weights are opened.)
Comment by sixdimensional 2 hours ago
A privately owned physical book and the public-domain work embodied in it are different things. Keeping an old book in a dresser, even without providing access to it, is also meaningfully different from deliberately acquiring the sole remaining copy in order to extract its content, destroy the artifact, and preserve exclusive commercial control over the only remaining usable version.
Destroying the sole surviving copy would not literally steal publicly owned property. But it could irreversibly remove a public-domain work from practical human access and destroy a unique piece of cultural heritage.
In that very narrow situation, preservation, archival deposit, digitization, or public-access requirements may be justified, especially when the material is acquired for commercial use that depends on destroying or withholding the only surviving source.
The company would not be reasserting copyright in the legal sense. It would, however, be creating copyright-like control over practical access.
The work would remain legally free for everyone to reproduce, while the company’s conduct made it impossible for anyone else to obtain it.
I know there are more pressing issues than this one in the world, particularly those involving immediate human life.
But I also believe that the value we assign to human life is connected to the value we assign to human knowledge, memory, invention, and culture. When those things are casually treated as disposable inputs for extraction, something important erodes.
Companies should be free to use public-domain material commercially. The issue arises when that use creates a negative externality by permanently destroying the public’s future opportunity to access and use the work.
Mining our cultural heritage and locking away what remains is still extraction and exclusivity.
Comment by Cider9986 4 hours ago
Support your local shadow library: https://annas-archive.pk/donate
Comment by Springtime 5 hours ago
> You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal.
Is there evidence of this? Since otherwise they could very well be describing what is only occurring to in-print or non-rare books. (This is a genuine question since their post doesn't shed any light on it.)
Comment by D13Fd 5 hours ago
Comment by samastur 5 hours ago
Comment by qingcharles 1 hour ago
Comment by RIMR 1 hour ago
If it is happening, it is an outrage. However, the 404 article doesn't actually provide any evidence of this; it just connects the shredding of digitized books and the digitizing of rare books to the assumed shredding of rare books, which isn't necessarily happening.
Comment by 4ndrewl 1 hour ago
Comment by pu_pe 5 hours ago
I am not sure that physically destructing one copy of this type of book to preserve its contents digitally is so bad. Pretty much anything that is still under copyright should be valuable only for its content, not for the physical medium it's printed on.
Comment by qingcharles 28 minutes ago
Comment by pfdietz 3 hours ago
These are increasingly available online, btw. Historical research is accelerated when historians have direct access to scans of relevant source material. Not destructively scanned, of course.
Comment by eru 5 hours ago
Well, there are special collectors editions with the signature of the author and gold pages or whatnot. But the AI companies are probably not using those.
Comment by lejalv 5 hours ago
For whom? is the relevant question
Comment by godshatter 4 hours ago
Comment by storus 4 hours ago
Comment by pessimizer 2 hours ago
Consequently, many/most books are more available now than they've ever been. This is mostly a copyright question, not a question of preservation. I say mostly, because those scans, without redundancy and fingerprinting, can be changed and bowdlerized in the future without remaining physical copies as a reference.
I own about 3500 print books, and started a project to find scans for all of them (that I need to get back to.) Average publication year is probably around 1975, and the bulk ranging from the 1940s to the 2000s. I made it through about 1500, and couldn't find maybe 40, most of them bad. e.g. self-published stuff like "My God Heals, Does Yours?" This was a few years ago, if I went through those 40 now, I bet I'd find half of them.
I hate what the AI companies are getting away with, but only because they get to violate copyright while being aggressive enforcers of copyright and DRM circumvention laws. Destructive scanning, however, is a cheap way to get a good copy of a book online. If that book were then put in a place where teveryone interested in its contents (or who are just hoarders) can get a hold of it, you'll be able to find a verifiable copy of it 1000 years from now.
As of now, Russia and annas-archive are just a few points of failure that can erase those words forever. It will be done with armed, uniformed men, and people will claim it is not dystopia, but justice. I don't want the only copy of the text of a book to lie in the interpretation of some privately trained LLM.
Comment by flipped 5 hours ago
Comment by aerodexis 4 hours ago
Now that I think about it, The Judge is an apt metaphor for AI : "Whatever in creation exists without my knowledge exists without my consent."
Comment by karahime 3 hours ago
Comment by aerodexis 3 hours ago
Comment by karahime 2 hours ago
Comment by TeMPOraL 2 hours ago
Comment by RIMR 1 hour ago
Comment by xboxnolifes 3 hours ago
Comment by mchusma 3 hours ago
Comment by silverlimetea 3 hours ago
Comment by 999900000999 4 hours ago
Leave it up for anyone to download and then compensate the copyright holders later.
In fact if ingesting these books for LLMs is fair use, us commoners should be able to read them for free. Maybe restrict commercial redistribution though.
Comment by streetfighter64 4 hours ago
> They should be forced to publicly release the books as an Ebook.
would be reasonable in any way? The books aren't theirs to release publicly. If I brought a copy of any given movie on DVD, ripped it and used it to train my own "LLM" located at /dev/null, should I then be allowed (or even forced) to release the movie publicly for anyone to watch for free?
Comment by 999900000999 4 hours ago
Set up a compensation fund for the rights holders. Anything is better than culture literally being sucked into the void.
Comment by cj 4 hours ago
Comment by pfdietz 3 hours ago
Comment by opem 21 minutes ago
Comment by thechao 6 hours ago
Comment by Cynddl 5 hours ago
> The attachment contained 3,000 English-language titles organized by ISBN number, including books such as Distinct Element Modelling in Geomechanics by K.R. Saxena (1999); Barrett's Traditional Fairy Tales (2021), an academic study of Irish folklore; and Laser Shock Peening of Advanced Ceramics by Pratik Shukla (2018).
Comment by infinite_spin 5 hours ago
How is a book from 2021 considered rare in this context? There's almost certainly a digital copy of it in existence prior to Anthropic purchasing a print edition.
Comment by ACCount37 5 hours ago
A digital copy would exist somewhere, of course. But for us, that only matters if we can buy or download it. And for AI companies, that only matters if they can get a digital copy DRM-free and licensed permissively enough.
Comment by sidewndr46 5 hours ago
Comment by eru 5 hours ago
Comment by jdub 4 hours ago
(which is terrible, but I would be delighted if they breached it and got thoroughly spanked)
Comment by ACCount37 2 hours ago
That atrocity of a law was a blight upon digital freedom since the day it came to exist. DRM should never have been given any legal protection - and I would push for numerous forms of DRM to be outlawed instead.
Comment by sidewndr46 3 hours ago
Comment by goldlimetea 2 hours ago
Typically DRMs are implemented server-side so you can control user access in that way.
Comment by pessimizer 2 hours ago
Comment by streetfighter64 4 hours ago
Because DRM is just a way to make "breaking copyright" more practically cumbersome. What's easier, breaking digital DRM for each and every E-book you find, or just establishing a single pipeline for scanning physical books?
Comment by sidewndr46 3 hours ago
Comment by TeMPOraL 2 hours ago
More importantly, it's also explicitly illegal. Destructive format-shifting is not. Thank copyright laws.
Comment by tokai 4 hours ago
Comment by timmmmmmay 1 hour ago
Comment by andrepd 4 hours ago
Bearded German man was right.
Comment by skeledrew 35 minutes ago
Comment by JodieBenitez 3 hours ago
It's a routine act in any publisher stock management.
Comment by Gander5739 3 hours ago
Comment by robertoandred 3 hours ago
Comment by SwtCyber 3 hours ago
Comment by 1970-01-01 3 hours ago
Was was the alternative? Order them to stop quietly destroying their property, it is making someone else very upset?
Comment by _m_p 5 hours ago
https://www.ala.org/tools/challengesupport/selectionpolicyto...
Comment by oliwarner 5 hours ago
Weeding is the natural process of disposing of less-demand books. Like the rest of us, libraries operate in finite space, so if they want new books, they have to remove ones their users aren't using. Most libraries will try to sell books before disposing of them in any destructive way.
What similarities do you see here?
Comment by Jweb_Guru 5 hours ago
Comment by pfdietz 3 hours ago
Comment by Jweb_Guru 3 hours ago
Comment by 03284782470 3 hours ago
Comment by streetfighter64 4 hours ago
Comment by emddudley 5 hours ago
Comment by dingaling 3 hours ago
They are most assuredly destroyed.
Comment by malfist 5 hours ago
Comment by dwedge 5 hours ago
Once that AI companies are really shredding 200 year old rare books, and once that libraries are only weeding mass market pulp fiction.
Comment by malfist 4 hours ago
Comment by streetfighter64 4 hours ago
Comment by malfist 4 hours ago
Comment by dwedge 2 hours ago
Also I had a family member work in libraries. They got rid of books for very small prices based on how many people checked them out. Not how easy they were to replace. One example was a 100 year old gold leafed book they sold for £5 to someone who removed every page to sell separately.
Comment by altcognito 5 hours ago
Comment by Ratelman 5 hours ago
Comment by juancn 1 hour ago
It's sad in a romantic kinda way, because of the lost artifact, but the information is what makes the book valuable, not really the medium.
The out of copyright books don't really need to be destroyed anyway for them to be fair use for AI training, and arguable, even if you needed to, you only need one copy per title per company at most.
So it's not a gigantic loss.
Comment by rubylimetea 1 hour ago
They aren’t verbatim uploading the text 1:1. They are creating vector embeddings and training documents from it, changing whatever they want since it’s in private and protected by NDA, and destroying the original source of information so that nobody knows what was originally recorded.
Comment by adamddev1 1 hour ago
Comment by hackernudes 5 hours ago
Comment by justthehuman 3 hours ago
Comment by coffepot77 4 hours ago
Comment by TSiege 3 hours ago
Horrendous stewardship of humanities collective knowledge all for profit and the race to have the one god computer to rule them all.
As more time passes it becomes clearer that America’s AI strategy should’ve been a public private partnership where the public owned the datasets and the underlying models and we’d leave the productionizing of LLMs to private businesses
Comment by justthehuman 3 hours ago
Also, they can then just recycle/dispose of the paper and don't have to worry about reselling/donating the books themselves. I suspect this is all about speed of data ingestion and anything else is a side effect they don't care about.
Comment by dclaw 1 hour ago
Comment by rubylimetea 1 hour ago
They’re privately putting info into their LLMs, changing whatever they want, sorry, “sanitizing” then destroying the original.
Comment by nh23423fefe 1 hour ago
Comment by D13Fd 5 hours ago
The fix here is to change the law to permit training AI without destroying the original materials. But that is going to be a heavy lift.
Comment by bronlund 3 hours ago
There is a word for that kind of behavior.
Comment by DarkIye 5 hours ago
Comment by rubylimetea 1 hour ago
They are creating embedding vectors and training documents - changing whatever they want - and destroying the original copy so nobody knows what was actually said.
It is just like book burning.
Comment by dpark 5 hours ago
You make it sound like they are running a second Project Gutenberg. They most definitely are not making these available for electronic searches. At least not searches the public can participate in.
Comment by coffeefirst 5 hours ago
Comment by imhoguy 5 hours ago
Comment by rubylimetea 1 hour ago
I would say “reduced”
Comment by swed420 5 hours ago
What guarantee do we have that the book contents will be served unfiltered and unaltered?
Comment by left-struck 5 hours ago
Comment by writtenone 4 hours ago
Comment by Cider9986 4 hours ago
Comment by musha68k 5 hours ago
This is highly disturbing news; is this standard practice? What did Google Books do before?
Comment by qsera 5 hours ago
Comment by regnull 4 hours ago
Comment by bryan_w 4 hours ago
Comment by simonw 5 hours ago
The segment that talks about rare books:
> One professional bookseller who specializes in selling foreign language books on these marketplaces told me that, starting in April, he and other booksellers noticed a historic spike in sales. [...]
> This bookseller said his inventory is full of rare, foreign language, and low circulation books, meaning that if they are destroyed in the process of becoming training data, they’ll be even harder to obtain.
Comment by skybrian 5 hours ago
> Bulk purchases also usually reflect interest in a specific topic, whereas the recent, very large purchases were of books that had little in common, except for the fact that they all had ISBNs. This seller also sells rare books that do not have ISBNs, and none of those were part of the bulk purchases.
Article is paywalled, but I saved a few quotes here:
Comment by dpedu 3 hours ago
Comment by xgulfie 3 hours ago
Comment by lagrange77 3 hours ago
Say i feed the largest LLM a book of an alien civilisation, that it definitely hasn't seen before. Then this tiny piece of text muds the vast ocean (latent space) of the model minimally. It will not be able to cite from that book reliably after that fine tuning. Especially for rare books, because them being rare implies, that there aren't 1000s of other books, that encode the same information.
It's general language modelling capabilities might get an iota better, of course. But for putting factual information into it, wouldn't RAG be a much more solid approach?
Comment by 59percentmore 3 hours ago
It's a real shame that no one ever got that book in front of Hayao Miyazaki's eyes.
Comment by greenlimetea 2 hours ago
I looked it up.
surveillance, sousveillance - sure, why not? Seems harmless.
But it's as if the more people know the word sousveillance, the word and even the act of surveillance actually loses a little bit of power (its embedding changes, if you will!)
All our fears about the end state of surveillance can now be countered by an end state of sousveillance.
Before I had sousveillance to think of, I could only think of surveillance (when thinking of veillances) - and it was more of a threat then than it is now.
The world is made of language, or as Terence McKenna said made out of words which sounds obviously false at first. But go looking for the inside of an atom and tell me what you find, and think about where the medium of reality actually implements itself.
Comment by mwigdahl 2 hours ago
Comment by geephroh 2 hours ago
Comment by graemep 5 hours ago
Comment by roywiggins 5 hours ago
Comment by cormacrelf 1 hour ago
https://www.404media.co/ai-companies-are-buying-tons-of-old-...
Comment by m4rtink 2 hours ago
Comment by azan_ 5 hours ago
> It is equivalent to book burning in the past. A form of thought control
Comment by nhinck3 4 hours ago
Comment by frozenseven 1 hour ago
Yeah, this isn't a thing. Throwing out books, even "rare" ones, is a regular procedure. And if you really want to read 'em, you can buy any of these books right now.
Comment by swed420 5 hours ago
> > It is equivalent to book burning in the past. A form of thought control
That's only bullshit if you trust AI companies to serve the book contents without alteration.
Comment by dpark 5 hours ago
Comment by swed420 5 hours ago
The point is that even under the best intentions, hallucinations occur. Then there's the fact that most models have an ideological bias programmed into them.
The only expectation I have is for companies or anybody else to not destroy rare books. Is that such a tall order?
Comment by dpark 4 hours ago
Sure, in the same sense that they regurgitate any other text they consume. LLMs by definition do not have the full training dataset available, though. It’s far larger than the resulting model. So they can’t reliably reproduce full text without an external source (or if it’s in the training data repeatedly). ChatGPT actually refused to give me a bible quote the other day, presumably because I ran into some general “book regurgitation” safety net.
> The only expectation I have is for companies or anybody else to not destroy rare books. Is that such a tall order?
Honestly, yeah. The idea that people or corporations should hold onto books forever because of a cultural “ick” about throwing out books is a bit ridiculous. Most books end up in landfills.
They aren’t feeding Da Vinci manuscripts into this pipeline. They are feeding still-in-copyright books.
Comment by pfdietz 2 hours ago
Is it? How many different books are we talking about, and how much information is that, after conversion to text and lossless compression? Images, maybe, but text?
Comment by dpark 1 hour ago
Comment by pfdietz 1 hour ago
Comment by dpark 38 minutes ago
Comment by pfdietz 14 minutes ago
Comment by swed420 4 hours ago
That makes it even worse, then. This proves the original point.
> The idea that people or corporations should hold onto books forever because of a cultural “ick” about throwing out books is a bit ridiculous.
If we're building black and white straw man arguments, then sure, let's not archive anything.
Comment by dpark 3 hours ago
I don’t know what the “original point” is here, but these AI companies are not providing “book excerpt services” and do not claim to. ChatGPT at least will refuse to provide detailed book excerpts (I hit a week or two ago myself).
> If we're building black and white straw man arguments, then sure, let's not archive anything.
It seems like you are the one creating the straw man. Do you have evidence that these companies are shredding actually rare books? The only cited concrete examples (in this thread anyway) are all rather boring. I seriously doubt they are shredding 200 year old books because why would they?
Comment by bwfan123 3 hours ago
Comment by frozenseven 1 hour ago
Comment by pfdietz 3 hours ago
Isn't that oxymoronic? If they can be bulk-bought they aren't rare.
Comment by computerphage 3 hours ago
Comment by pfdietz 58 minutes ago
Comment by qsera 5 hours ago
I wish...
Comment by johnxianren 5 hours ago
Comment by Arshad-Talpur 4 hours ago
Comment by tasuki 4 hours ago
Comment by pyrophane 4 hours ago
Comment by extra-AI 4 hours ago
Why would I spend hours creating original content if Google can extract it and present the answer directly in an AI Overview? What is the incentive to keep doing the work?
If creators stop producing high-quality original material, the information we get over the next few years will increasingly be based on recycled, low-quality garbage.
Comment by Traster 4 hours ago
Can Google steal it and present it in an AI overview? Well kinda. Today Google is doing a trick - they're saying "You can refuse to consent to being fed into the slop machine, but if you do we won't crawl you for Google so you'll get no search traffic. But you're not going to get search traffic anyway! So you might as well opt out of being fed into the slop machine. And companies are starting to do that [1]
It's really interesting, because essentially what it means is Google is turning into a walled garden, but there's nothing growing inside it so they have to continually import new plants to live in their walled garden and they're going to have to pay to do that. So soon Google will be paying news sites for the right to plumb their feed into the slop machine.
[1]: https://www.wsj.com/business/media/google-search-publishers-...
Comment by secretsatan 4 hours ago
Comment by Gander5739 3 hours ago
A spurious claim; they simply want to avoid model collapse.
Comment by secretsatan 3 hours ago
Comment by Cider9986 4 hours ago
I find going after shadow libraries to be much worse because law enforcement is trying to prevent discrimination of knowledge to the public.
The real blame here should be going onto copyright laws.
Anthropic could take more care by figuring out if the books are still affected by copyright.
But this is just a company trying its best in an unfortunate regulatory environment.
Support your local shadow library:
Comment by HelloUsername 4 hours ago
Comment by gowld 3 hours ago
Comment by HelloUsername 1 hour ago
Comment by storus 4 hours ago
Comment by paxys 5 hours ago
Weird to see so many of these "trust me bro" twitter stories make it to the front page and cause outrage when no one has any real information.
Comment by skeledrew 46 minutes ago
Comment by iamsaitam 4 hours ago
Comment by hagen8 4 hours ago
Comment by goldlimetea 3 hours ago
And the only real source for anything will be an LLM response.
Comment by bigbuppo 2 hours ago
Just one more book, bro, and we'll solve AGI forever. Trust us, bro, just one more book. Come on, let me have those words and we'll solve AGI forever.
Comment by mindslight 4 hours ago
This is merely the latest incarnation. We can imagine a slightly different process on a few fronts - AI companies pay to digitize books (still for their own purposes), but are prevented from destroying the physical copies and they have to openly shared the digitized results. We would view that situation much more favorably - perhaps even as ideal, right?
Those two dynamics could be backed up by court decisions or laws iff they weren't so plainly at odds with how copyright has been and is generally implemented and interpreted. For example, imagine them having to do this through some nonprofit library whose goals was preservation and dissemination. Instead, libraries have been sidelined as things that operate at the edge of the law rather than vital public institutions, whereas shredding books in secret is fully legally condoned.
Comment by Cider9986 4 hours ago
Copyright has done more than anything else to prevent preservation and dissemination of knowledge. And it's forcing Anthropic's hand now. Although they could take more effort to preserve the books.
Comment by mindslight 3 hours ago
What I'm indicting is the copyright regime being primarily focused on control and the prevention of dissemination. We can imagine a different world in which the publishers' suit against the Internet Archive went the other way (or was not even brought), and a public interest group sues Anthropic (et al) for destroying cultural commons, and gets a judgement saying all scanning must be done non-destructively and made available through institutions like the Internet Archive.
Comment by tokai 4 hours ago
Comment by dandelioness 4 hours ago
Comment by tannerr_dev 4 hours ago
Comment by petesergeant 5 hours ago
Comment by freejazz 5 hours ago
Comment by gowld 3 hours ago
Comment by nailer 4 hours ago
Comment by ertucetin 3 hours ago
Comment by enaaem 5 hours ago
Comment by RIMR 1 hour ago
It's a weird practice to destroy a book when you digitize it, but it's at least an understandable legal strategy to ensure that the digital copy "replaces" the physical one. However, this is only going to apply to books that have active copyrights.
This author suggests that a rare 18th-century botanical text could fall victim to the same fate, but I am somehow doubtful that this is the case. Non-destructive scanning is trivial, and these kinds of books are likely being processed in a quantity that would allow for it without backing up the pipeline.
I really like 404 media, but it doesn't really seem like the evidence points to the conclusion here. Yes, AI companies are shredding books that they digitize, and yes, AI companies are digitizing old, rare books. But the rationale for the book shredding doesn't exist for the old, rare books, so I would need more evidence than just "putting two-and-two together".
At the end of the day, old books with no resale value, including rare books, end up destroyed with some regularity by libraries and bookstores. While this may be an excessively generous take, at least this way the books are getting digitized before they become pulp. The real tragedy will be if the old, rare books that were digitized are never shared with the rest of us, because they were ONLY digitized to train AI, and not to actually preserve anything.
Comment by pmkary 3 hours ago
Comment by ChrisArchitect 4 hours ago
Comment by 1234letshaveatw 5 hours ago
Comment by pfdietz 4 hours ago
Books are information delivery vehicles. We mostly shouldn't care about them any more than we care about a particular set of bits on a disk.
Publishers also pulp large numbers of books themselves. This is a consequence of the Supreme Court's Thor Power Tools ruling, which clarified tax rules in the US so that keeping large inventories of unsold books was less economical.
Comment by Good4boothee 5 hours ago
Really? That sure wasn't a thing when one startup got sued for streaming from its wall od dvd-players, and they adhered to 1 disc = maximum 1 stream at same time.
Comment by fmaccomber 5 hours ago
Comment by retinaros 5 hours ago
Comment by azan_ 5 hours ago
Comment by qsera 5 hours ago
Comment by azan_ 4 hours ago
Comment by qsera 4 hours ago
Comment by azan_ 3 hours ago
Comment by qsera 3 hours ago
* their values are aligned with the best interests of humanity
* their values are not aligned with the best interests of humanity.
Now go ahead and read my mind!
Comment by pfdietz 3 hours ago
But that's a lie. They're not destroying the last copy of books.
Comment by 03284782470 3 hours ago
Comment by qsera 3 hours ago
Oh, I have a reputation here. Thanks for letting me know.
Comment by api 5 hours ago
But that would help competitors with training data, which I assume is why they don’t do this.
Comment by impsunrise 4 hours ago
Comment by qsera 5 hours ago
Comment by psychoslave 3 hours ago
Of course even disregarding this fact, this is utterly disgusting attack on humanity heritage, just as much as any group out there destroying what we should all cherish be it for the historical artifact they represent. Whatever how US judge name it, they don’t worth more than their same-behavior consorts that is terrorists and totalitarian governments.
https://en.wikipedia.org/wiki/Palimpsest
https://organiser.org/2025/07/24/304285/world/china-wages-wa...
https://www.historyexpose.com/things/demolition-afghanistans...
https://link.springer.com/chapter/10.1007/978-3-031-96432-9_...
Comment by emsign 3 hours ago
Luckily I don't live in NYC where book hoarders are being evicted because of "fire hazard". Just like in Fahrenheit 451.
Comment by tudou527 2 hours ago
Comment by SwtCyber 3 hours ago
Comment by z0ltan 3 hours ago
Comment by flipped 5 hours ago
Comment by fibobrain 5 hours ago
Comment by bebe8393jrir 5 hours ago
Comment by cluckindan 4 hours ago
Comment by vessenes 5 hours ago
If we use 'shredding' to mean A book is laid flat, its cover is removed, and then a paper cutter cuts through the binding to create a flat stack of sheets, which are then fed to a sheet feeder, then I could maybe imagine this is better than a page-turning scanner. But, sheet feeding old paper sucks shit, bro. It's not fun.
Upshot, I think we'd like to hear from an anonymous frontier lab employee here to see what's going on -- there are a lot of books in Anna's archive available at considerably less difficulty.
Comment by Jolter 5 hours ago
Comment by vessenes 5 hours ago
Comment by Jolter 3 hours ago
Comment by timcobb 5 hours ago
Comment by iandanforth 4 hours ago
Comment by pfdietz 4 hours ago
Comment by BoxOfRain 4 hours ago
Comment by pfdietz 4 hours ago
You know that pleasant used bookstore smell? It's paper slowly decomposing.
Comment by BoxOfRain 3 hours ago
Comment by pfdietz 2 hours ago
Modern acid-free paper might last 1000 years; 500 is more typical. Acid paper breaks down in less than a century.
Parchment was so expensive it was often scraped and reused; old texts can sometimes be recovered after being overwritten (palimpsests).
Comment by krunck 4 hours ago
These companies are regressive book burners.
Comment by tekne 4 hours ago
Comment by 03284782470 3 hours ago
Comment by genxy 3 hours ago
Maybe the Library of Alexandria didn't burn, it was digested.
The market has solved the what do with excess knowledge problem.
Comment by greenlimetea 3 hours ago
Like, Library of Alexandria or Council of Nicaea bad.
We may never be able to recover the information if, say, one of these AI companies copied or translated it wrong then destroyed the source material.
Maybe it's from bad OCR, or maybe from a bad actor - but there are a lot of ways history and information could change in this game-of-telephone like transfer of knowledge.
What is the point of destroying the source material? I don't buy the copyright thing.
Comment by TeMPOraL 2 hours ago
It is the copyright thing.
Despite what people say about scanning, the fact is, non-destructive scanning machines have been built and perfected long time ago. This was preferred in the past, back before some major kerfuffle with the publishers during COVID, but that incidentally happened to be before LLMs became a thing, so AI companies never had that option available.
Comment by rubyfruit 2 hours ago
We'd be batteries if you were the spokesman of The People xD
Even a 10 year old knows they are simply changing information and destroying the source, so that the lie is now in the LLM and you can't prove otherwise.
They're making the LLM "the source" and destroying the original.
Only a total midbrain would defend it and be unable to see the obvious.
Comment by 1970-01-01 3 hours ago
It is the cheapest way to get them scanned.
It is the fastest way to get them scanned.
It doesn't need to be safely archived for another century until it is resold to someone that has not yet been born.
Comment by rubyfruit 2 hours ago
Is that possible at all or nah?
Comment by stuartjohnson12 6 hours ago
Indeed, I think there's a high chance that this process increases preservation of the most relevant part of the media - the actual content!
There are tons of old rolls of film slowly rotting away in warehouses that were never digitised. Even for beloved media, the BBC occasionally tracks down an old lost episode of Dr Who.
For now these books are in corpuses of training data, but eventually I trust they will make their way to the rest of us.
Comment by Cynddl 5 hours ago
What makes you think they will? What would be the incentives for these companies to do so?
Comment by stuartjohnson12 5 hours ago
1. At some level of critical information-withholding mass, a leak or disclosure similar to SciHub is inevitable because of the commonly held opposition to hiding knowledge.
2. Availability via Google Books or similar.
3. Availability via AI model reference.
4. Failing any of the above, better AI models that are more capable of doing more things, at the expense of books that were likely to go unread (revealed preference, rare books are often rare for a reason). This will obviously be a nonstarter if you don't want this to happen, but I think it would be good for the world if it did.
I think category of old books that were going to be read or otherwise become important parts of human knowledge that have not yet been digitised and now will never become so because they are instead being shredded and will never make their way into the light because of AI company data hoarding is a small category.
Comment by b3lvedere 5 hours ago
Comment by breezybottom 5 hours ago
Comment by gowld 3 hours ago
Comment by breezybottom 1 hour ago
Comment by croes 5 hours ago
Comment by stuartjohnson12 5 hours ago
I am however OK with destroying one of the remaining 50 children's books of which only 300 copies were ever printed in a small town in Ohio in the 1970s as a test run for a failed book which was subsequently never commercialised.
Comment by croes 5 hours ago
What if the perception of those books change over time and are considered masterpieces later on?
Moby Dick was out of print when Melville died 1891 and not a huge success until it got a revival in the 1920s
Comment by stuartjohnson12 4 hours ago
This is the kind of rationalization hoarders use. The inability to get rid of things because it could turn out to be something we want in the future for reasons that we cannot currently describe.
It's a loss avertive instinct that I think is misplaced. Treating every printed book as priceless is intractable. It's not how we treat these books at the moment. Apparently, today we don't even care enough to spend a few hours per book nondestructively scanning them in.
Let's say we, instead of destructively scanning this books, nondestructively scanned them. What would you propose doing with the copies afterwards? Sell them? To who? They're valueless individually for the overwhelming part. Warehouse them? Why? For who? Do you want to go and look through them? Why haven't you done so already? Have you ever shown interest in consuming an undigitised book a single time in your life to date?
It feels like the anti-AI crowd here have to tie themselves in knots here to make the loss minimization work.
Comment by infinite_spin 5 hours ago
Comment by croes 5 hours ago
That’s the whole point because the book shredding is already declared legal.
Comment by infinite_spin 4 hours ago
Comment by Invictus0 5 hours ago
Comment by infinite_spin 5 hours ago