Rendered at 01:53:52 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
squidbeak 12 hours ago [-]
I've limited sympathy for the publishers.
It pisses me off to reflect that they can sit on works until copyright expires, keeping them out of print. There's no real need for any of these so-called rare books to be rare while they're under copyright.
And related to this, the books that are in print are mostly only in print in the shittiest way. I often see well-made books from the 17th or 18th centuries which are still in good nick. It's ridiculous that in the 21st century, publication standards have fallen to the point where for most works a disposable format is the only type available - where no amount of money could buy a truly decent hardback copy.
If we have to have copyright laws, I'd like to see two changes to them.
When a publisher has no incentive to keep an edition in print, it should be available to any other publisher to print, without compensation to the original publisher, and with renegotiated royalties for the author.
And if the publisher keeps a book in print - but only in bestseller-grade materials, bogroll paper that furrows in any humidity and perfect binding that molts its pages a couple of dry seasons later - and if it refuses to print a durable hardback copy with signatures, good paper and decent print - something that will still be readable in several generations' time - any other publisher keen to have a crack at it should be able to, again without any compensation for the original publisher, though perhaps in this case, with matching royalties for the author.
giancarlostoro 11 hours ago [-]
A friend proposed that copyright should just die with the author and / or their spouse and I'm left agreeing. I want books and music to be less strict on copyright. Some of my favorite YouTube channels break down music and songs, and go as far as recreating beats / tracks from famous hip hop songs, but someone at a record label company dings every one of their videos, they can barely sample a few seconds, its VERY CLEARLY fair use, and even gets me to listen to the songs more, but they are stealing all revenue for a fair use video that took a lot of time and money to make to begin with, they are profiting from work they didn't even do. This is silly to me, I'm sure there's loads of other channels and videos out there screwed over by the record industry over 15 second samples from a song... Which should 100% be fair use and not be given to the record labels.
I thought about this and what they are doing is writing the music industry out of the minds of the next generation who mostly uses Youtube etc. for entertainment.
Most Youtube videos do not contain any copyrighted music as it takes too much revenue as you state.
Again, its short term profit at the cost of long-term gain.
monknomo 10 hours ago [-]
20 years to make some money, and then we set the work free for the public benefit. If it's good enough for patents, I don't see why it isn't good enough for copyrights.
reorder9695 10 hours ago [-]
It also solves the issue of potentially not knowing when the author died, with a fixed period (the number isn't so important imo so long as it's sane), a work is out of copyright x years after the first known copy was published.
monknomo 9 hours ago [-]
reduces the ownership hunt problem, as well!
marcosdumay 10 hours ago [-]
Either that or some exponentially increasing tax so that Disney can keep their vault. (I'm perfectly fine with them keeping it if they pay some proper taxes.)
Ikatza 8 hours ago [-]
They pay taxes every time that they make money off Mickey Mouse, regardless of copyright status.
couchdive 4 hours ago [-]
Do they? scope out their federal tax liability in 2025.
Yes, and they should pay taxes whether they use it or not if they want to keep it out of the public domain.
monknomo 9 hours ago [-]
the important thing is that copyright serves the public good, I do not think it currently does, at least not as well as it could
Ekaros 7 hours ago [-]
I agree 20 or 25 years should be plenty of time to protect artistic work. No other industry or line of work has anything like that protection. A work published today by 10 year old could stay in protection for nearly 200 years. If that person was to live to 140, theoretically possible with medicine in future.
We do not continue to pay for most things once they are created. Unless they are continuous services. Artistic works should not be any different.
qwytw 9 hours ago [-]
For corporations sure. For individual authors that's certainly not fair. Especially since it makes it easier for corporations to exploit their work without paying them anything.
> If it's good enough for patents, I don't see why it isn't good enough for copyrights
Because there are fundamentally differ concepts and serve different purposes?
dspillett 2 hours ago [-]
> For individual authors that's certainly not fair.
How about lifetime of the author? Or "lifetime or 25 years whichever is longer" so those writing in their later years (or dying young) can pass on the time they didn't get chance to fully use.
> Especially since it makes it easier for corporations to exploit their work without paying them anything.
That ship has sailed. We've seems "big corp" commit mass piracy and get the lightest slap on the wrist, I doubt they'll get less brazen going forwards.
monknomo 7 hours ago [-]
I don't think they are very different in concept, other than one covers physical goods (and also procedures to make physical goods), and the other covers writings.
The purposes of both are: "To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries."
Like, this is a made up regime with a specific intent. The fact that we treat copyrights and patents differently is an accident of history. I think we could quite reasonably choose a different period of time (and in fact, have done so several times over the past few hundred years) and still promote progress.
I think it's very reasonable to say that one good idea should not be enough to let you coast your whole life, you should be prodded to cough up 3 good ideas. Further, it reduces corporate power at the other end by allowing individuals to play in coroporate properties after a relatively short time. You could be futzing around with, idk, a copyright free Cars under my proposed regime.
qwytw 5 hours ago [-]
It would also make it much easier for corporations to exploit authors.
They would generally make most of their money early (in the couple of years following when the content is released). Individual authors would be much more affected, it might take years for your book to become popular. Also imagine if a studio decides to make a movie or tv show just right after your copyright expires, they wouldn't pay the author anything and just have higher profit margins.
> The fact that we treat copyrights and patents differently is an accident of history
Patented inventions and technologies have some sort of direct practical value. Society does not really benefit much if anyone is allowed to created derived works based on any copyrighted content without compensating the author.
> a copyright free Cars under my proposed regime.
I don't think cars are copyrighted unless you want to make an exact copy of it you shouldn't run into any issues.
monknomo 5 hours ago [-]
If it is true that society does not benefit from creating derived works, then why is open source so useful?
altmanaltman 9 hours ago [-]
But this must then follow for all forms of works, why just books? Every company whose original creator dies must be converted into a public company within 20 years of founder's death for the public benefit as well if that's the case.
If I wrote it, I own the copyright on it, why should I or my future family give away something I worked really hard for? Why do only authors must care about public benefits?
tancop 6 hours ago [-]
> Every company whose original creator dies must be converted into a public company within 20 years of founder's death
if public company means employee owned (instead of publicly traded or state owned) then im all for it. if the founders family are good managers they can easily convince their workers to let them keep running things.
you dont deserve a job or a fortune just because your parents did a lot of hard work before you were an adult. you got to prove yourself and be better than the rest, thats what capitalism is all about right?
altmanaltman 5 hours ago [-]
What do you mean that's what capitalism is about? It's about capital. It doesn't asks where you get that capital - if you have $100,000 that you inherited vrs $100,000 you earned, JPM will still see it as just $100,000. So yes, by capitalism, if your parents did a lot of hard work and saved the capital and gave it to you, you have it. Where does it make a moral judgement that it's wrong to have inheritance or right to earn it by the bootstraps?
And no, I mean every trade secret of the company should become public and anyone should be able to create its products and brands. Basically, do to them what you propose to do to writers and artists - why must their families benefit from their work?
constantius 3 hours ago [-]
To your second point: a company (or any kind of organisation) is continuous work. It's a very obvious retort to your analogy. It would be akin to the author writing a book in a series every year for twenty years, then giving their rights to their child who keeps writing books every year, and every book only has 20 years of copyright. Seems fair to me.
To your first point: inheritance has a concentrating effect on wealth. Concentration of wealth is not a good thing, as it breeds inequality by definition.
If your grandparents gave you a house for free, and you used that enormous advantage to pur your money elsewhere and then own 2, 3 houses, and your child then gets 10, that's 10 houses that people could own themselves, instead of paying the highest rent that your child can get away with.
If you have any experience of poverty, you should understand how viscerally unfair the existence of 'rich kids' seem, and how damaging it is to society.
monknomo 6 hours ago [-]
So what? Why should the state grant you a monopoly at all?
The reason is to encourage people to create useful writings and make useful discoveries. But we need to balance this encouragement with the benefit we get by making these writings and discoveries available to everyone. By granting a time-bound monopoly the state rewards creators, and by ensuring this time period is not excessive, spreads the benefit among the populace.
Plus, with a time bound benefit, you have to keep on creating, which is good for everyone under this regime.
These aren't, like, inherent rights, they are contingent
altmanaltman 5 hours ago [-]
The point isn't why the state should grant me (or writers) a monopoly. But rather, why should the principle not be applied to all property and ownership and not just copyright for creative works? Why should one lose ownership over their said work if they are a writer but they shouldn't lose ownership of their trade secrets because they are a high-frequency trader? Or more so, why should only the creative field worry about the public better good when it comes to ownership and property? The conversation for public domain must extend to everything and not just works or art if one really believes in it.
lsaferite 3 hours ago [-]
Funny that you are comparing Copyrights and Trade Secrets. The monopoly granted on Copyright is predicated on it being published. Trade Secrets, by definition, are not published. Copyright is granted based on the value to the public provided by the author publishing their works. Trade Secrets have no such protection as they are not meant for the public good. They have other protection from theft, but once they are published, there is no protection of the information.
thdr 10 hours ago [-]
> copyright should just die with the author
That would have a few undesirable consequences... for example, you wouldn't hire a 70 year old writer for your commercial project no matter how brilliant they are.
The complexity of our legal system is in many cases justified. The problems are often the numbers (duration of copyright protection etc.)
fn-mote 9 hours ago [-]
> copyright should just die with the author
This makes sense when you’re thinking of a painting or a book.
Who owns the copyright to Windows or MacOS? A corporation. How do you deal with that?
> you wouldn't hire a 70 year old writer for your commercial project no matter how brilliant they are
Commercial projects are works-for-hire and the copyright is not owned by the person who does the work.
The proposal for the limit of copyright needs to be refined.
In the US, the current rule is:
> For … a work made for hire, the copyright endures for a term of 95 years from the year of its first publication or a term of 120 years from the year of its creation
this_was_posted 10 hours ago [-]
Then just make it the rule that the copyright expires after 30 years or when the author dies, whichever comes last.
delecti 9 hours ago [-]
Why not just 30 years? Patents get a flat 20.
Not that I'm arguing for 30 per se, just that I don't see what goals of copyright would be advanced more by adding an "or until death" complication.
qwytw 9 hours ago [-]
Well imagine you write a book in your 20s or 30s and it only becomes popular after a couple of decades. The publisher gets to pocket all the money.
Somebody decides to make a movie based on your book? You get nothing at all from it...
The movie bit would be problematic even for books that were reasonably popular at the time. e.g the Witcher adaption came out almost exactly 20 years after the last book, for GOT it wasn't that far from being the case as well (at least for the initial volumes). Studios would be incentivized just to wait a couple of years to avoid paying anything.
I think it could be reasonably to have a fixed limit if the rights are held by corporations, though.
constantius 3 hours ago [-]
This is an edge case. The vast majority of intellectual output loses value extremely fast. If a new Kafka comes along and their writing becomes popular 30 years after publication, they will have 0 issues getting a fat contract for a new book.
Justifications for copyright are always built on edge cases, seemingly moral justifications of an empirically immoral practice. Yes, it'd be nice if a single handicapped mother of 2 coild see her children rise out of poverty thanks to her writing talent after 50 years. In practice, this person doesn't exist and building society around that scenario is not a good thing.
These 'what ifs' have the same value and do the same damage as 'who will think of the children' do for human rights.
DrScientist 7 hours ago [-]
For inventions if you don't make money off it in the first 20 years you are unlikely to ever make any money off it - as the invention space moves on.
That's not the same for a work of fiction or a piece of music.
Case in point apparently books sales for the Odyssey are massively up - when it was originally written in 7-8 BC :-)
Also most books etc don't make much, if any money - an publisher/author might rely on a the cummulative effect of a number of revenue streams built over time.
Also the effect of exclusivity is different - for patents you are potentially blocking the area of innovation you have patented by your exclusivity.
That's not the same societal effect as somebody not being able to copy mickey mouse.
So they aren't exactly the same - however I'm not proposing a 3000 year copyright :-)
delecti 6 hours ago [-]
My personal preference is actually a flat 50 years. I think that if an author writes something at 25 and it doesn't blow up until they're 75, I think that's given them more than a fair chance to capitalize on it.
My real point though is that IMO, whatever duration we pick shouldn't depend on the creator. It would tend to undervalue their later creations, treats corporations differently from people in a way that doesn't seem relevant to copyright, and oddly might lead to the untimely demise of creators.
I find your Odyssey example to be relevant. Homer's death means people today can release their own translations or adaptations. I can find a public domain version from 100+ years ago, or a modern translator can profit from their work so that I can see their take. I can watch the Italian 1911 silent film version for free on youtube [0], or pay for Nolan's modern take. The expiry of copyright gives me options.
I am fine that estate gets to keep rights for whatever period is left. If estate is dissolved ofc rights would also end. So they can't be orphaned. Either someone has them or they are public domain.
wiether 10 hours ago [-]
> for example, you wouldn't hire a 70 year old writer for your commercial project no matter how brilliant they are
If you hire them, then you own the work you paid them to do, no?
graeme 10 hours ago [-]
Not in every country, and secondly if you're basing it on life of the author then that does't solve corporate copyright unless you tie it to the live of a particular employee.
You could do "life of author or X years, whichever is longer". Or include a period after death.
But you see how the complexities come in.
clort 10 hours ago [-]
The complexities are the problem. I've always thought a fixed term is best. Then, you can purchase a work - it says Copyright <dddd> on it. Then, you know that after dddd+term the copyright is lapsed. You don't need to hunt down the author to see if they died. No guessing, just written on the work that you purchased. No, don't have optional extensions - that just means you have to look it up. It should say it right there on the work you purchased when the copyright expires.
It is said that the vast vast majority of works don't earn anything significant after a few years in any case, meaning the only possible reason to have long copyrights is so that a very few people can get stinking rich. But those people already got rich, in the first few years.. society does not benefit from them getting richer.
20 years fixed term is my proposal.
TeMPOraL 10 hours ago [-]
Either that or if it didn't work, insurance industry would've created a product that makes it work for commissioned works. Nothing happens in isolation here.
Always42 10 hours ago [-]
This would create incentive to kill people
ktm5j 10 hours ago [-]
So does life insurance.. but we still have that.
fragmede 2 hours ago [-]
What we don't have anymore though because of that, is tontines.
dspillett 2 hours ago [-]
They are still around.
weberer 10 hours ago [-]
I would change it to "the expected natural life of the author" which is their birthday + 85 years. Regardless of whether they actually are still alive past that date.
AussieWog93 10 hours ago [-]
So if an 86 year old writes a book they can't copyright it?
duzer65657 9 hours ago [-]
This can't work: life expectancy (of humans) can differ more than 30 years around the world, plus there are significant differences between men and women, and corporations - which often own copyright on work for hire - can live in perpetuity.
giancarlostoro 10 hours ago [-]
I would accept that as well, but crank it down to like 62 years since their date of birth, presumably by the time they would retire.
fn-mote 9 hours ago [-]
What?
Think this through some more.
Authors and artists are still creating after that age.
inferniac 10 hours ago [-]
not really, when copyright disappears, so does the profit motive
sure, you don't have to pay royalties, but any other publisher can now publish too
cindyllm 10 hours ago [-]
[dead]
Ikatza 8 hours ago [-]
What if the copyright is owned by a corporation?
kasey_junk 12 hours ago [-]
Books in the 17th and 18th century often didnt get bound by the publisher. They were done by independent binders for _custom_ orders. You’d see whole libraries with the owners binding/cover standards rather than per book.
Books in that time were _luxury_ goods. Most people could not afford them. One of the ways that was changed was to introduce cheap, mass produced bindings that were lower quality than the bespoke artisianal bindings done by specialist craftsmen.
You can still get custom bindings done. There exists whole niches on the internet of crafters that will take a production run book and strip its binding and make you extremely high quality and custom bindings and covers.
ryanmcbride 12 hours ago [-]
Learning this fact is what got me interested into binding my own books!
With how the quality of things seems to have been degrading over the years (either real or just me getting older and experiencing the impermanence of all things) I've been trying to adopt an attitude of "if this practice existed before the industrial revolution, I can _probably_ do it" and it's been really great to learn how things were made before they had to be mass produced as cheaply as possible.
DrewADesign 12 hours ago [-]
I worked in an academic library that rebound pretty much every book they acquired. One of the biggest in the world, too.
kasey_junk 12 hours ago [-]
My wife worked for a company that specialized in rebinding paperbacks for schools and libraries. It makes more economic sense to do that for niche use cases rather than make all print runs more expensive.
DrewADesign 9 hours ago [-]
Yeah it depends on the library and collection. If it’s a general-purpose library, probably a pretty limited portion of the collection. If it’s a grad school library with a low-turnover collection, like the one I worked in, that’s totally different. The college I worked for had dozens of libraries and their use cases were all pretty different.
pfdietz 12 hours ago [-]
Nowadays, academic libraries are moving their books to archival storage. New publications are electronic only. You need to be a formal member of the university community to access it, unlike in times past when any member of the public could stroll the aisles of physical books and journals.
DrewADesign 10 hours ago [-]
Most academic libraries have large archival collections anyway, and the extent to which they de-emphasize their patron-facing stacks varies dramatically between institutions. Same with who can access them. You just can’t generalize like that.
Lutzb 12 hours ago [-]
One of my relatives is one of those micro publishers. Selected works are printed and bound to extremely high standards and materials in a way that their customers are willing to pay 1k-25k+ per book. These editions only have a couple of prints and are mostly made to customer requests.
There is a market for these type of books, albeit a very small one.
ascagnel_ 11 hours ago [-]
You don't even need custom bindings —- library bindings are a thing, even if they cost significantly more then a typical hard cover binding.
m4rtink 9 hours ago [-]
The Czech National libraries do that for newspapers, magazines and other periodicals - they bind the copies they get automatically for preservation to big books, so they can be better stored in their archives.
Apparently it is getting harder to find people who can do that as most schools no longer have book binding as a course you can study.
tookmund 12 hours ago [-]
> I often see well-made books from the 17th or 18th centuries which are still in good nick
At the risk of stating the obvious, any poorly made books from then wouldn’t have lasted this long and so you would never see them.
> It's ridiculous that in the 21st century, publication standards have fallen to the point where for most works a disposable format is the only type available - where no amount of money could buy a truly decent hardback copy.
This is purely a response to market demand. Publishers aren’t going to put in the extra expense of binding high quality versions of every book so it can occupy warehouse space while consumers everywhere buy the cheap paperback.
> And if the publisher keeps a book in print - but only in bestseller-grade materials, bogroll paper that furrows in any humidity and perfect binding that molts its pages a couple of dry seasons later - and if it refuses to print a durable hardback copy with signatures, good paper and decent print - something that will still be readable in several generations' time - any other publisher keen to have a crack at it should be able to, again without any compensation for the original publisher, though perhaps in this case, with matching royalties for the author.
This is all based on the idea that there is hidden demand for something, but publishers are choosing to deprive us all of it for reasons. That if we open up the laws, another company will come along and satisfy this hidden market opportunity and associated profits that publishers are declining to take.
The simpler explanation is that these high quality editions aren’t being published because the publishers have the data about demand for them. They know they won’t sell.
If the goal is preservation, laws forcing publishers to print on slightly nicer paper isn’t going to solve the problem. It needs to be a robust digital archive and it needs to exist somewhere other than in unsold warehouse inventory or some book collector’s shelf. You’re trying to solve a problem with last century’s technology.
abdullahkhalids 2 hours ago [-]
Large publishers are optimizing their overall revenue/profit, not the revenue/profit per book. Given that large publishers dictate a significant fraction of what appears on bookstore shelves, how concentrated publicity drives of new books/authors drives their sales at the expense of old books/authors, and how the emergence of second hand books outside their control drives sales in different categories, its extremely likely that the advantageous strategy often is to not print older titles even when there is demand for them.
ShinyLeftPad 12 hours ago [-]
I have even less sympathy for IP stealing LLM operators
phoghed 12 hours ago [-]
It’s been determined that training on lawfully acquired works is fair use. Presumably in this discussion of shredding physical books Dario and Sam are not pulling heists at the local library.
Might surprise people to learn but the law isn't the final arbiter on what is moral or just. The law only decides what is legal and as we've seen over our lived history as humans, many "legal" things may not be those we want society to uphold.
TeMPOraL 10 hours ago [-]
It's not really a surprise to many here, but we don't want to go back there too often, because that topic is well trodden - the intellectual property in its current form is far worse offender against "what is moral or just" than anything LLM companies ever did.
In fact, the whole problem of shredding books (destructive format-shifting) was created by copyright laws in the first place, and that in itself is a concession hard won against the IP establishment - and all that way before LLMs became a thing.
phoghed 11 hours ago [-]
Yes IP theft and fair use are discussing legality, not morality. Whether or not it’s moral for someone to make money off somebody else’s labor or Reddit comments, I don’t want to get into a discussion about personally.
roarcher 10 hours ago [-]
The comment you originally replied to was clearly discussing morality. You don't get to insert your little "erm, ackshually" and then pretend to be a neutral observer.
phoghed 6 hours ago [-]
lmao no it’s not
rc5150 10 hours ago [-]
Okay, then step aside while others discuss the morality.
phoghed 10 hours ago [-]
You don’t seem to be doing so, but I eagerly await your doubtless profound and unique insights on the topic.
> step aside
Isn’t that what I explicitly just did in the comment you replied to?
10 hours ago [-]
HedonicEscal8r 10 hours ago [-]
Yup, and IP theft is morally good. Are we not on HackerNews?
tancop 6 hours ago [-]
no, we have always been on law-abiding-citizen-except-when-its-white-collar-crime news. the only moral problem here is people cant decide if copyright landlords or ai scammers are better for their portfolio.
sofixa 11 hours ago [-]
No, courts so far in the jurisdictions which have heard such cases, have ruled it's fair use. There is plenty of ongoing litigation in many jurisdictions, so it's way too early to just decree "it's been determined". It likely won't be for years to come.
bscphil 11 hours ago [-]
Why did Anthropic settle with authors for 1.5 billion then? Surely their lawyers must have decided there's a pretty good chance of judges ultimately deciding that it is copyright infringement?
Anthropic settled because even though the training is fair use, Anthropic did not acquire all of the training material through legal means.
nativeit 10 hours ago [-]
If a case is settled by the parties, it cannot be cited as establishing precedent.
xienze 10 hours ago [-]
You could also ask why the other side agreed to that settlement. It's not a one-way street.
qwytw 9 hours ago [-]
It was about piracy, though? Not fair use.
xienze 6 hours ago [-]
What does that matter? They were sued and the group suing them accepted at $1.5B settlement instead of something theoretically much larger and precedent-setting.
qwytw 5 hours ago [-]
Because the judge ruled that using the books for training was "fair use". The issue was Anthropic downloading and storing pirated content and there is quite a bit of precedent regarding that already...
EnergyAmy 7 hours ago [-]
Did you manually set those utm_source parameters?
ButlerianJihad 12 hours ago [-]
You don’t know what “Fair Use” is.
Fair Use is not an activity that you engage in. Fair Use is not a category with criteria that you meet. Fair Use is not a precedent that paves the way for everything afterwards.
Fair Use is a defense that can be used in court when you’re named in a copyright lawsuit. Fair Use is how you justify your actions before the court finds infringement.
hn_acker 11 hours ago [-]
Fair use is a legal defense with a specific test to demonstrate whether a particular instance of use of a copyrighted work is not infringement. Although fair use is always evaluated on a case by case basis, fair use does produce some precedent (example at [1], not related to TFA), and fair use is not mere justification [2]:
> Notwithstanding the provisions of sections 106 and 106A, the fair use of a copyrighted work... is not an infringement of copyright.
Sure, everywhere it’s been legally tested so far and there’s been a conclusion, the outcome has been in favor of the LLM trainer. They just can’t steal the books, they have to pay for them. So claims by random HN users that it’s IP theft and copyright infringement and all that are currently incorrect.
Now if you look at how fair use is used colloquially, everyone understood what I meant except the autistic pedants.
gruez 12 hours ago [-]
>>It’s been determined that training on lawfully acquired works is fair use
>You don’t know what “Fair Use” is.
This isn't the opinion of some armchair HN commenter. Actual judges have affirmed this, as other commenters in this thread has pointed out.
Marsymars 10 hours ago [-]
Depends on jurisdiction, clearly.
e.g. Canada doesn't have fair use, but from wiki on fair dealing in Canada, "According to the Supreme Court of Canada, it is more than a simple defence; it is an integral part of the Copyright Act of Canada, providing balance between the rights of owners and users."
parineum 12 hours ago [-]
s/fair use/ip theft
The comment uses the same language as it's parent.
parineum 12 hours ago [-]
Being informed is much too high a barrier for the people who just want to pretend that everyone with money is lex luther.
arduanika 11 hours ago [-]
If you're going to call people ignorant, you should probably know how to spell the name of the character you're referencing.
parineum 7 hours ago [-]
I'm sorry I don't read comic books. Seems you got the idea though.
DonsDiscountGas 10 hours ago [-]
It's not copying (and hence not stealing) if you destroy the original book. That is where copyright law has brought us
MemoryHoleHQ 11 hours ago [-]
[dead]
kevin_thibedeau 12 hours ago [-]
> I often see well-made books from the 17th or 18th centuries which are still in good nick.
Those books predate the development of wood pulp paper. It isn't the publisher's fault they can't economically print on rag paper anymore.
ctolsen 10 hours ago [-]
I'd go for the less complex version: just make copyright last for like 20 years at most.
nyeah 11 hours ago [-]
Yeah it's as if the publishers are somehow strapped for cash compared to AI companies.
jolmg 10 hours ago [-]
> I've limited sympathy for the publishers.
This doesn't affect publishers.
It affects humanity as a whole by having large private parties hoard books to partly destroy them. The covers and spines can have historical value too. There's not even a reason for these companies to release their scans to the public once the copyright expires.
It prevents proper preservation by archivists and preservationists.
cataphract 9 hours ago [-]
The post doesn't call for the sympathy for the publishers, so I'm not sure what you're responding to.
trescenzi 10 hours ago [-]
No one reasonable is concerned for publishers they are concerned for the books. Assuming the allegations are correct historical artifacts are being destroyed with no trace. It doesn’t even matter if they republish it. The artifact itself and the time it has transited is the valuable thing. It is simply immoral to destroy stuff like this.
streetfighter64 12 hours ago [-]
> When a publisher has no incentive to keep an edition in print, it should be available to any other publisher to print
How do you adjudicate that? And wouldn't it just lead to loopholes such as "Ghost Printings" (cf. Ghost Flights https://en.wikipedia.org/wiki/Ghost_flight_(commercial_aviat... ) where the books are technically printed in the required volume but practically unavailable to customers through one method or another. Because, the cost of wastefully printing a few books to warehouse, is less than the potential losses of the IP rights, probably.
flir 11 hours ago [-]
I think putting the works in a print-on-demand catalogue at an unreasonable price would be more likely.
But there are worse outcomes.
eudamoniac 10 hours ago [-]
I've never read a coherent argument for why all of these "should"s should be. I understand you want to read the books. But I don't see how that desire results in the laws needing to be changed so that you can read everything you want, even things the author or owner doesn't want you to read.
There is not a shortage of books in print. I don't understand why people get so hung up on a few of them being out of print. It is okay for someone to own something cool and not let anyone see it. The cool thing doesn't suddenly become a societal necessity because it is words written down.
jolmg 10 hours ago [-]
> There is not a shortage of books in print. I don't understand why people get so hung up on a few of them being out of print.
Because the point isn't so much in reading whatever text as if all text is the same. The point is the spreading of knowledge. A single book can contain knowledge not present in any other.
> But I don't see how that desire results in the laws needing to be changed so that you can read everything you want
In order for your "people" (country, etc.) to do better, you want them to be educated. In order for people to understand one another, you want them to be able to see all the same various perspectives there are. It makes perfect sense for laws to aim for these goals. This is why libraries exist.
> It is okay for someone to own something cool and not let anyone see it. The cool thing doesn't suddenly become a societal necessity because it is words written down.
Books aren't simply trinkets, like an item you bought at a gift-shop.
> I understand you want to read the books.
I think you're looking at this too much as what people want for their own individual selves, when it's more of what people want for everyone. It's about what they believe is best for society as a whole. They don't need to want to read a book themselves.
eudamoniac 8 hours ago [-]
Which books are in copyright, not being published, and contain unique valuable knowledge? I've never seen this be the case. It's always some genre fiction or obsolete botanical text or something. I've never seen valuable knowledge locked away like you are describing.
TFNA 5 hours ago [-]
"Which books are in copyright, not being published, and contain unique valuable knowledge?" Most works in a large number of academic domains from the twentieth century.
In my own field (at the intersection of linguistics, history and archaeology) I am constantly, multiple times per day, referring to information in publications from e.g the 1960s or 1970s that never got republished later except as a brief citation to that old book. And guess what, nearly all that twentieth-century scholarship is still under copyright, often from the big German or Dutch publishers that enforce their claims fiercely. The academic community has been doing a lot of work to scan our institutional libraries and upload them to the shadow libraries, but this is still all illegal copyright violation.
tesseract 3 hours ago [-]
You're going to make me quote the "Dive Manual" scene from _Cryptonomicon_ again, aren't you.
(It features a software engineer introspecting about the fact that his line of work has caused him to dramatically overvalue recency when evaluating books about... pretty much any other technical field.)
tourmalinetaco 7 hours ago [-]
15 years, 5 year extension, that’s it. That’s all we ever need, and there is no logical reason to extend it outside of the kind of greed they warned about in The Bible. Additionally, any and all source code, assets, etc should all be entered into the copyright office and released when it enters the public domain. This should even include printing assets and high res copies of artwork.
Copyright can be intrinsic, sure, but only to maybe 5-10yrs, far less than the maximum to incentivize registration (and thus incentivizing archiving of culture).
I should be able to freely download every Disney film prior to 2006, compile Killer7 for the fun of it, buy a hardback of the LOTR trilogy from any printer I wish, and do so without any interference from the copyright holders.
However, because we don’t live in a utopia and Sonny Bozo became a politician, I’ll be waving a certain jolly flag for the foreseeable future.
nobodyandproud 8 hours ago [-]
I agree. However, AI companies have managed to make publishers and copyright owners saints by comparison.
andrepd 9 hours ago [-]
Thanks for your thoughts, but how is this is anyway relevant to the issue at hand? I'm sad that this is the top comment...
SwtCyber 10 hours ago [-]
[dead]
stellamariesays 12 hours ago [-]
[flagged]
trollbridge 13 hours ago [-]
We reprint old books after checking out copyrights (for all books, this means pre-1930, but for some (I'd actually say most) it also means ones published up to 1964 and 1973, depending on how the rightsholders did (or didn't) do the renewals).
We use a special guillotine type cutter to cut off the binding and then store the pages in a sealed plastic bag which goes in the archives; they're stored there indefinitely in case the book needs rescanned for some reason. We also keep the original, uncompressed copies of the books on magnetic disks.
We also go out of our way to try to find rare books published in 1931, 1932, etc. so they are ready to go once the copyright expires.
And no, no AI company has ever come to us and asked to run training on all of our scanned copies.
palmotea 12 hours ago [-]
> We reprint old books after checking out copyrights
Who is we?
> then store the pages in a sealed plastic bag which goes in the archives; they're stored there indefinitely
Is that the best thing for archival storage? Like could things like chemical breakdown increase the humidity in the sealed bag or concentrate corrosive chemical vapors? I was under the impression the best environment was an actively climate-controlled environment.
ComputerPerson 11 hours ago [-]
The issue is mold. Sealed plastic bags for old books is a recipe for mold. At least they're individually stored, so they would develop mold at independent rates. (And maybe the seal allows airflow.)
Most books have some mold by the time they're ~100 years old, it's just not enough to cause a serious problem. Sealed enclosures (wrappings, bags, tubs) are a nightmare situation. Even packing books too tighly on shelves accelerates mold growth to problematic levels.
starkparker 10 hours ago [-]
Dealing with mold in an archival environment is a nightmare, and yeah, the only thing sealing the books in plastic does is make it easier to either throw them away, or to freeze them to try to kill and then manually clean out the mold. https://www.nedcc.org/free-resources/preservation-leaflets/3...
pfdietz 10 hours ago [-]
If I did that I'd seal them with silica gel to keep the humidity down.
8 hours ago [-]
qingcharles 9 hours ago [-]
I would think this totally dries the pages out and then they just turn to dust, from experience.
ComputerPerson 8 hours ago [-]
A well-organized system derived from binder clips works for small/medium scale collections of debound books. It's a quick and easy, and relatively cheap, way to "rebind" them in parts. The clips even provide good surfaces for labels.
Archivists recommend standard ambient conditions or a little drier for long term storage. As you've said, too dry and the pages fall apart; permanent damage.
I suppose the fancy silica gels that maintain specific humidities would work in bags.
pfdietz 4 hours ago [-]
My understanding is that the fibers in paper can become brittle if they are too dry. However, this doesn't cause the paper to break down, it just makes it vulnerable to damage if flexed. So storing paper dry is fine if it is then humidified before being stressed.
qingcharles 3 hours ago [-]
It's possible. My only experience is with vintage newspapers that were allowed to dry out, and then when trying to open them just shatter like a carbonized volcano scroll.
rolandog 3 hours ago [-]
It's truly humbling how much knowledge professionals from different fields have, and how easy it is to fall into a Dunning–Kruger effect trap where a bit of knowledge — extremely extrapolated — would've led me down the path of hubris and guessing so many wrong answers. Thanks all.
qingcharles 1 hours ago [-]
One of the best things about this site is how many people there are on here who are way smarter than me, with much deeper domain knowledge.
palmotea 8 hours ago [-]
>> If I did that I'd seal them with silica gel to keep the humidity down.
> I would think this totally dries the pages out and then they just turn to dust, from experience.
They used to sell Boveda two-way humidity control packs that would keep a bag at a constant low-ish 32% humidity, but last I looked those were discontinued (it looks like they've pivoted pretty heavily to marijuana storage and higher-humidity products).
Aurornis 10 hours ago [-]
> And no, no AI company has ever come to us and asked to run training on all of our scanned copies
The value of most very old books for AI training is very low. You don’t really want your AI training data to start biasing toward outdated writing styles. Most of the valuable knowledge has been covered again in modern texts in more depth and detail.
There is interesting value in old texts and it’s important to have them archived. It’s less valuable for stirring into the giant pot of AI training data, though.
sosodev 10 hours ago [-]
If you only care about facts, maybe. Even then I'm sure there are countless facts not described outside of old books.
I have a hard time believing that text valuable to humans would not be valuable to AI.
dgellow 9 hours ago [-]
Agreed, in addition LLMs are trained on a lot of useless data, such as the whole Reddit dataset which has to be at least 90% Reddit garbage, and that doesn’t seem to be a problem. I don’t see why ai labs wouldn’t want to also train on older data
5 hours ago [-]
vander_elst 12 hours ago [-]
Some source or citation or context is needed here, is this the work of a 2 person no profit or a trillion valued pre IPO company?
11 hours ago [-]
butlike 12 hours ago [-]
Why cut off the spines? Isn't that how you end up getting unattributed 'dead sea scrolls'?
remus 12 hours ago [-]
It's easier to get good quality scans from individual pages than it is from pages in a complete book. Imagine laying a book flat, then the page is distorted in the area around the spine. You can work around this (either by trying to correct for the distortion in software or with clever scanners that position the book more advantageously) but it adds complexity compared to chopping off the spine and just dealing with flat sheets of paper.
ryukoposting 12 hours ago [-]
Speed and cost. You can run the book through a typical sheet-fed scanner instead of using a contraption like this: https://linearbookscanner.org/
As for consumer-grade solutions, look for the Fujitsu SV600.
phasefactor 12 hours ago [-]
Easier to send them through a duplex scanner (or put them in a flatbed one if they are fragile). Cheaper than buying the automated ones with the page turning robot arm.
I have done it at home for my books since the mid-00s.
open-paren 12 hours ago [-]
Do you purchase every book two times, or do you have a home devoid of physical books that you enjoy? Genuinely curious
JKCalhoun 10 hours ago [-]
Not OP, but I do similar. I only de-spine a book if it is in poor condition (and I have a nice copy). Some of the material I scan is stapled instead of glued and I can remove the staples, scan and replace with new (not rusted!) staples to return the book to its original form.
I have books but would have 5x as many if I could not capture them digitally. (When I go to move or go through a purge, some of the books I scanned do go to a used book store.)
So I scan in part to keep my physical book-footprint smaller, but also my scans all get cleaned up and uploaded to archive.org. Mainly I scan young-adult science books from the 50's and 60's (since they were so influential and have all but disappeared except on eBay and the like).
tourmalinetaco 7 hours ago [-]
Thank you for your service, genuinely.
flir 11 hours ago [-]
Not OP, but: If it's something I want to work from, I get a printer to slice the spine off and rebind it with a ring binding. That way it lays flat on the desk.
You've trashed the book (and its lifespan) but some books are for using, not keeping.
htrp 12 hours ago [-]
> And no, no AI company has ever come to us and asked to run training on all of our scanned copies
Yet
yorwba 12 hours ago [-]
Most likely they don't know you exist. If you contacted them first, maybe something could be arranged...
freejazz 12 hours ago [-]
Geez - their LLMs couldn't figure it out?
phoghed 12 hours ago [-]
Can you?
freejazz 11 hours ago [-]
I don't know which company the OP belongs to in particular, but searching online reveals plenty of similar cos.
ck2 12 hours ago [-]
or just send the books/scans to a country that doesn't recognize US copyright
goldlimetea 4 hours ago [-]
[flagged]
est31 13 hours ago [-]
> You can reprint a bestseller. You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal. So it's going to accelerate.
Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright?
IDK about the shredding, it's not nice, but it's more a problem with copyright law than AI companies.
Scanning books you own should be legal from a copyright point of view, and not require shredding.
Second, one should think about abandoned property provisions for copyright works published more than 50 years ago and in danger of being forgotten: once challenged, either you as the owner have to prove that the work is preserved for future generations (e.g. in various libraries around the world), or you have to authorize further copies, or you give up copyright on the work.
ACCount37 13 hours ago [-]
Scanning books by taking them apart into singular pages and scanning those pages is faster and cheaper. AI training is a numbers game, so they want faster and cheaper.
What happens to the pages after? No one needs them anymore, so they get mulched and recycled.
That would be the dominant scanning method even if copyright wasn't a thing. But then again - if copyright wasn't a thing, there would be much less need to scan any physical media.
The reason why OpenAI can't just go on Amazon, buy a "digital edition" of a 2018 book and use that is that it would violate the license in ten ways, and then the DMCA laws that forbid breaking DRM on top of it.
voakbasda 12 hours ago [-]
Think about that last point for a moment. Our “rights to read” are diminished significantly with digital works as compared to printed works. Right of resale. Right to lend.
In the end, digital publishing just isn’t right and will lead to massive gap in our historical records. They require active curation and cannot be preserved simply by resting on a dusty shelf.
nomel 5 hours ago [-]
> Our “rights to read” are diminished significantly with digital works as compared to printed works.
I see it as opposite. I can hand someone the complete contents of a public library on a thumb drive. Delivering that to their door is going to be much trickier.
Reading and distribution (pirate like) has been made MUCH easier, with non-physical distribution. I'm always carrying a 1" thick book in my pocket, that I can read wherever I am. I was ecstatic when I switched over to digital. I could read anywhere!!!
butlike 12 hours ago [-]
Every innovation since the microprocessor isn't worth saving in the grand scheme of things.
When today's algae evolve enough into tomorrow's sentient creatures, they're really only going to need up to the industrial revolution and should probably stop right before that.
compass_copium 10 hours ago [-]
I'm personally a fan of more than 50% of children surviving past the age of 6, something that didn't happen until the 20th century.
WorldPeas 7 hours ago [-]
the mass-market-ification of: clothing(world clothelessness used to be a severe problem, paper clothes existed for a reason), AC, light vehicles, solar: all things after the 20th (and even 21st century) that are worth saving. Even with the seeming doom, things do get better.
TeMPOraL 10 hours ago [-]
Machines for non-destructively scanning books were developed and perfected long ago. The destructive scanning is neither technological limitation nor an issue of expedience. It's an issue of copyright law and fair use.
k8si 7 hours ago [-]
Why is the shredding a result of the fair use stuff? I actually don't understand
tourmalinetaco 6 hours ago [-]
They’re not legally allowed to keep a physical and digital copy at one time because pf copyright law, and fair use doesn’t cover it as an exception.
edoloughlin 11 hours ago [-]
> What happens to the pages after? No one needs them anymore, so they get mulched and recycled.
Strictly speaking, no one needs the Sistine Chapel or the Pietà etc. It would be a shame if they were mulched and recycled, though.
doublerabbit 10 hours ago [-]
Same with the magna carta and the American constitution.
ChatGPT know them, i'd count that as digitalised why keep the originals?
dragonwriter 12 hours ago [-]
> if copyright wasn't a thing, there would be much less need to scan any physical media.
Because there’d be much less content created in any media to capture in the first place.
eru 12 hours ago [-]
Empirically, probably not. We had lots and lots of content before copyright, and people seem to produce lots of content even in jurisdictions with weaker copyright.
rstuart4133 3 hours ago [-]
It's kinda amusing copyright always seems to attract discussions about authors rights. It was never about author or their rights. It was and is a set of rules that allow money to be made by publishing books and songs. Authors have a role in that, but so do publishers.
Turns out most authors and publishers suck at their respective jobs. Most books produced by authors are things nobody wants to read. If a publisher's job is defined to be finding works the public likes, they suck at it too. Instead they just publish lots of stuff, mostly at a loss. They make their money from the occasional hit. But only way they can make money from it is if they have a monopoly over publishing it for a while - which is exactly what copyright gives them. Looked at in another way, this "publish lots of stuff and see what sticks" is the way our society discovers what new works are popular. And copyright funds it.
If you take the view that copyright is paying for the discovery and publishing of new works, then consumers paying monopoly prices for 70 years or more is a bad idea. By far the majority of works are commercially dead 1 year after being published. The publishers tend to make their money from the remaining 1%. They might last five years. Very, very few last 20. I don't see how copyright lasting beyond 20 years can be justified with anything other than: "because of the good work he did 20 years ago, he deserves to be paid for sitting on his arse for the rest of his life". In reality, he didn't sit on his arse - he donated to a few politicians - but the outcome is almost the same.
classified 12 hours ago [-]
It's the law, logic doesn't enter into it.
graemep 12 hours ago [-]
> You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal.
An 18th century book would be out of copyright so why would it be illegal to keep the original and scan it?
11 hours ago [-]
sethops1 13 hours ago [-]
> Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright?
It's cheaper to scan the books if you do it destructively. Cost. That's why they're shredding irreplaceable texts. Nothing to do with copyright.
One doesn’t need to pulp the pages after scanning though. After scanning, they could be rebound and put into a library.
dpark 10 hours ago [-]
That would be a clear case of copyright infringement under current law. You can’t make a copy of a book and then give the original to someone else.
breakyerself 9 hours ago [-]
If it's copyright is expired why not?
dpark 8 hours ago [-]
The books in question are not antique. The Twitter post makes this claim but so far as I can tell it’s not based in fact.
404 Media published a story about this as well and cites a bookseller who notes that all of the books there have sold have had ISBNs (and are thus from 1967 or later and generally would have active copyright).
”very large purchases were of books that had little in common, except for the fact that they all had ISBNs. This seller also sells rare books that do not have ISBNs, and none of those were part of the bulk purchases”
This is the necessary context distilled down to be concise. Thank you.
mc32 13 hours ago [-]
Old rare books where there are single digit copies should enjoy some sort of patrimonial protection just like museum pieces. You can own them but have the state have the option to buy it if you’re about to significantly deface it or destroy it.
soco 12 hours ago [-]
The only issue I see is, how could you tell which are those books?
qingcharles 9 hours ago [-]
You can't, without spending a fortune to investigate the scarcity of some of the titles. I do a little work in this space and I don't know of anything destructively scanned that has zero other copies, but definitely some with single digit known copies, where no previous scan exists. Now there is one less physical copy, but a digital copy that none of us can access (except by tricking an LLM).
It's not just the big players buying up these archives, either. There are a lot of smaller players, especially in the OCR space, who are buying up huge swathes of works in languages which have much smaller digital footprints, e.g. Arabic.
9dev 12 hours ago [-]
We manage to do this for endangered wildlife too without anyone counting every single specimen; why shouldn’t we be able to estimate how rare a book is?
qingcharles 8 hours ago [-]
The same techniques roughly apply. You can look at all the book marketplaces and see how many of each title are available and it'll give you a rough estimate. If you look across all the marketplaces, all the library catalogs, plus eBay etc and zero copies surface, even going into Worthpoint to look at the last several years of eBay sales, then you know you have a problem. That's how I normally estimate it.
After that you are diving into forum posts etc to see if you can find anyone who has even mentioned owning a copy or having seen a copy.
I don't know what happens to some works. Supposedly thousands, tens of thousands, or sometimes apparently a million or more copies published and yet not a single copy surfaces for years.
eru 12 hours ago [-]
It's pretty expensive for the wildlife.
Most rare books are rare because no one cared enough about them. Ie most rare books are rubbish.
phoghed 12 hours ago [-]
And if you instituted this, the commenters of this very website would surely decry it as a prime example of government overstep and waste.
shimman 11 hours ago [-]
Commentators on this web site work for some of the most evil organizations on the planet and have beliefs that 95% of the population rejects. You can safely ignore the YC cohort of devs and be fine.
ctoth 9 hours ago [-]
> Commentators on this web site work for some of the most evil organizations on the planet and have beliefs that 95% of the population rejects. You can safely ignore the YC cohort of devs and be fine.
1. Why are you here?
2. What is the purpose of this comment?
croes 13 hours ago [-]
Books that are shredded can’t be scanned by competitors.
jfyi 10 hours ago [-]
Yeah, this is the point. I don't understand the bulk of this conversation. Copyright doesn't matter, the books themselves don't matter. All that matters is that their corpus of training data grows faster than their competitors.
lousken 13 hours ago [-]
That's why archive.org should have never been sued for lending books they had physical copy of. This is the result. Publishers should be more careful what they wish for.
kingstnap 13 hours ago [-]
The archive.org story was more nuanced than that. If I recall correctly the full story was that they used to lend digital versions of books they physically bought and scanned with DRM to enforce a sort of one to one at a time restriction.
But during covid archive.org decided to just remove the limit and lend unlimited copies concurrently which started the debacle with the publishers.
cvadict 12 hours ago [-]
> But during covid archive.org decided to just remove the limit and lend unlimited copies concurrently
IIRC, this was 100% it. Lending one digital version of one physical asset was likely already a violation copyright. Lending UNLIMITED digital versions of one physical copy was DEFINITELY a blatant violation of copyright.
phasefactor 12 hours ago [-]
Correct, it was switching to unlimited lending instead of one lend per physical book that got them in trouble.
butlike 12 hours ago [-]
How is lending one digital version of one physical asset a violation of copyright? Since I MAY be able to lend out the physical as well?
ndiddy 12 hours ago [-]
From the court decision:
"IA maintains that it delivers each Work “only to one already entitled to view [it]”―i.e., the one person who would be entitled to check out the physical copy of each Work. But this characterization confuses IA’s practices with traditional library lending of print books. IA does not perform the traditional functions of a library; it prepares derivatives of Publishers’ Works and delivers those derivatives to its users in full. That Section 108 allows libraries to make a small number of copies for preservation and replacement purposes does not mean that IA can prepare and distribute derivative works en masse and assert that it is simply performing the traditional functions of a library. 17 U.S.C. § 108; see also, e.g., ReDigi, 910 F.3d at 658 (“We are not free to disregard the terms of the statute merely because the entity performing an unauthorized reproduction makes efforts to nullify its consequences by the counterbalancing destruction of the preexisting phonorecords.”)."
card_zero 10 hours ago [-]
Derivative works? That seems to translate as "because the books are digitized, it's strange and new and we can't allow it".
ndiddy 9 hours ago [-]
[dead]
Cthulhu_ 12 hours ago [-]
Yeah that was it; if I got this right, US libraries got the right to lend out one digital version of a book that they had in their inventory. Archive.org combined those digital versions so that people could check out a digital book if any library in the US had it (digitally) available. But during the 'rona they removed this limit and just lent out books regardless of it being "checked out" digitally from a library.
This wasn't a very smart move of them. I get why they did it but they put themselves at a huge legal risk.
lousken 10 hours ago [-]
If i recall correctly, they had permission from physical libraries to use their copies as well, so it wasn't just a single copy but many copies, just one of them was converted into digital. Still wasn't enough apparently...
Incipient 13 hours ago [-]
Publishers don't care if rare books get shredded?
the-grump 13 hours ago [-]
And, regrettably, The Archive lent books regardless of physical possession.
Publishers had accepted the prior arrangement before The Archive decided to push it, if not explicitly then implicitly by not suing.
I'm a believer in The Archive's mission, and I wish they had treated the goodwill they'd accumulated as something worth preserving and not a currency to be spent.
It has been stated by many before me: lending books should have been handled by a separate entity, especially when they removed the physical backing requirement.
kmeisthax 12 hours ago [-]
Just to be clear, publishers hadn't accepted the "controlled digital lending" (CDL) premise, not even with the one-to-one ratio. Their position was always "first sale ends when the atoms do". There was even controlling precedent: a few years before IA tried their online lending library thing, there was an "MP3 resale" company called ReDigi that had lost on very similar grounds. The publishers suing IA even made sure to sue in the same venue that had decided the ReDigi case so it'd be controlling precedent.
Furthermore, in the discovery for the Internet Archive case, publishers had already found a case where IA had lent out books despite knowing their partner libraries wasn't actually withdrawing loaned-out copies from circulation. The CDL premise was always just a suggestion, and IA would have still lost their case if they hadn't done the National Emergency Library (NEL) stunt or if they'd been sued in another venue that hadn't had the ReDigi case as precedent.
It's important to note that whenever a company decides to sue for copyright, it is often late, because the company is banking infringements up to the 3-year statute of limitations and because building a meritorious case takes time. The lack of a timely lawsuit proves almost nothing about the intent of a publisher with a valid case against you.
The thing is, I don't even think the whole stunt damaged much of the IA's goodwill? I know of a few people who withheld donations to IA, but that was mainly under the assumption that publishers would be getting a billion-dollar damage award that would immediately bankrupt IA and result in it's archives being sold off to Lexis-Nexis or something. The funny thing is, IA wound up settling for a sum so small they had to promise never to reveal it, and the danger is gone, so the only thing people complain about now is just that the NEL stunt maybe pushed them "above the radar" or something.
It's still insane that shredding books for AI training is legal, but this isn't.
TeMPOraL 10 hours ago [-]
The big insanity is tying this to AI. Shredding books is about format shifting; it's a concession hard-won from copyright establishment, which would otherwise be more than happy to deny you the option to convert the media you owned from physical to digital.
AI training happens to be one of the fields exercising that option, but since it's the current favorite topic for people to hate on, here we are.
kmeisthax 7 hours ago [-]
You don't need to shred books to scan them. They make book scanners that will "rip" a fully-bound book no-problem, and even correct for the curvature of the page and binding to give you an equivalent image. In fact, the Internet Archive specifically built nondestructive book scanners[0] for exactly the purpose of which AI companies are now shredding books. The smart / savvy thing to do would be to buy those machines off IA and use them to read the books they're interested in.
The reason why AI companies don't do this is that they're cheap and desperate for training tokens. Same reason why they have scrapers that will happily overload web interfaces for Git repos following links to everything, even though you can just Git clone the repo with far less stress on the host. The AI people are ultimately there just to pillage as much knowledge as they can as fast as possible. Their scraping practices are slap-dash garbage.
Again, the entire point I'm making is that the choice is dictated by copyright law. They can't legally scan books non-destructively. They can legally format-shift them, i.e. scan them destructively. So this is what they're doing.
Your comment seems to also be regurgitating common misconceptions (to put it charitably) about AI and web scraping.
the-grump 4 hours ago [-]
Sure, that's why I said implicitly.
Publishers will always state a maximalist position but the truth of what they accept is what they tolerate without suing.
azan_ 13 hours ago [-]
Yeah, why would it be bad for publishers? If anything they'd most likely encourage more book shredding!
lousken 10 hours ago [-]
There are many reasons but it's just rare books - in general: they are setting a precedent for people to pirate instead of using libraries. In general companies are trying really hard to make people switch back to torrenting, pirate sites, sharing media etc.
infinite_spin 13 hours ago [-]
of the rare books, which was the rarest of them all? What year was it published?
WorldPeas 7 hours ago [-]
> Publishers should be more careful what they wish for.
you think they don't like people destroying their books? if you destroy a book, that increases the value of the others and the unpublished holdings (unfortunately)
redsocksfan45 12 hours ago [-]
[dead]
JumpCrisscross 12 hours ago [-]
> That's why archive.org should have never been sued for lending books they had physical copy of. This is the result
How are these things remotely related? If anything, Archive.org’s callous, thoughtless approach nuked the hands of legitimate archival efforts.
ACCount37 13 hours ago [-]
The publishers sued AI companies for training on shadow library data, hoping to negotiate content deals for big $$$ down the line. Instead, they got analog hole'd.
Turns out that buying an old book for $5 and destructively scanning it for $25 is way cheaper than paying extortion fees to the copyright-mongers.
What I don't buy is it being "rare, precious books". First, they're not after ancient texts - they're after the books that there's still copyright on. Second, when it comes to books, "old" doesn't mean "valuable" - plenty of libraries destroy old books because there's no demand for them, and storage costs you. This is how those scanning companies get books for so cheap.
tencentshill 13 hours ago [-]
So they're not valuable... except to AI companies. They should pay a fair amount.
ACCount37 13 hours ago [-]
They are paying a fair amount. In the ballpark of $5 per book.
You know, piracy online is nice and simple - but it's kind of hard to get physical media without paying what the previous owner considers "a fair amount" to part with it.
chii 13 hours ago [-]
> They should pay a fair amount.
they should pay the marginal value that the next buyer would buy.
Do you also think that a person dying of thirst ought to pay the maximum price they could possibly pay for water?
qwytw 8 hours ago [-]
> the marginal value that the next buyer would buy
Well the argument is that they should pay for the right to produce derivative works not for the physical copy.
TeMPOraL 6 hours ago [-]
Given the specific of the "derived works" being produced, do you feel obliged to pay royalties every time you cite or use some knowledge acquired from a book during your daily life?
qwytw 6 hours ago [-]
Well no, but I'm a human and not a commercial computer program so it's not particularly relevant. I'm would not be making loads of money making those citations either but that's only secondary.
> derived works
It's not evident that any works are being produced in the legal sense. Since LLM outputs are not considered copyrightable it's probably closer to using a search engine.
soco 12 hours ago [-]
It might be me but I fail to see the equivalence of AI companies with thirst deaths. That', or maybe because it's a strawman.
Minor49er 11 hours ago [-]
The point is that you can buy a book for a few bucks, read it, and have its insights completely available to you. But people are demanding that LLM companies pay far more to publishers for the same copies
qwytw 8 hours ago [-]
You can also buy a book produce a large number of copies after jumbling around some chapters sell it to other people. That's illegal to do if you are person but not if you an AI company.
I suppose an argument might be made for fair use if they released their model weights publicly without financially profiting from it.
bitwize 11 hours ago [-]
If Elon Musk is dying of thirst then the price of that glass of water is set at (pinky to side of mouth) one million dollars.
TeMPOraL 10 hours ago [-]
Doing that would be considered deeply immoral and sociopathic in any sane culture. Do you really want to live in a society that finds such attitudes acceptable?
Instead of blindly jumping on a manipulated outrage bandwagon, people would do well to maybe read some of those old books, not even the rare ones - they tend to contain plenty of parables and stories explaining basics of morality and civilized conduct. We used to teach that to kids at homes and in primary education...
qwytw 8 hours ago [-]
Money is relative, though. I guess the question is if it's acceptable to charge a tiny symbolic fee for that cup of water if the other person's pockets are visibly overflowing with cash. If yes then it's more than reasonably to charge Musk a few million.
Turning it the other way around is it deeply immoral and sociopathic for someone to hoard massive amounts of money/resources if there are people dying or suffering around them (even if not on the spot but e.g. due to poor access to healthcare)?
glaslong 10 hours ago [-]
Given his approximate Trillionaire status, this is just what every vending machine doing income-based algorithmic pricing should charge him for EVERY bottle of water.
Catloafdev 10 hours ago [-]
Why are you assuming they aren't? The people who have those books are choosing to sell them for a price. What isn't fair?
Analemma_ 11 hours ago [-]
They're paying the market price for a used book and getting the same rights out of that purchase as you or I would if we bought a used book.
SirFatty 13 hours ago [-]
I see.. so the various AI companies are in the right on this?
infinite_spin 13 hours ago [-]
I think they are in the legal sense of right, and I think they only discarded the remains of these dissected books because previous rulings (e.g. archive.org's lending practices of digital copies of books they physically owned) gave rise to a situation where destruction bore less legal risk. As for the moral case, I don't have much to say on that, we all have our own lines in that sand.
jerf 12 hours ago [-]
I would say that especially for books out of copyright, they are unquestionably legally in the right. There is no legal standing for "I like books and it makes me feel squicky when someone disrespects them". Nothing stops you from buying old books and using them as decoupage[1] fodder, using them to start a campfire, or making paper airplanes out of their paper.
Moreover, while it is emotionally appealing to some people to want to add some sort of "responsibility to society" to people who own the old books, it's a very emotional plea that can't really be manifested in the real world. As already pointed out in other places, "an old book" itself doesn't really mean much in terms of what its value is in any particular dimension. Plus, I am always deeply suspicious of anything that expands to "Other people, who are not me, should expend vast quantities of resources so that in the next five or ten times I think about this issue for the rest of my life I feel slightly better about this issue" which is what this really amounts to. I think that as superficially appealing as that may be, it's really a very hostile and demanding position to take.
Personally, to the extent that I would want to lay a "social responsibility" on the AI companies, I'd like to see something like they are either obligated, or ideally, just do it of their own free will, to make the scans of the books that are out of copyright available for some reasonable fee (ideally, "free because we like the PR", but given the scope demanding it be free is not reasonable), and without them trying to lay any further claims on the public-domain results. Trading "one old book somewhere, inaccessible to the world" for "a scan of the book and an OCR of it" I would judge a net win for society for rather a lot of these old books, which are by no means "worthless" sitting in some old collection somewhere but would be a lot more useful for being available.
They could give a copy of the data once to some third party organisation that then seeds it in bittorrent or something like that. Basically, what I want to say is that this doesn't need to be an ongoing obligation for the scanner to be worthwhile for society.
11 hours ago [-]
master_crab 13 hours ago [-]
It can be the case that everyone in “a fight” is wrong.
13 hours ago [-]
yeoyeo42 12 hours ago [-]
yes. in the sense that they probably wouldnt want to do this but the law forces them to do an extremely dumb thing.
they're paying for the books, no shady things going on there. whether the publishers should deserve more than a single copy's worth is a separate question.
having the law such that its illegal to scan a book and then keep it, but legal to scan it and destroy it - gg no re there, law people retardmaxxed themselves as they tend to do with anything related to digital data.
margalabargala 12 hours ago [-]
Law aside, destructive scanning is cheaper and easier. Cutting the book into individual pages would be happening regardless, because that is the easiest, fastest, cheapest way to scan it.
TeMPOraL 10 hours ago [-]
Not necessarily, since machines for non-destructive scanning were invented and perfected a long time ago. If this is still more expensive, this very discussion shows that improved optics would've made it more than worth it to pay that extra costs. Alas, the copyright laws took away the non-destructive option.
butlike 12 hours ago [-]
Modern texts become ancient texts; given enough time.
pfdietz 10 hours ago [-]
Modern texts become piles of dust in time, given the acid paper they're printed on.
pessimizer 10 hours ago [-]
> Turns out that buying an old book for $5 and destructively scanning it for $25
More like destructively scanning it for $0.25, if you include the wear and tear on the machine, the salary of the guy who unjams it when it chokes, and recycling fees.
silverlimetea 11 hours ago [-]
This is such propaganda lol
They change the info (many buyers), destroy the source material, and now the lie is in the LLM.
That's it!
wwweston 12 hours ago [-]
Buying up old books is legal. Digitizing and distilling them arguably is too. And also:
> paying extortion fees to the copyright-mongers
Yes yes, greedy fat cat publishing oligarchs treading on the poor put-upon scrappy AI underdogs. /s
Back in reality, the fraction of people who got into publishing books to get rich collecting rents is… not large. There’s so many other fields that are likely to reward participants with more wealth that it’s absurd — even with all the passion for the work in tech it’s probably relatively less pure.
And whatever the excesses of copyright have been, the whole bargain has always been on more pro-social foundations and stronger intellectual foundations than “extortion” sneers. It recognizes that incentives matter and work that’s valuable should be rewarded and incentivized.
A culture that takes a Robin Hood approach to low marginal cost billing points but fawns over the hypercapitalized distribution King Johns isn’t creating a freer or richer society or fighting the real cartel center, it’s indulging resentment and caricature.
pessimizer 10 hours ago [-]
> Back in reality, the fraction of people who got into publishing books to get rich collecting rents is… not large.
That's because the business is buying copyrights in bulk from those people. Copyright-mongers ≠ publishers or writers.
sherr 13 hours ago [-]
I see mentions of Bradbury's "Fahrenheit 451" in that thread but what this really seems to be mostly like is Vernor Vinge's "shred and scan" factory in his novel "Rainbows End".
vessenes 13 hours ago [-]
Perhaps the last great near-term predictor. I often wish he'd written more. To remind us all, he predicted shred and scan would be a short stop over done by villains on the way to nondestructive scanning.
That said, supporting Anna's archive is one of the best things you could do for humanity long term in my opinion.
he didn't? stan ulam credits 'singularity' to von neumann
vinge wrote his singularity piece in the 80s I believe
dwohnitmok 8 hours ago [-]
Vinge addresses this in the linked article.
> Stan Ulam [28] paraphrased John von Neumann as saying:
>> One conversation centered on the ever accelerating progress of technology and changes in the mode of human life, which gives the appearance of approaching some essential singularity in the history of the race beyond which human affairs, as we know them, could not continue.
> Von Neumann even uses the term singularity, though it appears he is thinking of normal progress, not the creation of superhuman intellect. (For me, the superhumanity is the essence of the Singularity. Without that we would get a glut of technical riches, never properly absorbed (see [25]).)
which hews much closer to how people who are fans of the term use it (the Singularity explicitly refers to the rise of superhumanly intelligent systems, not just the progress of technology overall).
orthoxerox 13 hours ago [-]
Were they the villains? I remember the rogue three-letter-agency executive being the BBEG.
59percentmore 10 hours ago [-]
Part of the (relatively) happy denouement included the Chinese coming in with the nondestructive process and everyone being quite happy with that. So yes, the Shredders were the villains.
vessenes 13 hours ago [-]
I think some villainy is implied by "billionaire-backed-woodchipping of a library," but it's just my interpretation, no actual knowledge of Vinge's perspective.
clickety_clack 13 hours ago [-]
That’s digital though, so it requires the continued survival of readers for the data that is stored. The best thing you could for the long term is probably to buy a few hundred physical books to keep in a bookcase in your home.
trollbridge 12 hours ago [-]
Speaking as someone with dozens of bookshelves and tens of thousands of books... I kind of prefer the idea that continued survival means getting a bunch of 14TB drives and, you know, hosting certain files obtained from certain places. The reality is that most people's book collections are simply going into dumpsters, speaking as one of the people who go and try to buy these book collections at estate sales. (We can't do anything with the sheer volume of these books so most of them go into the dumpster. I cannot store hundreds of thousands of books, or millions, and no libraries want them.)
Also, holy cow, hard disks (as in the magnetic oxide kind) got a lot more expensive.
vessenes 12 hours ago [-]
Yeah agreed. I have on my long-term project list a hardware 'oracle' that would have everything and a local good model as a librarian/assistant and be solar powered in a pinch.
cyberrock 12 hours ago [-]
Well it's close to the author's intention for Fahrenheit 451, but just not what everyone wants it to mean.
JumpCrisscross 12 hours ago [-]
Maybe we can kill two birds with one stone: digitize rare books and reverse the damage from Authors Guild v. Google [1].
Let AI companies do this. But require them to make the digital copies public. Maybe with a multi-year delay, to give the original scanner advantage to doing it.
Let's say one of the books to be digitized and destroyed is the sole remaining copy of a book from 1850, which is now considered public domain.
On one hand, hoarding such a book, stealing its content from the public domain, locking its content behind a for-profit machine, and destroying the only remaining copy is clearly wrong. It's equivalent to stealing a public resource, just like mining minerals or oil on public lands without a permit or mineral rights. Pure extraction.
On the other hand, taking care to digitize the copy and making it available for free in perpetuity, as well as being required through regulation to provide access to that content through, let's say a public utility LLM/AI available for free through libraries and online... and perhaps after fair due diligence being required to preserve physical copies in a public archive of rare books of which there are no known remaining physical copies...
That seems much more reasonable to me at least. I can imagine there are many who would not see it that way though. Do we see it happening or gaining regulatory, moral and/or public support?
s1artibartfast 11 hours ago [-]
I think this misconstrus what public domain is.
It provides a freedom to circulate, but not access to the material. It is not a public owned resource.
Turning a copy over to the public or state might be an interesting requirement for obtaining a copyright, but instituting that fix for new works now would have a 70 year lag time.
Think of it this way, if I copyright a book and put it in my dresser for 70 years, that doesn't give the public the right to access it or come into my house and scan it after expiry
Legend2440 10 hours ago [-]
>Turning a copy over to the public or state might be an interesting requirement for obtaining a copyright
Oh, that is fascinating! I stand corrected. Then the question is really about non copyrighted works
JumpCrisscross 10 hours ago [-]
> It provides a freedom to circulate, but not access to the material. It is not a public owned resource
It increasingly looks like a grand compromise around copyright and AI is needed. Expanding public domain when it comes to AI companies in this respect seems merited. (The other bits are fair use if weights are opened.)
sixdimensional 10 hours ago [-]
I think you are right about the narrow legal distinction, although I was not suggesting that the public has a right to enter someone’s home and scan a book merely because its copyright has expired.
A privately owned physical book and the public-domain work embodied in it are different things. Keeping an old book in a dresser, even without providing access to it, is also meaningfully different from deliberately acquiring the sole remaining copy in order to extract its content, destroy the artifact, and preserve exclusive commercial control over the only remaining usable version.
Destroying the sole surviving copy would not literally steal publicly owned property. But it could irreversibly remove a public-domain work from practical human access and destroy a unique piece of cultural heritage.
In that very narrow situation, preservation, archival deposit, digitization, or public-access requirements may be justified, especially when the material is acquired for commercial use that depends on destroying or withholding the only surviving source.
The company would not be reasserting copyright in the legal sense. It would, however, be creating copyright-like control over practical access.
The work would remain legally free for everyone to reproduce, while the company’s conduct made it impossible for anyone else to obtain it.
I know there are more pressing issues than this one in the world, particularly those involving immediate human life.
But I also believe that the value we assign to human life is connected to the value we assign to human knowledge, memory, invention, and culture. When those things are casually treated as disposable inputs for extraction, something important erodes.
Companies should be free to use public-domain material commercially. The issue arises when that use creates a negative externality by permanently destroying the public’s future opportunity to access and use the work.
Mining our cultural heritage and locking away what remains is still extraction and exclusivity.
Cider9986 11 hours ago [-]
Can only blame AI companies to a limited extent. This is apparently the legal way to do things because of stupid copyright laws.
It seems a key contention of theirs is the possibility that rare books are being destroyed this way, yet the things they cite don't seem to suggest this (based on their paraphrasing), they just throw the following at the end to make it seem like it's occurring to irreplaceable books:
> You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal.
Is there evidence of this? Since otherwise they could very well be describing what is only occurring to in-print or non-rare books. (This is a genuine question since their post doesn't shed any light on it.)
D13Fd 13 hours ago [-]
There is zero reason to shred 18th century books. Any such books are out of copyright.
samastur 12 hours ago [-]
They are, but using the same process for all books is simpler and cheaper.
qingcharles 8 hours ago [-]
Though they're not shredded simply due to copyright, they're also shredded due to cost, speed and the quality of the scans.
RIMR 9 hours ago [-]
Yes, that is what the person you are responding to is saying. That's why they are questioning the assumption that this practice extends to 18th-century books that are solidly in the public domain.
If it is happening, it is an outrage. However, the 404 article doesn't actually provide any evidence of this; it just connects the shredding of digitized books and the digitizing of rare books to the assumed shredding of rare books, which isn't necessarily happening.
4ndrewl 8 hours ago [-]
Just sociopaths doing sociopath things. Bugger civilization.
pu_pe 13 hours ago [-]
From what I understand, the rare books in question are not some historically relevant medieval manuscripts, but rather some relatively recent books (still under copyright) for which there are few print copies available for purchase.
I am not sure that physically destructing one copy of this type of book to preserve its contents digitally is so bad. Pretty much anything that is still under copyright should be valuable only for its content, not for the physical medium it's printed on.
eru 12 hours ago [-]
> Pretty much anything that is still under copyright should be valuable only for its content, not for the physical medium it's printed on.
Well, there are special collectors editions with the signature of the author and gold pages or whatnot. But the AI companies are probably not using those.
pfdietz 10 hours ago [-]
> historically relevant medieval manuscripts,
These are increasingly available online, btw. Historical research is accelerated when historians have direct access to scans of relevant source material. Not destructively scanned, of course.
qingcharles 7 hours ago [-]
This works only if the digital copy is made publically available. I actually agree with destructively scanning a copy of a work that is in single digits as long as that excellent digital scan is made available.
lejalv 13 hours ago [-]
"To preserve them digitally"
For whom? is the relevant question
godshatter 11 hours ago [-]
Is slightly modifying some weights in a markov chain somewhere really preserving it's contents?
storus 11 hours ago [-]
It's not really Markov chain as you need full P(x_t|x_{t-1}, x_{t-2}... x_1) instead of just P(x_t|x_{t-1}).
pessimizer 10 hours ago [-]
They used to put them all up on books.google.com until they were ordered not to if they were less that 100 years old. Most of the stuff that was accessible was copied to archive.org, and almost all of that stuff ended up on annas-archive and library genesis.
Consequently, many/most books are more available now than they've ever been. This is mostly a copyright question, not a question of preservation. I say mostly, because those scans, without redundancy and fingerprinting, can be changed and bowdlerized in the future without remaining physical copies as a reference.
I own about 3500 print books, and started a project to find scans for all of them (that I need to get back to.) Average publication year is probably around 1975, and the bulk ranging from the 1940s to the 2000s. I made it through about 1500, and couldn't find maybe 40, most of them bad. e.g. self-published stuff like "My God Heals, Does Yours?" This was a few years ago, if I went through those 40 now, I bet I'd find half of them.
I hate what the AI companies are getting away with, but only because they get to violate copyright while being aggressive enforcers of copyright and DRM circumvention laws. Destructive scanning, however, is a cheap way to get a good copy of a book online. If that book were then put in a place where teveryone interested in its contents (or who are just hoarders) can get a hold of it, you'll be able to find a verifiable copy of it 1000 years from now.
As of now, Russia and annas-archive are just a few points of failure that can erase those words forever. It will be done with armed, uniformed men, and people will claim it is not dystopia, but justice. I don't want the only copy of the text of a book to lie in the interpretation of some privately trained LLM.
flipped 13 hours ago [-]
[dead]
aerodexis 12 hours ago [-]
Reminds me of Blood Meridian where The Judge meticulously sketches the rock glyphs that he comes across, and then destroys the original.
Now that I think about it, The Judge is an apt metaphor for AI : "Whatever in creation exists without my knowledge exists without my consent."
karahime 11 hours ago [-]
No, it isn't. Nothing about training a model requires this. You're thinking of copyright, which does treat information this way.
aerodexis 11 hours ago [-]
Sure, I'm conflating the tool w/ the people building the tool - but imo muddling the distinction is good, because the hard distinction is what tricks people into an unthinking mode.
karahime 10 hours ago [-]
No. Again, it's copyright that says "create a symbolic version and destroy the source, so that the thing may enter a controlled regime". Copyright is, in literal terms, a restriction on your right to copy.
TeMPOraL 10 hours ago [-]
On the contrary. The problem exists solely in copyright; AI companies are just users that happen to be popular to hate on, but this same legal situation applies to everyone else, including individuals. The issue under discussion is older than LLMs.
RIMR 9 hours ago [-]
Well, AI and Copyright are both related things that people created, so it's certainly relevant. I don't think the person you're responding to misspoke.
xboxnolifes 11 hours ago [-]
Isn't this the legal requirement for digitizing books usually? I dont like that its being done, but I feel the direction of anger for this one is misplaced. Or at least partially misplaced.
mchusma 11 hours ago [-]
Yes part of the legal defense used to good effect is that they are copying one for one (not really making a copy but rather converting it). I agree with many in this thread that copyright reform is the best solution here. Shorter copyrights. Lose copyright if you stop using it (printing it), more clear digitization and preservation guidelines.
silverlimetea 11 hours ago [-]
[flagged]
999900000999 12 hours ago [-]
They should be forced to publicly release the books as an Ebook.
Leave it up for anyone to download and then compensate the copyright holders later.
In fact if ingesting these books for LLMs is fair use, us commoners should be able to read them for free. Maybe restrict commercial redistribution though.
streetfighter64 12 hours ago [-]
There's a bunch of details regarding LLMs and copyright, but I don't see how
> They should be forced to publicly release the books as an Ebook.
would be reasonable in any way? The books aren't theirs to release publicly. If I brought a copy of any given movie on DVD, ripped it and used it to train my own "LLM" located at /dev/null, should I then be allowed (or even forced) to release the movie publicly for anyone to watch for free?
999900000999 11 hours ago [-]
We’re in a brave new world now. If theirs reason to believe your destroying the last copy of a book( or another copy isn’t easy to obtain) then you should make a copy of it available.
Set up a compensation fund for the rights holders. Anything is better than culture literally being sucked into the void.
11 hours ago [-]
cj 11 hours ago [-]
The alternative (throwing books away for no good reason) isn't much more reasonable.
pfdietz 10 hours ago [-]
It happens on a massive scale, so it's completely reasonable. Books are disposable information delivery vehicles now. Publishers pulp great quantities of unsold books too.
thechao 13 hours ago [-]
Which book that was rare was destroyed? I'm interested to know a few titles.
> The attachment contained 3,000 English-language titles organized by ISBN number, including books such as Distinct Element Modelling in Geomechanics by K.R. Saxena (1999); Barrett's Traditional Fairy Tales (2021), an academic study of Irish folklore; and Laser Shock Peening of Advanced Ceramics by Pratik Shukla (2018).
infinite_spin 13 hours ago [-]
> Barrett's Traditional Fairy Tales (2021)
How is a book from 2021 considered rare in this context? There's almost certainly a digital copy of it in existence prior to Anthropic purchasing a print edition.
ACCount37 13 hours ago [-]
Niche text. It's not impossible that there was only ever under a thousand of them printed and released into circulation.
A digital copy would exist somewhere, of course. But for us, that only matters if we can buy or download it. And for AI companies, that only matters if they can get a digital copy DRM-free and licensed permissively enough.
sidewndr46 12 hours ago [-]
This argument doesn't make any sense. All manner of AI companies just ingest whatever random text they can find on the internet to train their data, including copyrighted publications. Why would DRM on a digital copy of a book matter?
eru 12 hours ago [-]
There might be special rules around DRM that go beyond normal copyright?
jdub 12 hours ago [-]
Ladies and gentlemen, the Digital Millennium Copyright Act
(which is terrible, but I would be delighted if they breached it and got thoroughly spanked)
ACCount37 9 hours ago [-]
I would be delighted if AI companies got together and thoroughly dismantled DMCA.
That atrocity of a law was a blight upon digital freedom since the day it came to exist. DRM should never have been given any legal protection - and I would push for numerous forms of DRM to be outlawed instead.
sidewndr46 11 hours ago [-]
so if I base64 encode my blog, have some Javascript that 'validates' an authorized viewer and then decodes the base64 into HTML which is added to the DOM does that constitute DRM ?
pessimizer 9 hours ago [-]
Yes.
goldlimetea 9 hours ago [-]
I would say no because you need to retain the control over the IP from your end.
Typically DRMs are implemented server-side so you can control user access in that way.
efreak 2 hours ago [-]
The DMCA doesn't care exactly how the authorization works, it only cares that it's needed.
Quote from Wikipedia[0] of DMCA section 103:
> No person shall circumvent a technological measure that effectively controls access to a work protected under this title.
> "circumvent a technological measure" means to descramble a scrambled work, to decrypt an encrypted work, or otherwise to avoid, bypass, remove, deactivate, or impair a technological measure, without the authority of the copyright owner; and
> a technological measure "effectively controls access to a work" if the measure, in the ordinary course of its operation, requires the application of information, or a process or a treatment, with the authority of the copyright owner, to gain access to the work.
> Why would DRM on a digital copy of a book matter?
Because DRM is just a way to make "breaking copyright" more practically cumbersome. What's easier, breaking digital DRM for each and every E-book you find, or just establishing a single pipeline for scanning physical books?
sidewndr46 11 hours ago [-]
breaking DRM is so easy my generation was doing it as kids, there is no technical obstacle there.
TeMPOraL 10 hours ago [-]
It's a manual process.
More importantly, it's also explicitly illegal. Destructive format-shifting is not. Thank copyright laws.
tokai 12 hours ago [-]
Non of those are rare. All a available in libraries for ILL.
timmmmmmay 8 hours ago [-]
so... not rare, then? another fear mongering article lying to everyone.
13 hours ago [-]
andrepd 12 hours ago [-]
Rare or not, destroying books is in itself a morally repugnant act. I don't know man, it's not so long ago we used to view nazi book burnings as an archetype of evil. Today companies are offering book-burning-as-a-service and it hardly causes a stir.
Bearded German man was right.
JodieBenitez 10 hours ago [-]
> destroying books is in itself a morally repugnant act
It's a routine act in any publisher stock management.
Gander5739 11 hours ago [-]
Well, there's a difference between destroying books to erase knowledge and control people, and destroying books because it's cheaper.
skeledrew 8 hours ago [-]
Is it destruction or format shift? Technically the books still exist; just that they're now digital. There's no content loss (what I imagine comprises a "book") unless any digitization is then deleted.
robertoandred 11 hours ago [-]
Easy claim to make when you’re not responsible for storing those books.
SwtCyber 10 hours ago [-]
[dead]
_m_p 12 hours ago [-]
Librarians were already doing this at scale in a process euphemistically called "weeding":
Weeding is the natural process of disposing of less-demand books. Like the rest of us, libraries operate in finite space, so if they want new books, they have to remove ones their users aren't using. Most libraries will try to sell books before disposing of them in any destructive way.
What similarities do you see here?
Jweb_Guru 12 hours ago [-]
People are very desperate to try to claim that something that's somewhat obviously morally wrong is actually highly nuanced, because it makes them feel uncomfortable.
pfdietz 10 hours ago [-]
No, we're just telling you your moral intuition is absurd.
Jweb_Guru 10 hours ago [-]
Yes, clearly it's completely absurd to think that shredding rare books to more cheaply train LLMs is anything but a moral good, which is why there is this entire thread is full of people jumping through hoops to explain how it's technically legal (and therefore fine) and "you wouldn't have bought those books anyway" (I suppose we don't have the choice anymore!) and, most amusingly, "they're actually becoming digitized and searchable this way" (are LLMs stochastic parrots or aren't they?).
03284782470 11 hours ago [-]
People are very desperate to try to claim that something that's highly nuanced is actually obviously morally wrong, because it makes them feel self-satisfied.
streetfighter64 11 hours ago [-]
I would disagree on the "obviously morally wrong". What's the alternative for the books if they were "saved" from the fate of being scanned and shredded? How many rare books have you personally brought in the last year? Not everything that's ever created needs to be preserved forever.
9 hours ago [-]
emddudley 12 hours ago [-]
Libraries are not archives, and weeding does not mean destruction.
dingaling 11 hours ago [-]
Libraries in my area dispose of 'weeded' books by covert means, so that public ire isn't aroused by finding dumpsters full of discarded books.
They are most assuredly destroyed.
malfist 12 hours ago [-]
This is really dishonest framing, unless you really, honestly can't tell a difference between pulping a mass market paperback romance novel that there's 3 million of in circulation, and shredding an 18th century botanical text that there's only 2 copies of in existence.
dwedge 12 hours ago [-]
You call of dishonest framing, but you're begging the question twice.
Once that AI companies are really shredding 200 year old rare books, and once that libraries are only weeding mass market pulp fiction.
malfist 11 hours ago [-]
That is literally what this article is about.
streetfighter64 11 hours ago [-]
The "article" in question is a tweet. The "rare botanical text" is just an example the author of the tweet made up in order to generate sympathy.
malfist 11 hours ago [-]
The tweet is about a 404 Media article about the practice. It's linked elsewhere here.
dwedge 10 hours ago [-]
Does the 404media article mention that rare botanical book? I saw an excerpt here elsewhere of their examples and it wasn't in that list.
Also I had a family member work in libraries. They got rid of books for very small prices based on how many people checked them out. Not how easy they were to replace. One example was a 100 year old gold leafed book they sold for £5 to someone who removed every page to sell separately.
opem 7 hours ago [-]
Its ironic how the whole systems reacts when someone like open library tries to actually do digital preservation
D13Fd 13 hours ago [-]
This is the result of our copyright law in the United States, which is extremely tilted to favor authors and publishers. The judge made exactly the right call and the companies are following the law.
The fix here is to change the law to permit training AI without destroying the original materials. But that is going to be a heavy lift.
1970-01-01 11 hours ago [-]
>And the judge said it's legal.
Was was the alternative? Order them to stop quietly destroying their property, it is making someone else very upset?
altcognito 13 hours ago [-]
Is there any proof of this at all beyond this random message?
Ratelman 12 hours ago [-]
Yeah - it does feel a bit overly dramatic, books mentioned from a potentially related article are from like 2018 (https://nltimes.nl/2026/06/25/rare-book-dealers-fear-tech-fi...). Let's not pump the drama more than we need to, Sam Altman made another mention of the singularity over the weekend so enough of that going around.
From 1984: "Every record has been destroyed or falsified, every book rewritten, every picture has been repainted, every statue and street building has been renamed, every date has been altered. And the process is continuing day by day and minute by minute. History has stopped. Nothing exists except an endless present in which the Party is always right."
It’s unclear whether ISBNdb will scan books without ISBN’s, which were invented in the late 1960’s. Customers appear to be ordering books to be scanned by ISBN? Here is one book seller’s experience:
> Bulk purchases also usually reflect interest in a specific topic, whereas the recent, very large purchases were of books that had little in common, except for the fact that they all had ISBNs. This seller also sells rare books that do not have ISBNs, and none of those were part of the bulk purchases.
Article is paywalled, but I saved a few quotes here:
I had some old computer/unix/etc books for which I couldn't find digital copies. I was considering paying to have them scanned because lugging them around each time I moved was getting to be a pain in the butt. When I found out they destroyed the book in the process I could just never go through with it. I later found out there are non-destructive scanning machines (even open source ones!) but ended up selling the books before I ever went that route.
TSiege 10 hours ago [-]
How much of this shredding them isn’t just copyright but rather they don’t want anyone else having this information in their datasets?
Horrendous stewardship of humanities collective knowledge all for profit and the race to have the one god computer to rule them all.
As more time passes it becomes clearer that America’s AI strategy should’ve been a public private partnership where the public owned the datasets and the underlying models and we’d leave the productionizing of LLMs to private businesses
justthehuman 10 hours ago [-]
Honestly, I think its just the fastest way to scan them (there are slower non destructive methods available too!)
Also, they can then just recycle/dispose of the paper and don't have to worry about reselling/donating the books themselves. I suspect this is all about speed of data ingestion and anything else is a side effect they don't care about.
bronlund 10 hours ago [-]
You can’t champion on the greed that is copyright and then be sour when Anthropic tries to work within this framework.
There is a word for that kind of behavior.
musha68k 13 hours ago [-]
At the very least why not upload the scanned books to the internet archive while already at it?
This is highly disturbing news; is this standard practice? What did Google Books do before?
qsera 12 hours ago [-]
Yes, this is very disturbing to me. I recently discovered how great old, out of circulation/print can be. Now old abandoned libraries are like a treasure trove to me. If these books are not digitaly preserved, that sounds very sad to me.
> One professional bookseller who specializes in selling foreign language books on these marketplaces told me that, starting in April, he and other booksellers noticed a historic spike in sales. [...]
> This bookseller said his inventory is full of rare, foreign language, and low circulation books, meaning that if they are destroyed in the process of becoming training data, they’ll be even harder to obtain.
azan_ 13 hours ago [-]
The comments there are absolutely unhinged. There are some good reasons for being anti-AI, but why dilute it with this kind of bullshit:
> It is equivalent to book burning in the past. A form of thought control
nhinck3 12 hours ago [-]
You're right, it is arguably worse than book burning, not only are they seeking to deprive others of the books, they also want to profit off it.
frozenseven 9 hours ago [-]
>they seeking to deprive others of the books
Yeah, this isn't a thing. Throwing out books, even "rare" ones, is a regular procedure. And if you really want to read 'em, you can buy any of these books right now.
swed420 13 hours ago [-]
> why dilute it with this kind of bullshit:
> > It is equivalent to book burning in the past. A form of thought control
That's only bullshit if you trust AI companies to serve the book contents without alteration.
dpark 12 hours ago [-]
I don’t trust or expect AI companies to serve books at all. That’s not what they are scanning them for.
swed420 12 hours ago [-]
Not in a traditional sense, but obviously on the surface, they're using the info to regurgitate in some fashion and serve back.
The point is that even under the best intentions, hallucinations occur. Then there's the fact that most models have an ideological bias programmed into them.
The only expectation I have is for companies or anybody else to not destroy rare books. Is that such a tall order?
dpark 12 hours ago [-]
> they're using the info to regurgitate in some fashion and serve back.
Sure, in the same sense that they regurgitate any other text they consume. LLMs by definition do not have the full training dataset available, though. It’s far larger than the resulting model. So they can’t reliably reproduce full text without an external source (or if it’s in the training data repeatedly). ChatGPT actually refused to give me a bible quote the other day, presumably because I ran into some general “book regurgitation” safety net.
> The only expectation I have is for companies or anybody else to not destroy rare books. Is that such a tall order?
Honestly, yeah. The idea that people or corporations should hold onto books forever because of a cultural “ick” about throwing out books is a bit ridiculous. Most books end up in landfills.
They aren’t feeding Da Vinci manuscripts into this pipeline. They are feeding still-in-copyright books.
pfdietz 9 hours ago [-]
> It’s far larger than the resulting model.
Is it? How many different books are we talking about, and how much information is that, after conversion to text and lossless compression? Images, maybe, but text?
dpark 8 hours ago [-]
These models are trained on way more than just books. GPT-3 was trained on about half a terabyte of filtered plaintext and the training corpuses have grown significantly by then by all accounts.
pfdietz 8 hours ago [-]
I imagine that compresses by ~90%, and current top commercial models have a couple of trillion parameters, don't they?
dpark 8 hours ago [-]
They aren’t trained on compressed plaintext so I’m not sure of the relevance there. But regardless it’s my understanding that’s modern models are trained with orders of magnitude more storage than their parameters require. But it’s possible I’m incorrect. This is getting to the fringe of my knowledge of concrete LLM details.
pfdietz 7 hours ago [-]
The relevance is because the LLMs are storing information, not the explicit text, so we want to know how much actual information they need to store (this being an information theoretic argument). The representation in the parameters doesn't necessarily need 1 parameter per character, if the text is highly redundant.
swed420 11 hours ago [-]
> LLMs by definition do not have the full training dataset available, though.
That makes it even worse, then. This proves the original point.
> The idea that people or corporations should hold onto books forever because of a cultural “ick” about throwing out books is a bit ridiculous.
If we're building black and white straw man arguments, then sure, let's not archive anything.
dpark 10 hours ago [-]
> That makes it even worse, then. This proves the original point.
I don’t know what the “original point” is here, but these AI companies are not providing “book excerpt services” and do not claim to. ChatGPT at least will refuse to provide detailed book excerpts (I hit a week or two ago myself).
> If we're building black and white straw man arguments, then sure, let's not archive anything.
It seems like you are the one creating the straw man. Do you have evidence that these companies are shredding actually rare books? The only cited concrete examples (in this thread anyway) are all rather boring. I seriously doubt they are shredding 200 year old books because why would they?
tasuki 12 hours ago [-]
Maybe we should not have forced AI companies to shred rare books so they can use them for training?
13 hours ago [-]
graemep 12 hours ago [-]
Is this true? Rare books would very often be out of copyright for a start. What the the actual ruling that says you can scan if you destroy the original? ISBNdb is a database of book information, as you would guess from the name.
roywiggins 12 hours ago [-]
It's cheaper to destructively scan.
dpedu 10 hours ago [-]
Digitizing books, even if it means the destruction of the original, means more people end up having access to the knowledge within, and is a good thing. Full stop.
xgulfie 10 hours ago [-]
They are digitized for private training and not for public use. Not a single human is reading these.
Cider9986 11 hours ago [-]
It's not like these books were available to everyone before. If they are destroying one physical copy that's not accessible to the public and replacing it with a scanned copy that isn't accessible to the public, that's not a huge change. Except that will probably last longer digitized. Ideally they wouldn't destroy the physical copies, but this isn't anywhere as bad as book burning in Nazi Germany.
I find going after shadow libraries to be much worse because law enforcement is trying to prevent discrimination of knowledge to the public.
The real blame here should be going onto copyright laws.
Anthropic could take more care by figuring out if the books are still affected by copyright.
But this is just a company trying its best in an unfortunate regulatory environment.
Could someone explain to me, why rare books are so precious for them?
Say i feed the largest LLM a book of an alien civilisation, that it definitely hasn't seen before. Then this tiny piece of text muds the vast ocean (latent space) of the model minimally. It will not be able to cite from that book reliably after that fine tuning. Especially for rare books, because them being rare implies, that there aren't 1000s of other books, that encode the same information.
It's general language modelling capabilities might get an iota better, of course. But for putting factual information into it, wouldn't RAG be a much more solid approach?
DarkIye 13 hours ago [-]
This is the opposite of a book burning. These books which only a few would ever know the names of, let alone find, let alone read, are being digitised so they can be found in electronic searches.
dpark 12 hours ago [-]
> being digitised so they can be found in electronic searches
You make it sound like they are running a second Project Gutenberg. They most definitely are not making these available for electronic searches. At least not searches the public can participate in.
coffeefirst 12 hours ago [-]
Correct. If they were doing this with Gutenberg/Smithsonian/some library so there's both a public archive and training the language model, it wouldn't have the ick factor.
imhoguy 12 hours ago [-]
And knowledge from these books can be "imprinted" into a model to be used by much more people or even survive this planet once sent into space in a probe.
rubylimetea 8 hours ago [-]
> “imprinted”
I would say “reduced”
swed420 13 hours ago [-]
> are being digitised so they can be found in electronic searches
What guarantee do we have that the book contents will be served unfiltered and unaltered?
left-struck 12 hours ago [-]
Are they being digitised? That word implies that the book would have a digital representation of the original, which is not what an LLM is.
Can they be found in electronic searches?
I mean I kinda get what you mean, the knowledge was potentially forgotten and now it might not be, but also the authors who created that knowledge get no credit, no reward.
rubylimetea 8 hours ago [-]
They aren’t “digitizing them” as if they are making them available online.
They are creating embedding vectors and training documents - changing whatever they want - and destroying the original copy so nobody knows what was actually said.
It is just like book burning.
johnxianren 13 hours ago [-]
I have zero proof for this, but just a what if: what if Anthropic's strict anti-China stance actually means the Chinese training corpus is way more valuable than people realize?
qsera 13 hours ago [-]
I will make a robot scanner for books. I will then scan all the books in my state libraries and make a digital copy of them (without destroying them) before these things come for them.
I wish...
59percentmore 10 hours ago [-]
Literally Rainbows End. Vinge was an oracle.
It's a real shame that no one ever got that book in front of Hayao Miyazaki's eyes.
mwigdahl 10 hours ago [-]
Well, an oracle except for that part in the same book where unlimited access to the internet made all children prodigies self-motivated by their love of learning.
59percentmore 3 hours ago [-]
I don't know that that's an accurate description of the circumstances in the books. They were born into a world of ever-present and almost-ubiquitous connectivity, and heavily encouraged to string together modular black boxes to build new things. This didn't make them prodigies, it just kept them from falling behind. They lived in a world where even the physical shape of their schools wasn't stable and reliable.
When you compare to kids learning from home during the pandemic, on Chromebooks and iPads, building TikTok videos with CapCut instead of an actual local NLE application... you can start to hear the echoes.
geephroh 9 hours ago [-]
"Librareome Project"
greenlimetea 9 hours ago [-]
It's amazing what reading a word in that book like sousveillance can do.
But it's as if the more people know the word sousveillance, the word and even the act of surveillance actually loses a little bit of power (its embedding changes, if you will!)
All our fears about the end state of surveillance can now be countered by an end state of sousveillance.
Before I had sousveillance to think of, I could only think of surveillance (when thinking of veillances) - and it was more of a threat then than it is now.
The world is made of language, or as Terence McKenna said made out of words which sounds obviously false at first. But go looking for the inside of an atom and tell me what you find, and think about where the medium of reality actually implements itself.
59percentmore 3 hours ago [-]
What.
greenlimetea 2 hours ago [-]
[dead]
juancn 9 hours ago [-]
I dunno.
It's sad in a romantic kinda way, because of the lost artifact, but the information is what makes the book valuable, not really the medium.
The out of copyright books don't really need to be destroyed anyway for them to be fair use for AI training, and arguable, even if you needed to, you only need one copy per title per company at most.
So it's not a gigantic loss.
adamddev1 9 hours ago [-]
But is the information really scanned and preserved for direct access? Or is it just trained into an LLM so that we can only get fuzzy answers about the information?
rubylimetea 8 hours ago [-]
The gigantic loss is that they are destroying the original copy.
They aren’t verbatim uploading the text 1:1. They are creating vector embeddings and training documents from it, changing whatever they want since it’s in private and protected by NDA, and destroying the original source of information so that nobody knows what was originally recorded.
bwfan123 11 hours ago [-]
There needs to be wider reporting of this if it is true. Where are the journalists ? If this is true, it is scorched earth on the world's knowledge base. And the business models are not proven yet. I can now see why some ai companies paints a dire picture of the future - they are basically annihilating all knowledge sources at the altar of their ai gods. I can also see why there is growing ai-skepticism.
frozenseven 9 hours ago [-]
Because there's nothing to report. Books like this are routinely thrown out. By publishers, libraries, and regular people alike.
pfdietz 10 hours ago [-]
> bulk-buying rare books
Isn't that oxymoronic? If they can be bulk-bought they aren't rare.
computerphage 10 hours ago [-]
Each individual book is rare, the purchase includes many different rare books
pfdietz 8 hours ago [-]
If they were rare, they'd be valuable, and would not be sold as part of a bulk purchase.
Arshad-Talpur 12 hours ago [-]
May be I am old schools but rare books must remain rare and LLMs shouldnt have access to those books
extra-AI 12 hours ago [-]
And this is the beginning of the end for content creators.
Why would I spend hours creating original content if Google can extract it and present the answer directly in an AI Overview? What is the incentive to keep doing the work?
If creators stop producing high-quality original material, the information we get over the next few years will increasingly be based on recycled, low-quality garbage.
Traster 12 hours ago [-]
Take a look at most successful journalism today. It's behind a paywall. You get paid by the people who are interested and value your work.
Can Google steal it and present it in an AI overview? Well kinda. Today Google is doing a trick - they're saying "You can refuse to consent to being fed into the slop machine, but if you do we won't crawl you for Google so you'll get no search traffic. But you're not going to get search traffic anyway! So you might as well opt out of being fed into the slop machine. And companies are starting to do that [1]
It's really interesting, because essentially what it means is Google is turning into a walled garden, but there's nothing growing inside it so they have to continually import new plants to live in their walled garden and they're going to have to pay to do that. So soon Google will be paying news sites for the right to plumb their feed into the slop machine.
It’s even noted they’re looking for text pre 2022 as afterward, it’s tainted by their own shit, they don’t believe in the crap they’re making.
Gander5739 11 hours ago [-]
> they don’t believe in the crap they’re making
A spurious claim; they simply want to avoid model collapse.
secretsatan 10 hours ago [-]
Simply? It’s flawed if you can’t train on new data past 2022 without first checking ai hasn’t tainted it. Dismissing that as simply avoiding model collapse seems to miss the point.
pyrophane 11 hours ago [-]
Edging a little closer to the Krazam video "rare data hunters."
13 hours ago [-]
HelloUsername 11 hours ago [-]
Are the AI companies also destroying the digital scans they made?
gowld 11 hours ago [-]
Almost certainly not, because they need that data for model-building and rebuilding., and the storage cost is completely trivial.
HelloUsername 9 hours ago [-]
So the X post is not entirely true; there still exist copies of the original
paxys 12 hours ago [-]
People keep bringing up these rare books but never share what they actually are. What are their names? When were they written? How many copies were in existence? Were they in libraries, or locked up in vaults? Did people have access to them prior to being scanned for AI? How many such books have actually been destroyed?
Weird to see so many of these "trust me bro" twitter stories make it to the front page and cause outrage when no one has any real information.
mindslight 12 hours ago [-]
The design of copyright has always been to restrict the dissemination of knowledge so that somebody can turn a profit. The handwaved justification is that the profit encourages the creation of new knowledge whose dissemination can be restricted, but that still doesn't erase the fundamental dynamic.
This is merely the latest incarnation. We can imagine a slightly different process on a few fronts - AI companies pay to digitize books (still for their own purposes), but are prevented from destroying the physical copies and they have to openly shared the digitized results. We would view that situation much more favorably - perhaps even as ideal, right?
Those two dynamics could be backed up by court decisions or laws iff they weren't so plainly at odds with how copyright has been and is generally implemented and interpreted. For example, imagine them having to do this through some nonprofit library whose goals was preservation and dissemination. Instead, libraries have been sidelined as things that operate at the edge of the law rather than vital public institutions, whereas shredding books in secret is fully legally condoned.
Cider9986 11 hours ago [-]
This is the correct take.
Copyright has done more than anything else to prevent preservation and dissemination of knowledge. And it's forcing Anthropic's hand now. Although they could take more effort to preserve the books.
mindslight 11 hours ago [-]
I'm not going to absolve Anthropic here (even though I am a satisfied customer). They could probably get a court decision that it's fair use to non-destructively scan books and keep the physical books, legal inability to directly distribute the results notwithstanding. And they certainly have enough money to pay for some old salt mines or whatever, or fund a library-type institution to do so.
What I'm indicting is the copyright regime being primarily focused on control and the prevention of dissemination. We can imagine a different world in which the publishers' suit against the Internet Archive went the other way (or was not even brought), and a public interest group sues Anthropic (et al) for destroying cultural commons, and gets a judgement saying all scanning must be done non-destructively and made available through institutions like the Internet Archive.
storus 11 hours ago [-]
That's one way to pull the ladder if regulatory capture fails...
cormacrelf 9 hours ago [-]
This story was reported by 404 media. OP linked some random person’s tweet.
If true this is basically unforgivable. The wanton destruction of history is the stuff of the Third Reich and the Taliban. You simply cannot profess to care about knowledge, culture, or civilization while destroying the physical manifestations of the same.
pfdietz 12 hours ago [-]
Yes, actually you can. The fetish of physical book worship is a fossil of an age when information storage and retrieval was much harder.
BoxOfRain 12 hours ago [-]
I strongly disagree with this take, a book might remain readable thousands of years from now but very few if not zero of our digital data formats likely will. We shouldn't be so quick to throw away diversity in the way information which may be useful for future generations is stored for the long term.
pfdietz 11 hours ago [-]
Paper doesn't easily survive for thousands of years.
You know that pleasant used bookstore smell? It's paper slowly decomposing.
BoxOfRain 10 hours ago [-]
I'd be willing to bet that there's a lot more readable paper from thousands of years ago than there will be readable digital data from today in thousands of years. Digital information has to be actively preserved in every case, whereas paper can be passively preserved in some cases.
pfdietz 9 hours ago [-]
Texts from thousands of years ago survive only because they were repeatedly copied. There are some extreme examples like the Dead Sea Scrolls but they are the exception.
Modern acid-free paper might last 1000 years; 500 is more typical. Acid paper breaks down in less than a century.
Parchment was so expensive it was often scraped and reused; old texts can sometimes be recovered after being overwritten (palimpsests).
petesergeant 12 hours ago [-]
So the premise here is that the books are valueless enough that they're being sold by weight, but they're also rare enough that someone will one day wonder where they went, and also that the AI companies aren't also saving the text somewhere much safer than physical media in a warehouse. k.
freejazz 12 hours ago [-]
Why would you need to shred a book from the 1800s when it is in the public domain? Something is fishy here.
dclaw 9 hours ago [-]
It's just modern book burning. The contents don't even matter, you will never see the text and images in these books again.
nh23423fefe 8 hours ago [-]
Pearl clutching nonsense
rubylimetea 8 hours ago [-]
I agree, this is bad.
They’re privately putting info into their LLMs, changing whatever they want, sorry, “sanitizing” then destroying the original.
regnull 12 hours ago [-]
The source for this is "I've heard it from some guy". Before we freak out, perhaps make sure it's actually happening?
bryan_w 12 hours ago [-]
This is quite a reasonable take.
iamsaitam 12 hours ago [-]
If you consider that AI companies should pay the same for a book, as a person does, you should read about royalties.
hagen8 12 hours ago [-]
Where are the sources for that?
goldlimetea 10 hours ago [-]
The irony in asking for a source here. I guess the idea is that eventually you won't get to ask for a source because they'll all have been destroyed.
And the only real source for anything will be an LLM response.
gowld 11 hours ago [-]
Everyone who refused to buy these books in the past was voting for the books' destruction by default. Authors don't owe you a copy of their work. Bookstores don't you real estate to hold a book you don't want and will never want.
m4rtink 10 hours ago [-]
Destoying books, for any reason, is a crime and it should not be done, ever!
tokai 12 hours ago [-]
I don't buy this at all. Looking up some of the mentioned rare books, in other sources, on worldcat shows every single book available somewhere. Most of them in dozen or even hundreds of libraries.
you wouldn't believe how much shredding your local library does in the name of space conservation and to address changing borrower preferences and demographics. I doubt any AI company comes close to the annual combined library turnover
pfdietz 12 hours ago [-]
Where I live there's a large library-associated biannual book sale where books are (effectively) reverse auctioned over a period of a few weeks. At the end, anything (with a few exceptions, like the Collector's Corner) that isn't sold is disposed of, with a large 18-wheel truck sized dumpster filled with items to be sent for pulping and recycling. The price at the end is $1 for a grocery bag full of books, so things that don't sell truly are perceived as worthless.
Books are information delivery vehicles. We mostly shouldn't care about them any more than we care about a particular set of bits on a disk.
Publishers also pulp large numbers of books themselves. This is a consequence of the Supreme Court's Thor Power Tools ruling, which clarified tax rules in the US so that keeping large inventories of unsold books was less economical.
skeledrew 8 hours ago [-]
Another consequence of the existence of that thing called "copyright".
13 hours ago [-]
ertucetin 10 hours ago [-]
They are playing dirty.
RIMR 9 hours ago [-]
I feel outraged, but I also worry this article isn't necessarily responding to what's actually happening.
It's a weird practice to destroy a book when you digitize it, but it's at least an understandable legal strategy to ensure that the digital copy "replaces" the physical one. However, this is only going to apply to books that have active copyrights.
This author suggests that a rare 18th-century botanical text could fall victim to the same fate, but I am somehow doubtful that this is the case. Non-destructive scanning is trivial, and these kinds of books are likely being processed in a quantity that would allow for it without backing up the pipeline.
I really like 404 media, but it doesn't really seem like the evidence points to the conclusion here. Yes, AI companies are shredding books that they digitize, and yes, AI companies are digitizing old, rare books. But the rationale for the book shredding doesn't exist for the old, rare books, so I would need more evidence than just "putting two-and-two together".
At the end of the day, old books with no resale value, including rare books, end up destroyed with some regularity by libraries and bookstores. While this may be an excessively generous take, at least this way the books are getting digitized before they become pulp. The real tragedy will be if the old, rare books that were digitized are never shared with the rest of us, because they were ONLY digitized to train AI, and not to actually preserve anything.
bigbuppo 10 hours ago [-]
I can't help but make a comparison to one of those road planning memes.
Just one more book, bro, and we'll solve AGI forever. Trust us, bro, just one more book. Come on, let me have those words and we'll solve AGI forever.
> A federal judge ruled the practice is fair use because eliminating the original means only one copy exists at a time.
Really? That sure wasn't a thing when one startup got sued for streaming from its wall od dvd-players, and they adhered to 1 disc = maximum 1 stream at same time.
fmaccomber 13 hours ago [-]
That's not an accurate characterization of the ruling
impsunrise 12 hours ago [-]
ironically, this tweet and account appear to be complete AI slop
if they do this you can foresee what else they can do
azan_ 13 hours ago [-]
What else they can do based on this?
qsera 13 hours ago [-]
If they are destroying old books, then it shows where their values lie...
azan_ 12 hours ago [-]
Where? Could you stop with the vague posting?
qsera 11 hours ago [-]
[flagged]
azan_ 11 hours ago [-]
I genuinely want to know what exact point you are trying to make, I'm not a mind reader.
qsera 10 hours ago [-]
This entity is destroying books and make them only accessible via the "intelligence" exposed by their LLMs. So I ll give you two options...
* their values are aligned with the best interests of humanity
* their values are not aligned with the best interests of humanity.
Now go ahead and read my mind!
pfdietz 10 hours ago [-]
> make them only accessible
But that's a lie. They're not destroying the last copy of books.
qsera 10 hours ago [-]
TFA says so...
pfdietz 8 hours ago [-]
I don't see how that makes sense. The last copy of a book would be valuable and would not go into a bulk buy.
03284782470 11 hours ago [-]
Yeah, didn't expect you retard to present any argument. qed
qsera 10 hours ago [-]
>retard
Oh, I have a reputation here. Thanks for letting me know.
api 12 hours ago [-]
People would have at least somewhat less of a problem with this if they also put up an archive of PDFs of all these rare books if they are out of copyright.
But that would help competitors with training data, which I assume is why they don’t do this.
emsign 11 hours ago [-]
That's exactly why I have been hoarding physical media for years now. I knew they would be buying them up, besides the usual dumpster fire they would end up in when people throw them away. CDs, DVDs, BluRays, books, vinyl, everything. And as they are physical they are not prone to retro-editing. Just like in 1984.
Luckily I don't live in NYC where book hoarders are being evicted because of "fire hazard". Just like in Fahrenheit 451.
qsera 12 hours ago [-]
All those AI powers, and they cannot find a way to do it without destroying the books? Pathetic.
psychoslave 10 hours ago [-]
Guys never heard of palimpsests and how your regular scans won’t necessarily catch all data there is in the material?
Of course even disregarding this fact, this is utterly disgusting attack on humanity heritage, just as much as any group out there destroying what we should all cherish be it for the historical artifact they represent. Whatever how US judge name it, they don’t worth more than their same-behavior consorts that is terrorists and totalitarian governments.
They change the info then destroy the source. Now the lie is in the LLM.
bebe8393jrir 13 hours ago [-]
[flagged]
greenlimetea 11 hours ago [-]
This is potentially very bad.
Like, Library of Alexandria or Council of Nicaea bad.
We may never be able to recover the information if, say, one of these AI companies copied or translated it wrong then destroyed the source material.
Maybe it's from bad OCR, or maybe from a bad actor - but there are a lot of ways history and information could change in this game-of-telephone like transfer of knowledge.
What is the point of destroying the source material? I don't buy the copyright thing.
TeMPOraL 10 hours ago [-]
> What is the point of destroying the source material? I don't buy the copyright thing.
It is the copyright thing.
Despite what people say about scanning, the fact is, non-destructive scanning machines have been built and perfected long time ago. This was preferred in the past, back before some major kerfuffle with the publishers during COVID, but that incidentally happened to be before LLMs became a thing, so AI companies never had that option available.
rubyfruit 9 hours ago [-]
[flagged]
1970-01-01 11 hours ago [-]
They're destroyed for 3 reasons:
It is the cheapest way to get them scanned.
It is the fastest way to get them scanned.
It doesn't need to be safely archived for another century until it is resold to someone that has not yet been born.
rubyfruit 9 hours ago [-]
Or they change info then destroy the source.
Is that possible at all or nah?
ianm218 11 hours ago [-]
> What is the point of destroying the source material? I don't buy the copyright thing.
The main reason is the machines they use to scan books at scale destroy the books in the process.
rubyfruit 9 hours ago [-]
How convenient
cluckindan 12 hours ago [-]
This is not that different from having a book burning march. The fervent neocon tech overlords and their funding circles want a monopoly; not only on truth, but on ideas and thought in general.
vessenes 13 hours ago [-]
I don't see any proof of shredding here. Most book scanners I'm aware of are from Google's scanning days, and those had cameras plus page turning.
If we use 'shredding' to mean A book is laid flat, its cover is removed, and then a paper cutter cuts through the binding to create a flat stack of sheets, which are then fed to a sheet feeder, then I could maybe imagine this is better than a page-turning scanner. But, sheet feeding old paper sucks shit, bro. It's not fun.
Upshot, I think we'd like to hear from an anonymous frontier lab employee here to see what's going on -- there are a lot of books in Anna's archive available at considerably less difficulty.
Jolter 12 hours ago [-]
The idea here is that the frontier labs are no longer willing to risk using the likes of Anna’s Archive, because they have already been held legally liable for that in the past.
vessenes 12 hours ago [-]
Are you referring to the META lawsuit? I think the current landscape is not bad for the frontier labs -- it's settled law that it's legal to 'read' and ingest this data. Anthropic went ahead and just settled a licensing deal for content. The open issue in that META suit is whether or not any distribution happened, as I understand it. I'm certain they all have full backups of the archive somewhere in the org.
Jolter 10 hours ago [-]
Yes, that’s the case I’m on about. Training notwithstanding, I’m pretty sure downloading the torrent was a crime to begin with.
krunck 11 hours ago [-]
In a society that values science and knowledge, preservation of knowledge - and no, slurping text up into an LLM is not preservation - is more important than profits.
These companies are regressive book burners.
tekne 11 hours ago [-]
You know, I'm pretty sure post-slurping the text is still there...
03284782470 11 hours ago [-]
Burn them at the stake! Grab your pitchforks!
genxy 11 hours ago [-]
I love it. Literally burning down civilization to generate losses from slop. This reminds me of the Simpsons Halloween episode where Homer gets a taste for his own flesh and eats himself to death.
Maybe the Library of Alexandria didn't burn, it was digested.
The market has solved the what do with excess knowledge problem.
timcobb 13 hours ago [-]
This kind of reads like a blood libel. My guess is they're buying all those books that University libraries are throwing away these days (to see other HM threads for that), kinda sad but probably aren't "rare books" in the way people are thinking
stuartjohnson12 13 hours ago [-]
I think people are, on the whole, too precious about old things. In the case of books produced after major commercial printing began, I don't believe it is the paper that imbues the book with historical value.
Indeed, I think there's a high chance that this process increases preservation of the most relevant part of the media - the actual content!
There are tons of old rolls of film slowly rotting away in warehouses that were never digitised. Even for beloved media, the BBC occasionally tracks down an old lost episode of Dr Who.
For now these books are in corpuses of training data, but eventually I trust they will make their way to the rest of us.
Cynddl 13 hours ago [-]
> For now these books are in corpuses of training data, but eventually I trust they will make their way to the rest of us.
What makes you think they will? What would be the incentives for these companies to do so?
stuartjohnson12 13 hours ago [-]
Well, you probably weren't going to go and find any rare, non-digitised books to physically go and read (unless you were going to, in which case, rock on), so we can start by benchmarking relative probability there.
1. At some level of critical information-withholding mass, a leak or disclosure similar to SciHub is inevitable because of the commonly held opposition to hiding knowledge.
2. Availability via Google Books or similar.
3. Availability via AI model reference.
4. Failing any of the above, better AI models that are more capable of doing more things, at the expense of books that were likely to go unread (revealed preference, rare books are often rare for a reason). This will obviously be a nonstarter if you don't want this to happen, but I think it would be good for the world if it did.
I think category of old books that were going to be read or otherwise become important parts of human knowledge that have not yet been digitised and now will never become so because they are instead being shredded and will never make their way into the light because of AI company data hoarding is a small category.
That's only true if you think digital storage has more longevity than paper. It doesn't.
gowld 11 hours ago [-]
Then why are people so worried about the paper not surviving?
breezybottom 9 hours ago [-]
Because its longevity depends on not being deliberately destroyed.
croes 13 hours ago [-]
So if the Mona Lisa is part of a model we can burn it?
stuartjohnson12 13 hours ago [-]
I pre-empted this - my argument does not apply to texts where the physical object is a major part of the historical value of the thing. No, I'm not saying to destroy one of the four remaining Magna Cartas that were meticulously copied by hand. But even if I was, we're only dealing with texts here that are irrelevant enough to have never been digitised already - we tend to digitise most things of value and so the Mona Lisa and Magna Carta would never have been part of this discussion in the first place.
I am however OK with destroying one of the remaining 50 children's books of which only 300 copies were ever printed in a small town in Ohio in the 1970s as a test run for a failed book which was subsequently never commercialised.
croes 12 hours ago [-]
> I am however OK with destroying one of the remaining 50 children's books of which only 300 copies …
What if the perception of those books change over time and are considered masterpieces later on?
Moby Dick was out of print when
Melville died 1891 and not a huge success until it got a revival in the 1920s
stuartjohnson12 12 hours ago [-]
When was the last time a formerly undigitised book that was printed but fell out of print suddenly became a success after being discovered in the last, say, 50 years? Is once in 100 years the sort of level of frequency we're talking about?
This is the kind of rationalization hoarders use. The inability to get rid of things because it could turn out to be something we want in the future for reasons that we cannot currently describe.
It's a loss avertive instinct that I think is misplaced. Treating every printed book as priceless is intractable. It's not how we treat these books at the moment. Apparently, today we don't even care enough to spend a few hours per book nondestructively scanning them in.
Let's say we, instead of destructively scanning this books, nondestructively scanned them. What would you propose doing with the copies afterwards? Sell them? To who? They're valueless individually for the overwhelming part. Warehouse them? Why? For who? Do you want to go and look through them? Why haven't you done so already? Have you ever shown interest in consuming an undigitised book a single time in your life to date?
It feels like the anti-AI crowd here have to tie themselves in knots here to make the loss minimization work.
infinite_spin 13 hours ago [-]
If you purchased the Mona Lisa (or some rare book), in this hypothetical, you can burn it.
croes 12 hours ago [-]
There is a difference between legal and right.
That’s the whole point because the book shredding is already declared legal.
infinite_spin 11 hours ago [-]
You asked a question about whether you can do something, not whether I found the behavior moral. I'm not interested in moral debates.
Invictus0 13 hours ago [-]
> For instance, a painter may insist on proper attribution of their painting, and in some instances may sue the owner of the physical painting for destroying the painting even if the owner of the painting lawfully owned it.[1]
The Mona Lisa's painter isn't alive, they can't sue, and this act doesn't apply to printed books. You're expanding the hypothetical to say "also you're under a specific set of laws that forbid exactly the thing I questioned".. that's moving the goalpost.
It pisses me off to reflect that they can sit on works until copyright expires, keeping them out of print. There's no real need for any of these so-called rare books to be rare while they're under copyright.
And related to this, the books that are in print are mostly only in print in the shittiest way. I often see well-made books from the 17th or 18th centuries which are still in good nick. It's ridiculous that in the 21st century, publication standards have fallen to the point where for most works a disposable format is the only type available - where no amount of money could buy a truly decent hardback copy.
If we have to have copyright laws, I'd like to see two changes to them.
When a publisher has no incentive to keep an edition in print, it should be available to any other publisher to print, without compensation to the original publisher, and with renegotiated royalties for the author.
And if the publisher keeps a book in print - but only in bestseller-grade materials, bogroll paper that furrows in any humidity and perfect binding that molts its pages a couple of dry seasons later - and if it refuses to print a durable hardback copy with signatures, good paper and decent print - something that will still be readable in several generations' time - any other publisher keen to have a crack at it should be able to, again without any compensation for the original publisher, though perhaps in this case, with matching royalties for the author.
Issue is small streamers have no legal defense.
Edit, adding the channel I am referring to:
https://www.youtube.com/@diggingthegreats/videos
Most Youtube videos do not contain any copyrighted music as it takes too much revenue as you state.
Again, its short term profit at the cost of long-term gain.
https://insidethemagic.net/2026/04/new-report-shows-that-the...
We do not continue to pay for most things once they are created. Unless they are continuous services. Artistic works should not be any different.
> If it's good enough for patents, I don't see why it isn't good enough for copyrights
Because there are fundamentally differ concepts and serve different purposes?
How about lifetime of the author? Or "lifetime or 25 years whichever is longer" so those writing in their later years (or dying young) can pass on the time they didn't get chance to fully use.
> Especially since it makes it easier for corporations to exploit their work without paying them anything.
That ship has sailed. We've seems "big corp" commit mass piracy and get the lightest slap on the wrist, I doubt they'll get less brazen going forwards.
The purposes of both are: "To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries."
Like, this is a made up regime with a specific intent. The fact that we treat copyrights and patents differently is an accident of history. I think we could quite reasonably choose a different period of time (and in fact, have done so several times over the past few hundred years) and still promote progress.
I think it's very reasonable to say that one good idea should not be enough to let you coast your whole life, you should be prodded to cough up 3 good ideas. Further, it reduces corporate power at the other end by allowing individuals to play in coroporate properties after a relatively short time. You could be futzing around with, idk, a copyright free Cars under my proposed regime.
They would generally make most of their money early (in the couple of years following when the content is released). Individual authors would be much more affected, it might take years for your book to become popular. Also imagine if a studio decides to make a movie or tv show just right after your copyright expires, they wouldn't pay the author anything and just have higher profit margins.
> The fact that we treat copyrights and patents differently is an accident of history
Patented inventions and technologies have some sort of direct practical value. Society does not really benefit much if anyone is allowed to created derived works based on any copyrighted content without compensating the author.
> a copyright free Cars under my proposed regime.
I don't think cars are copyrighted unless you want to make an exact copy of it you shouldn't run into any issues.
If I wrote it, I own the copyright on it, why should I or my future family give away something I worked really hard for? Why do only authors must care about public benefits?
if public company means employee owned (instead of publicly traded or state owned) then im all for it. if the founders family are good managers they can easily convince their workers to let them keep running things.
you dont deserve a job or a fortune just because your parents did a lot of hard work before you were an adult. you got to prove yourself and be better than the rest, thats what capitalism is all about right?
And no, I mean every trade secret of the company should become public and anyone should be able to create its products and brands. Basically, do to them what you propose to do to writers and artists - why must their families benefit from their work?
To your first point: inheritance has a concentrating effect on wealth. Concentration of wealth is not a good thing, as it breeds inequality by definition.
If your grandparents gave you a house for free, and you used that enormous advantage to pur your money elsewhere and then own 2, 3 houses, and your child then gets 10, that's 10 houses that people could own themselves, instead of paying the highest rent that your child can get away with.
If you have any experience of poverty, you should understand how viscerally unfair the existence of 'rich kids' seem, and how damaging it is to society.
The reason is to encourage people to create useful writings and make useful discoveries. But we need to balance this encouragement with the benefit we get by making these writings and discoveries available to everyone. By granting a time-bound monopoly the state rewards creators, and by ensuring this time period is not excessive, spreads the benefit among the populace.
Plus, with a time bound benefit, you have to keep on creating, which is good for everyone under this regime.
These aren't, like, inherent rights, they are contingent
That would have a few undesirable consequences... for example, you wouldn't hire a 70 year old writer for your commercial project no matter how brilliant they are.
The complexity of our legal system is in many cases justified. The problems are often the numbers (duration of copyright protection etc.)
This makes sense when you’re thinking of a painting or a book.
Who owns the copyright to Windows or MacOS? A corporation. How do you deal with that?
> you wouldn't hire a 70 year old writer for your commercial project no matter how brilliant they are
Commercial projects are works-for-hire and the copyright is not owned by the person who does the work.
The proposal for the limit of copyright needs to be refined.
In the US, the current rule is:
> For … a work made for hire, the copyright endures for a term of 95 years from the year of its first publication or a term of 120 years from the year of its creation
Not that I'm arguing for 30 per se, just that I don't see what goals of copyright would be advanced more by adding an "or until death" complication.
Somebody decides to make a movie based on your book? You get nothing at all from it... The movie bit would be problematic even for books that were reasonably popular at the time. e.g the Witcher adaption came out almost exactly 20 years after the last book, for GOT it wasn't that far from being the case as well (at least for the initial volumes). Studios would be incentivized just to wait a couple of years to avoid paying anything.
I think it could be reasonably to have a fixed limit if the rights are held by corporations, though.
Justifications for copyright are always built on edge cases, seemingly moral justifications of an empirically immoral practice. Yes, it'd be nice if a single handicapped mother of 2 coild see her children rise out of poverty thanks to her writing talent after 50 years. In practice, this person doesn't exist and building society around that scenario is not a good thing.
These 'what ifs' have the same value and do the same damage as 'who will think of the children' do for human rights.
That's not the same for a work of fiction or a piece of music. Case in point apparently books sales for the Odyssey are massively up - when it was originally written in 7-8 BC :-)
Also most books etc don't make much, if any money - an publisher/author might rely on a the cummulative effect of a number of revenue streams built over time.
Also the effect of exclusivity is different - for patents you are potentially blocking the area of innovation you have patented by your exclusivity.
That's not the same societal effect as somebody not being able to copy mickey mouse.
So they aren't exactly the same - however I'm not proposing a 3000 year copyright :-)
My real point though is that IMO, whatever duration we pick shouldn't depend on the creator. It would tend to undervalue their later creations, treats corporations differently from people in a way that doesn't seem relevant to copyright, and oddly might lead to the untimely demise of creators.
I find your Odyssey example to be relevant. Homer's death means people today can release their own translations or adaptations. I can find a public domain version from 100+ years ago, or a modern translator can profit from their work so that I can see their take. I can watch the Italian 1911 silent film version for free on youtube [0], or pay for Nolan's modern take. The expiry of copyright gives me options.
[0] https://www.youtube.com/watch?v=ZbR97hqfG2o
You could do "life of author or X years, whichever is longer". Or include a period after death.
But you see how the complexities come in.
It is said that the vast vast majority of works don't earn anything significant after a few years in any case, meaning the only possible reason to have long copyrights is so that a very few people can get stinking rich. But those people already got rich, in the first few years.. society does not benefit from them getting richer.
20 years fixed term is my proposal.
Think this through some more.
Authors and artists are still creating after that age.
sure, you don't have to pay royalties, but any other publisher can now publish too
Books in that time were _luxury_ goods. Most people could not afford them. One of the ways that was changed was to introduce cheap, mass produced bindings that were lower quality than the bespoke artisianal bindings done by specialist craftsmen.
You can still get custom bindings done. There exists whole niches on the internet of crafters that will take a production run book and strip its binding and make you extremely high quality and custom bindings and covers.
With how the quality of things seems to have been degrading over the years (either real or just me getting older and experiencing the impermanence of all things) I've been trying to adopt an attitude of "if this practice existed before the industrial revolution, I can _probably_ do it" and it's been really great to learn how things were made before they had to be mass produced as cheaply as possible.
There is a market for these type of books, albeit a very small one.
Apparently it is getting harder to find people who can do that as most schools no longer have book binding as a course you can study.
At the risk of stating the obvious, any poorly made books from then wouldn’t have lasted this long and so you would never see them.
This is purely a response to market demand. Publishers aren’t going to put in the extra expense of binding high quality versions of every book so it can occupy warehouse space while consumers everywhere buy the cheap paperback.
> And if the publisher keeps a book in print - but only in bestseller-grade materials, bogroll paper that furrows in any humidity and perfect binding that molts its pages a couple of dry seasons later - and if it refuses to print a durable hardback copy with signatures, good paper and decent print - something that will still be readable in several generations' time - any other publisher keen to have a crack at it should be able to, again without any compensation for the original publisher, though perhaps in this case, with matching royalties for the author.
This is all based on the idea that there is hidden demand for something, but publishers are choosing to deprive us all of it for reasons. That if we open up the laws, another company will come along and satisfy this hidden market opportunity and associated profits that publishers are declining to take.
The simpler explanation is that these high quality editions aren’t being published because the publishers have the data about demand for them. They know they won’t sell.
If the goal is preservation, laws forcing publishers to print on slightly nicer paper isn’t going to solve the problem. It needs to be a robust digital archive and it needs to exist somewhere other than in unsold warehouse inventory or some book collector’s shelf. You’re trying to solve a problem with last century’s technology.
I’m sure there’s ongoing litigation, and better sources than this, but fair use was determined in June 2025 in a sf federal district court https://www.goodwinlaw.com/en/insights/publications/2025/06/...
Similar conclusion vs meta https://www.jw.com/news/insights-kadrey-meta-bartz-anthropic...
And the more recent $1.5B settlement did not overturn it https://www.reuters.com/world/us-judge-approves-anthropics-1...
In fact, the whole problem of shredding books (destructive format-shifting) was created by copyright laws in the first place, and that in itself is a concession hard won against the IP establishment - and all that way before LLMs became a thing.
> step aside
Isn’t that what I explicitly just did in the comment you replied to?
Fair Use is not an activity that you engage in. Fair Use is not a category with criteria that you meet. Fair Use is not a precedent that paves the way for everything afterwards.
Fair Use is a defense that can be used in court when you’re named in a copyright lawsuit. Fair Use is how you justify your actions before the court finds infringement.
> Notwithstanding the provisions of sections 106 and 106A, the fair use of a copyrighted work... is not an infringement of copyright.
[1] https://en.wikipedia.org/wiki/Google_LLC_v._Oracle_America,_....
[2] https://www.law.cornell.edu/uscode/text/17/107
Now if you look at how fair use is used colloquially, everyone understood what I meant except the autistic pedants.
>You don’t know what “Fair Use” is.
This isn't the opinion of some armchair HN commenter. Actual judges have affirmed this, as other commenters in this thread has pointed out.
e.g. Canada doesn't have fair use, but from wiki on fair dealing in Canada, "According to the Supreme Court of Canada, it is more than a simple defence; it is an integral part of the Copyright Act of Canada, providing balance between the rights of owners and users."
The comment uses the same language as it's parent.
Those books predate the development of wood pulp paper. It isn't the publisher's fault they can't economically print on rag paper anymore.
This doesn't affect publishers.
It affects humanity as a whole by having large private parties hoard books to partly destroy them. The covers and spines can have historical value too. There's not even a reason for these companies to release their scans to the public once the copyright expires.
It prevents proper preservation by archivists and preservationists.
How do you adjudicate that? And wouldn't it just lead to loopholes such as "Ghost Printings" (cf. Ghost Flights https://en.wikipedia.org/wiki/Ghost_flight_(commercial_aviat... ) where the books are technically printed in the required volume but practically unavailable to customers through one method or another. Because, the cost of wastefully printing a few books to warehouse, is less than the potential losses of the IP rights, probably.
But there are worse outcomes.
There is not a shortage of books in print. I don't understand why people get so hung up on a few of them being out of print. It is okay for someone to own something cool and not let anyone see it. The cool thing doesn't suddenly become a societal necessity because it is words written down.
Because the point isn't so much in reading whatever text as if all text is the same. The point is the spreading of knowledge. A single book can contain knowledge not present in any other.
> But I don't see how that desire results in the laws needing to be changed so that you can read everything you want
In order for your "people" (country, etc.) to do better, you want them to be educated. In order for people to understand one another, you want them to be able to see all the same various perspectives there are. It makes perfect sense for laws to aim for these goals. This is why libraries exist.
> It is okay for someone to own something cool and not let anyone see it. The cool thing doesn't suddenly become a societal necessity because it is words written down.
Books aren't simply trinkets, like an item you bought at a gift-shop.
> I understand you want to read the books.
I think you're looking at this too much as what people want for their own individual selves, when it's more of what people want for everyone. It's about what they believe is best for society as a whole. They don't need to want to read a book themselves.
In my own field (at the intersection of linguistics, history and archaeology) I am constantly, multiple times per day, referring to information in publications from e.g the 1960s or 1970s that never got republished later except as a brief citation to that old book. And guess what, nearly all that twentieth-century scholarship is still under copyright, often from the big German or Dutch publishers that enforce their claims fiercely. The academic community has been doing a lot of work to scan our institutional libraries and upload them to the shadow libraries, but this is still all illegal copyright violation.
(It features a software engineer introspecting about the fact that his line of work has caused him to dramatically overvalue recency when evaluating books about... pretty much any other technical field.)
Copyright can be intrinsic, sure, but only to maybe 5-10yrs, far less than the maximum to incentivize registration (and thus incentivizing archiving of culture).
I should be able to freely download every Disney film prior to 2006, compile Killer7 for the fun of it, buy a hardback of the LOTR trilogy from any printer I wish, and do so without any interference from the copyright holders.
However, because we don’t live in a utopia and Sonny Bozo became a politician, I’ll be waving a certain jolly flag for the foreseeable future.
We use a special guillotine type cutter to cut off the binding and then store the pages in a sealed plastic bag which goes in the archives; they're stored there indefinitely in case the book needs rescanned for some reason. We also keep the original, uncompressed copies of the books on magnetic disks.
We also go out of our way to try to find rare books published in 1931, 1932, etc. so they are ready to go once the copyright expires.
And no, no AI company has ever come to us and asked to run training on all of our scanned copies.
Who is we?
> then store the pages in a sealed plastic bag which goes in the archives; they're stored there indefinitely
Is that the best thing for archival storage? Like could things like chemical breakdown increase the humidity in the sealed bag or concentrate corrosive chemical vapors? I was under the impression the best environment was an actively climate-controlled environment.
Most books have some mold by the time they're ~100 years old, it's just not enough to cause a serious problem. Sealed enclosures (wrappings, bags, tubs) are a nightmare situation. Even packing books too tighly on shelves accelerates mold growth to problematic levels.
Archivists recommend standard ambient conditions or a little drier for long term storage. As you've said, too dry and the pages fall apart; permanent damage.
I suppose the fancy silica gels that maintain specific humidities would work in bags.
> I would think this totally dries the pages out and then they just turn to dust, from experience.
They used to sell Boveda two-way humidity control packs that would keep a bag at a constant low-ish 32% humidity, but last I looked those were discontinued (it looks like they've pivoted pretty heavily to marijuana storage and higher-humidity products).
The value of most very old books for AI training is very low. You don’t really want your AI training data to start biasing toward outdated writing styles. Most of the valuable knowledge has been covered again in modern texts in more depth and detail.
There is interesting value in old texts and it’s important to have them archived. It’s less valuable for stirring into the giant pot of AI training data, though.
I have a hard time believing that text valuable to humans would not be valuable to AI.
As for consumer-grade solutions, look for the Fujitsu SV600.
I have done it at home for my books since the mid-00s.
I have books but would have 5x as many if I could not capture them digitally. (When I go to move or go through a purge, some of the books I scanned do go to a used book store.)
So I scan in part to keep my physical book-footprint smaller, but also my scans all get cleaned up and uploaded to archive.org. Mainly I scan young-adult science books from the 50's and 60's (since they were so influential and have all but disappeared except on eBay and the like).
You've trashed the book (and its lifespan) but some books are for using, not keeping.
Yet
Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright?
IDK about the shredding, it's not nice, but it's more a problem with copyright law than AI companies.
Scanning books you own should be legal from a copyright point of view, and not require shredding.
Second, one should think about abandoned property provisions for copyright works published more than 50 years ago and in danger of being forgotten: once challenged, either you as the owner have to prove that the work is preserved for future generations (e.g. in various libraries around the world), or you have to authorize further copies, or you give up copyright on the work.
What happens to the pages after? No one needs them anymore, so they get mulched and recycled.
That would be the dominant scanning method even if copyright wasn't a thing. But then again - if copyright wasn't a thing, there would be much less need to scan any physical media.
The reason why OpenAI can't just go on Amazon, buy a "digital edition" of a 2018 book and use that is that it would violate the license in ten ways, and then the DMCA laws that forbid breaking DRM on top of it.
In the end, digital publishing just isn’t right and will lead to massive gap in our historical records. They require active curation and cannot be preserved simply by resting on a dusty shelf.
I see it as opposite. I can hand someone the complete contents of a public library on a thumb drive. Delivering that to their door is going to be much trickier.
Reading and distribution (pirate like) has been made MUCH easier, with non-physical distribution. I'm always carrying a 1" thick book in my pocket, that I can read wherever I am. I was ecstatic when I switched over to digital. I could read anywhere!!!
When today's algae evolve enough into tomorrow's sentient creatures, they're really only going to need up to the industrial revolution and should probably stop right before that.
Strictly speaking, no one needs the Sistine Chapel or the Pietà etc. It would be a shame if they were mulched and recycled, though.
ChatGPT know them, i'd count that as digitalised why keep the originals?
Because there’d be much less content created in any media to capture in the first place.
Turns out most authors and publishers suck at their respective jobs. Most books produced by authors are things nobody wants to read. If a publisher's job is defined to be finding works the public likes, they suck at it too. Instead they just publish lots of stuff, mostly at a loss. They make their money from the occasional hit. But only way they can make money from it is if they have a monopoly over publishing it for a while - which is exactly what copyright gives them. Looked at in another way, this "publish lots of stuff and see what sticks" is the way our society discovers what new works are popular. And copyright funds it.
If you take the view that copyright is paying for the discovery and publishing of new works, then consumers paying monopoly prices for 70 years or more is a bad idea. By far the majority of works are commercially dead 1 year after being published. The publishers tend to make their money from the remaining 1%. They might last five years. Very, very few last 20. I don't see how copyright lasting beyond 20 years can be justified with anything other than: "because of the good work he did 20 years ago, he deserves to be paid for sitting on his arse for the rest of his life". In reality, he didn't sit on his arse - he donated to a few politicians - but the outcome is almost the same.
An 18th century book would be out of copyright so why would it be illegal to keep the original and scan it?
It's cheaper to scan the books if you do it destructively. Cost. That's why they're shredding irreplaceable texts. Nothing to do with copyright.
https://www.404media.co/ai-companies-are-buying-tons-of-old-...
404 Media published a story about this as well and cites a bookseller who notes that all of the books there have sold have had ISBNs (and are thus from 1967 or later and generally would have active copyright).
”very large purchases were of books that had little in common, except for the fact that they all had ISBNs. This seller also sells rare books that do not have ISBNs, and none of those were part of the bulk purchases”
https://archive.ph/9MQrK
It's not just the big players buying up these archives, either. There are a lot of smaller players, especially in the OCR space, who are buying up huge swathes of works in languages which have much smaller digital footprints, e.g. Arabic.
After that you are diving into forum posts etc to see if you can find anyone who has even mentioned owning a copy or having seen a copy.
I don't know what happens to some works. Supposedly thousands, tens of thousands, or sometimes apparently a million or more copies published and yet not a single copy surfaces for years.
Most rare books are rare because no one cared enough about them. Ie most rare books are rubbish.
1. Why are you here?
2. What is the purpose of this comment?
But during covid archive.org decided to just remove the limit and lend unlimited copies concurrently which started the debacle with the publishers.
IIRC, this was 100% it. Lending one digital version of one physical asset was likely already a violation copyright. Lending UNLIMITED digital versions of one physical copy was DEFINITELY a blatant violation of copyright.
"IA maintains that it delivers each Work “only to one already entitled to view [it]”―i.e., the one person who would be entitled to check out the physical copy of each Work. But this characterization confuses IA’s practices with traditional library lending of print books. IA does not perform the traditional functions of a library; it prepares derivatives of Publishers’ Works and delivers those derivatives to its users in full. That Section 108 allows libraries to make a small number of copies for preservation and replacement purposes does not mean that IA can prepare and distribute derivative works en masse and assert that it is simply performing the traditional functions of a library. 17 U.S.C. § 108; see also, e.g., ReDigi, 910 F.3d at 658 (“We are not free to disregard the terms of the statute merely because the entity performing an unauthorized reproduction makes efforts to nullify its consequences by the counterbalancing destruction of the preexisting phonorecords.”)."
This wasn't a very smart move of them. I get why they did it but they put themselves at a huge legal risk.
Publishers had accepted the prior arrangement before The Archive decided to push it, if not explicitly then implicitly by not suing.
I'm a believer in The Archive's mission, and I wish they had treated the goodwill they'd accumulated as something worth preserving and not a currency to be spent.
It has been stated by many before me: lending books should have been handled by a separate entity, especially when they removed the physical backing requirement.
Furthermore, in the discovery for the Internet Archive case, publishers had already found a case where IA had lent out books despite knowing their partner libraries wasn't actually withdrawing loaned-out copies from circulation. The CDL premise was always just a suggestion, and IA would have still lost their case if they hadn't done the National Emergency Library (NEL) stunt or if they'd been sued in another venue that hadn't had the ReDigi case as precedent.
It's important to note that whenever a company decides to sue for copyright, it is often late, because the company is banking infringements up to the 3-year statute of limitations and because building a meritorious case takes time. The lack of a timely lawsuit proves almost nothing about the intent of a publisher with a valid case against you.
The thing is, I don't even think the whole stunt damaged much of the IA's goodwill? I know of a few people who withheld donations to IA, but that was mainly under the assumption that publishers would be getting a billion-dollar damage award that would immediately bankrupt IA and result in it's archives being sold off to Lexis-Nexis or something. The funny thing is, IA wound up settling for a sum so small they had to promise never to reveal it, and the danger is gone, so the only thing people complain about now is just that the NEL stunt maybe pushed them "above the radar" or something.
It's still insane that shredding books for AI training is legal, but this isn't.
AI training happens to be one of the fields exercising that option, but since it's the current favorite topic for people to hate on, here we are.
The reason why AI companies don't do this is that they're cheap and desperate for training tokens. Same reason why they have scrapers that will happily overload web interfaces for Git repos following links to everything, even though you can just Git clone the repo with far less stress on the host. The AI people are ultimately there just to pillage as much knowledge as they can as fast as possible. Their scraping practices are slap-dash garbage.
[0] https://ones-and-zeroes.ghost.io/scanning-all-the-books-the-...
Your comment seems to also be regurgitating common misconceptions (to put it charitably) about AI and web scraping.
Publishers will always state a maximalist position but the truth of what they accept is what they tolerate without suing.
How are these things remotely related? If anything, Archive.org’s callous, thoughtless approach nuked the hands of legitimate archival efforts.
Turns out that buying an old book for $5 and destructively scanning it for $25 is way cheaper than paying extortion fees to the copyright-mongers.
What I don't buy is it being "rare, precious books". First, they're not after ancient texts - they're after the books that there's still copyright on. Second, when it comes to books, "old" doesn't mean "valuable" - plenty of libraries destroy old books because there's no demand for them, and storage costs you. This is how those scanning companies get books for so cheap.
You know, piracy online is nice and simple - but it's kind of hard to get physical media without paying what the previous owner considers "a fair amount" to part with it.
they should pay the marginal value that the next buyer would buy.
Do you also think that a person dying of thirst ought to pay the maximum price they could possibly pay for water?
Well the argument is that they should pay for the right to produce derivative works not for the physical copy.
> derived works
It's not evident that any works are being produced in the legal sense. Since LLM outputs are not considered copyrightable it's probably closer to using a search engine.
I suppose an argument might be made for fair use if they released their model weights publicly without financially profiting from it.
Instead of blindly jumping on a manipulated outrage bandwagon, people would do well to maybe read some of those old books, not even the rare ones - they tend to contain plenty of parables and stories explaining basics of morality and civilized conduct. We used to teach that to kids at homes and in primary education...
Turning it the other way around is it deeply immoral and sociopathic for someone to hoard massive amounts of money/resources if there are people dying or suffering around them (even if not on the spot but e.g. due to poor access to healthcare)?
Moreover, while it is emotionally appealing to some people to want to add some sort of "responsibility to society" to people who own the old books, it's a very emotional plea that can't really be manifested in the real world. As already pointed out in other places, "an old book" itself doesn't really mean much in terms of what its value is in any particular dimension. Plus, I am always deeply suspicious of anything that expands to "Other people, who are not me, should expend vast quantities of resources so that in the next five or ten times I think about this issue for the rest of my life I feel slightly better about this issue" which is what this really amounts to. I think that as superficially appealing as that may be, it's really a very hostile and demanding position to take.
Personally, to the extent that I would want to lay a "social responsibility" on the AI companies, I'd like to see something like they are either obligated, or ideally, just do it of their own free will, to make the scans of the books that are out of copyright available for some reasonable fee (ideally, "free because we like the PR", but given the scope demanding it be free is not reasonable), and without them trying to lay any further claims on the public-domain results. Trading "one old book somewhere, inaccessible to the world" for "a scan of the book and an OCR of it" I would judge a net win for society for rather a lot of these old books, which are by no means "worthless" sitting in some old collection somewhere but would be a lot more useful for being available.
[1]: https://en.wikipedia.org/wiki/Decoupage , since I imagine a number of people won't know what that is.
They could give a copy of the data once to some third party organisation that then seeds it in bittorrent or something like that. Basically, what I want to say is that this doesn't need to be an ongoing obligation for the scanner to be worthwhile for society.
they're paying for the books, no shady things going on there. whether the publishers should deserve more than a single copy's worth is a separate question.
having the law such that its illegal to scan a book and then keep it, but legal to scan it and destroy it - gg no re there, law people retardmaxxed themselves as they tend to do with anything related to digital data.
More like destructively scanning it for $0.25, if you include the wear and tear on the machine, the salary of the guy who unjams it when it chokes, and recycling fees.
They change the info (many buyers), destroy the source material, and now the lie is in the LLM.
That's it!
> paying extortion fees to the copyright-mongers
Yes yes, greedy fat cat publishing oligarchs treading on the poor put-upon scrappy AI underdogs. /s
Back in reality, the fraction of people who got into publishing books to get rich collecting rents is… not large. There’s so many other fields that are likely to reward participants with more wealth that it’s absurd — even with all the passion for the work in tech it’s probably relatively less pure.
And whatever the excesses of copyright have been, the whole bargain has always been on more pro-social foundations and stronger intellectual foundations than “extortion” sneers. It recognizes that incentives matter and work that’s valuable should be rewarded and incentivized.
A culture that takes a Robin Hood approach to low marginal cost billing points but fawns over the hypercapitalized distribution King Johns isn’t creating a freer or richer society or fighting the real cartel center, it’s indulging resentment and caricature.
That's because the business is buying copyrights in bulk from those people. Copyright-mongers ≠ publishers or writers.
That said, supporting Anna's archive is one of the best things you could do for humanity long term in my opinion.
vinge wrote his singularity piece in the 80s I believe
> Stan Ulam [28] paraphrased John von Neumann as saying:
>> One conversation centered on the ever accelerating progress of technology and changes in the mode of human life, which gives the appearance of approaching some essential singularity in the history of the race beyond which human affairs, as we know them, could not continue.
> Von Neumann even uses the term singularity, though it appears he is thinking of normal progress, not the creation of superhuman intellect. (For me, the superhumanity is the essence of the Singularity. Without that we would get a glut of technical riches, never properly absorbed (see [25]).)
which hews much closer to how people who are fans of the term use it (the Singularity explicitly refers to the rise of superhumanly intelligent systems, not just the progress of technology overall).
Also, holy cow, hard disks (as in the magnetic oxide kind) got a lot more expensive.
Let AI companies do this. But require them to make the digital copies public. Maybe with a multi-year delay, to give the original scanner advantage to doing it.
[1] https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....
Let's say one of the books to be digitized and destroyed is the sole remaining copy of a book from 1850, which is now considered public domain.
On one hand, hoarding such a book, stealing its content from the public domain, locking its content behind a for-profit machine, and destroying the only remaining copy is clearly wrong. It's equivalent to stealing a public resource, just like mining minerals or oil on public lands without a permit or mineral rights. Pure extraction.
On the other hand, taking care to digitize the copy and making it available for free in perpetuity, as well as being required through regulation to provide access to that content through, let's say a public utility LLM/AI available for free through libraries and online... and perhaps after fair due diligence being required to preserve physical copies in a public archive of rare books of which there are no known remaining physical copies...
That seems much more reasonable to me at least. I can imagine there are many who would not see it that way though. Do we see it happening or gaining regulatory, moral and/or public support?
It provides a freedom to circulate, but not access to the material. It is not a public owned resource.
Turning a copy over to the public or state might be an interesting requirement for obtaining a copyright, but instituting that fix for new works now would have a 70 year lag time.
Think of it this way, if I copyright a book and put it in my dresser for 70 years, that doesn't give the public the right to access it or come into my house and scan it after expiry
This has been a requirement in the US since 1790. It is called mandatory deposit: https://www.copyright.gov/help/faq/mandatory_deposit.html
They don't keep every book though.
It increasingly looks like a grand compromise around copyright and AI is needed. Expanding public domain when it comes to AI companies in this respect seems merited. (The other bits are fair use if weights are opened.)
A privately owned physical book and the public-domain work embodied in it are different things. Keeping an old book in a dresser, even without providing access to it, is also meaningfully different from deliberately acquiring the sole remaining copy in order to extract its content, destroy the artifact, and preserve exclusive commercial control over the only remaining usable version.
Destroying the sole surviving copy would not literally steal publicly owned property. But it could irreversibly remove a public-domain work from practical human access and destroy a unique piece of cultural heritage.
In that very narrow situation, preservation, archival deposit, digitization, or public-access requirements may be justified, especially when the material is acquired for commercial use that depends on destroying or withholding the only surviving source.
The company would not be reasserting copyright in the legal sense. It would, however, be creating copyright-like control over practical access.
The work would remain legally free for everyone to reproduce, while the company’s conduct made it impossible for anyone else to obtain it.
I know there are more pressing issues than this one in the world, particularly those involving immediate human life.
But I also believe that the value we assign to human life is connected to the value we assign to human knowledge, memory, invention, and culture. When those things are casually treated as disposable inputs for extraction, something important erodes.
Companies should be free to use public-domain material commercially. The issue arises when that use creates a negative externality by permanently destroying the public’s future opportunity to access and use the work.
Mining our cultural heritage and locking away what remains is still extraction and exclusivity.
Support your local shadow library: https://annas-archive.pk/donate
> You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal.
Is there evidence of this? Since otherwise they could very well be describing what is only occurring to in-print or non-rare books. (This is a genuine question since their post doesn't shed any light on it.)
If it is happening, it is an outrage. However, the 404 article doesn't actually provide any evidence of this; it just connects the shredding of digitized books and the digitizing of rare books to the assumed shredding of rare books, which isn't necessarily happening.
I am not sure that physically destructing one copy of this type of book to preserve its contents digitally is so bad. Pretty much anything that is still under copyright should be valuable only for its content, not for the physical medium it's printed on.
Well, there are special collectors editions with the signature of the author and gold pages or whatnot. But the AI companies are probably not using those.
These are increasingly available online, btw. Historical research is accelerated when historians have direct access to scans of relevant source material. Not destructively scanned, of course.
For whom? is the relevant question
Consequently, many/most books are more available now than they've ever been. This is mostly a copyright question, not a question of preservation. I say mostly, because those scans, without redundancy and fingerprinting, can be changed and bowdlerized in the future without remaining physical copies as a reference.
I own about 3500 print books, and started a project to find scans for all of them (that I need to get back to.) Average publication year is probably around 1975, and the bulk ranging from the 1940s to the 2000s. I made it through about 1500, and couldn't find maybe 40, most of them bad. e.g. self-published stuff like "My God Heals, Does Yours?" This was a few years ago, if I went through those 40 now, I bet I'd find half of them.
I hate what the AI companies are getting away with, but only because they get to violate copyright while being aggressive enforcers of copyright and DRM circumvention laws. Destructive scanning, however, is a cheap way to get a good copy of a book online. If that book were then put in a place where teveryone interested in its contents (or who are just hoarders) can get a hold of it, you'll be able to find a verifiable copy of it 1000 years from now.
As of now, Russia and annas-archive are just a few points of failure that can erase those words forever. It will be done with armed, uniformed men, and people will claim it is not dystopia, but justice. I don't want the only copy of the text of a book to lie in the interpretation of some privately trained LLM.
Now that I think about it, The Judge is an apt metaphor for AI : "Whatever in creation exists without my knowledge exists without my consent."
Leave it up for anyone to download and then compensate the copyright holders later.
In fact if ingesting these books for LLMs is fair use, us commoners should be able to read them for free. Maybe restrict commercial redistribution though.
> They should be forced to publicly release the books as an Ebook.
would be reasonable in any way? The books aren't theirs to release publicly. If I brought a copy of any given movie on DVD, ripped it and used it to train my own "LLM" located at /dev/null, should I then be allowed (or even forced) to release the movie publicly for anyone to watch for free?
Set up a compensation fund for the rights holders. Anything is better than culture literally being sucked into the void.
> The attachment contained 3,000 English-language titles organized by ISBN number, including books such as Distinct Element Modelling in Geomechanics by K.R. Saxena (1999); Barrett's Traditional Fairy Tales (2021), an academic study of Irish folklore; and Laser Shock Peening of Advanced Ceramics by Pratik Shukla (2018).
How is a book from 2021 considered rare in this context? There's almost certainly a digital copy of it in existence prior to Anthropic purchasing a print edition.
A digital copy would exist somewhere, of course. But for us, that only matters if we can buy or download it. And for AI companies, that only matters if they can get a digital copy DRM-free and licensed permissively enough.
(which is terrible, but I would be delighted if they breached it and got thoroughly spanked)
That atrocity of a law was a blight upon digital freedom since the day it came to exist. DRM should never have been given any legal protection - and I would push for numerous forms of DRM to be outlawed instead.
Typically DRMs are implemented server-side so you can control user access in that way.
Quote from Wikipedia[0] of DMCA section 103:
> No person shall circumvent a technological measure that effectively controls access to a work protected under this title. > "circumvent a technological measure" means to descramble a scrambled work, to decrypt an encrypted work, or otherwise to avoid, bypass, remove, deactivate, or impair a technological measure, without the authority of the copyright owner; and > a technological measure "effectively controls access to a work" if the measure, in the ordinary course of its operation, requires the application of information, or a process or a treatment, with the authority of the copyright owner, to gain access to the work.
[0]: https://en.wikipedia.org/wiki/Anti-circumvention_laws#Circum...
Because DRM is just a way to make "breaking copyright" more practically cumbersome. What's easier, breaking digital DRM for each and every E-book you find, or just establishing a single pipeline for scanning physical books?
More importantly, it's also explicitly illegal. Destructive format-shifting is not. Thank copyright laws.
Bearded German man was right.
It's a routine act in any publisher stock management.
https://www.ala.org/tools/challengesupport/selectionpolicyto...
Weeding is the natural process of disposing of less-demand books. Like the rest of us, libraries operate in finite space, so if they want new books, they have to remove ones their users aren't using. Most libraries will try to sell books before disposing of them in any destructive way.
What similarities do you see here?
They are most assuredly destroyed.
Once that AI companies are really shredding 200 year old rare books, and once that libraries are only weeding mass market pulp fiction.
Also I had a family member work in libraries. They got rid of books for very small prices based on how many people checked them out. Not how easy they were to replace. One example was a 100 year old gold leafed book they sold for £5 to someone who removed every page to sell separately.
The fix here is to change the law to permit training AI without destroying the original materials. But that is going to be a heavy lift.
Was was the alternative? Order them to stop quietly destroying their property, it is making someone else very upset?
> Bulk purchases also usually reflect interest in a specific topic, whereas the recent, very large purchases were of books that had little in common, except for the fact that they all had ISBNs. This seller also sells rare books that do not have ISBNs, and none of those were part of the bulk purchases.
Article is paywalled, but I saved a few quotes here:
https://skybrian-links.exe.xyz/post/1026
Horrendous stewardship of humanities collective knowledge all for profit and the race to have the one god computer to rule them all.
As more time passes it becomes clearer that America’s AI strategy should’ve been a public private partnership where the public owned the datasets and the underlying models and we’d leave the productionizing of LLMs to private businesses
Also, they can then just recycle/dispose of the paper and don't have to worry about reselling/donating the books themselves. I suspect this is all about speed of data ingestion and anything else is a side effect they don't care about.
There is a word for that kind of behavior.
This is highly disturbing news; is this standard practice? What did Google Books do before?
The segment that talks about rare books:
> One professional bookseller who specializes in selling foreign language books on these marketplaces told me that, starting in April, he and other booksellers noticed a historic spike in sales. [...]
> This bookseller said his inventory is full of rare, foreign language, and low circulation books, meaning that if they are destroyed in the process of becoming training data, they’ll be even harder to obtain.
> It is equivalent to book burning in the past. A form of thought control
Yeah, this isn't a thing. Throwing out books, even "rare" ones, is a regular procedure. And if you really want to read 'em, you can buy any of these books right now.
> > It is equivalent to book burning in the past. A form of thought control
That's only bullshit if you trust AI companies to serve the book contents without alteration.
The point is that even under the best intentions, hallucinations occur. Then there's the fact that most models have an ideological bias programmed into them.
The only expectation I have is for companies or anybody else to not destroy rare books. Is that such a tall order?
Sure, in the same sense that they regurgitate any other text they consume. LLMs by definition do not have the full training dataset available, though. It’s far larger than the resulting model. So they can’t reliably reproduce full text without an external source (or if it’s in the training data repeatedly). ChatGPT actually refused to give me a bible quote the other day, presumably because I ran into some general “book regurgitation” safety net.
> The only expectation I have is for companies or anybody else to not destroy rare books. Is that such a tall order?
Honestly, yeah. The idea that people or corporations should hold onto books forever because of a cultural “ick” about throwing out books is a bit ridiculous. Most books end up in landfills.
They aren’t feeding Da Vinci manuscripts into this pipeline. They are feeding still-in-copyright books.
Is it? How many different books are we talking about, and how much information is that, after conversion to text and lossless compression? Images, maybe, but text?
That makes it even worse, then. This proves the original point.
> The idea that people or corporations should hold onto books forever because of a cultural “ick” about throwing out books is a bit ridiculous.
If we're building black and white straw man arguments, then sure, let's not archive anything.
I don’t know what the “original point” is here, but these AI companies are not providing “book excerpt services” and do not claim to. ChatGPT at least will refuse to provide detailed book excerpts (I hit a week or two ago myself).
> If we're building black and white straw man arguments, then sure, let's not archive anything.
It seems like you are the one creating the straw man. Do you have evidence that these companies are shredding actually rare books? The only cited concrete examples (in this thread anyway) are all rather boring. I seriously doubt they are shredding 200 year old books because why would they?
I find going after shadow libraries to be much worse because law enforcement is trying to prevent discrimination of knowledge to the public.
The real blame here should be going onto copyright laws.
Anthropic could take more care by figuring out if the books are still affected by copyright.
But this is just a company trying its best in an unfortunate regulatory environment.
Support your local shadow library:
https://annas-archive.pk/donate
Say i feed the largest LLM a book of an alien civilisation, that it definitely hasn't seen before. Then this tiny piece of text muds the vast ocean (latent space) of the model minimally. It will not be able to cite from that book reliably after that fine tuning. Especially for rare books, because them being rare implies, that there aren't 1000s of other books, that encode the same information.
It's general language modelling capabilities might get an iota better, of course. But for putting factual information into it, wouldn't RAG be a much more solid approach?
You make it sound like they are running a second Project Gutenberg. They most definitely are not making these available for electronic searches. At least not searches the public can participate in.
I would say “reduced”
What guarantee do we have that the book contents will be served unfiltered and unaltered?
They are creating embedding vectors and training documents - changing whatever they want - and destroying the original copy so nobody knows what was actually said.
It is just like book burning.
I wish...
It's a real shame that no one ever got that book in front of Hayao Miyazaki's eyes.
When you compare to kids learning from home during the pandemic, on Chromebooks and iPads, building TikTok videos with CapCut instead of an actual local NLE application... you can start to hear the echoes.
I looked it up.
surveillance, sousveillance - sure, why not? Seems harmless.
But it's as if the more people know the word sousveillance, the word and even the act of surveillance actually loses a little bit of power (its embedding changes, if you will!)
All our fears about the end state of surveillance can now be countered by an end state of sousveillance.
Before I had sousveillance to think of, I could only think of surveillance (when thinking of veillances) - and it was more of a threat then than it is now.
The world is made of language, or as Terence McKenna said made out of words which sounds obviously false at first. But go looking for the inside of an atom and tell me what you find, and think about where the medium of reality actually implements itself.
It's sad in a romantic kinda way, because of the lost artifact, but the information is what makes the book valuable, not really the medium.
The out of copyright books don't really need to be destroyed anyway for them to be fair use for AI training, and arguable, even if you needed to, you only need one copy per title per company at most.
So it's not a gigantic loss.
They aren’t verbatim uploading the text 1:1. They are creating vector embeddings and training documents from it, changing whatever they want since it’s in private and protected by NDA, and destroying the original source of information so that nobody knows what was originally recorded.
Isn't that oxymoronic? If they can be bulk-bought they aren't rare.
Why would I spend hours creating original content if Google can extract it and present the answer directly in an AI Overview? What is the incentive to keep doing the work?
If creators stop producing high-quality original material, the information we get over the next few years will increasingly be based on recycled, low-quality garbage.
Can Google steal it and present it in an AI overview? Well kinda. Today Google is doing a trick - they're saying "You can refuse to consent to being fed into the slop machine, but if you do we won't crawl you for Google so you'll get no search traffic. But you're not going to get search traffic anyway! So you might as well opt out of being fed into the slop machine. And companies are starting to do that [1]
It's really interesting, because essentially what it means is Google is turning into a walled garden, but there's nothing growing inside it so they have to continually import new plants to live in their walled garden and they're going to have to pay to do that. So soon Google will be paying news sites for the right to plumb their feed into the slop machine.
[1]: https://www.wsj.com/business/media/google-search-publishers-...
A spurious claim; they simply want to avoid model collapse.
Weird to see so many of these "trust me bro" twitter stories make it to the front page and cause outrage when no one has any real information.
This is merely the latest incarnation. We can imagine a slightly different process on a few fronts - AI companies pay to digitize books (still for their own purposes), but are prevented from destroying the physical copies and they have to openly shared the digitized results. We would view that situation much more favorably - perhaps even as ideal, right?
Those two dynamics could be backed up by court decisions or laws iff they weren't so plainly at odds with how copyright has been and is generally implemented and interpreted. For example, imagine them having to do this through some nonprofit library whose goals was preservation and dissemination. Instead, libraries have been sidelined as things that operate at the edge of the law rather than vital public institutions, whereas shredding books in secret is fully legally condoned.
Copyright has done more than anything else to prevent preservation and dissemination of knowledge. And it's forcing Anthropic's hand now. Although they could take more effort to preserve the books.
What I'm indicting is the copyright regime being primarily focused on control and the prevention of dissemination. We can imagine a different world in which the publishers' suit against the Internet Archive went the other way (or was not even brought), and a public interest group sues Anthropic (et al) for destroying cultural commons, and gets a judgement saying all scanning must be done non-destructively and made available through institutions like the Internet Archive.
https://www.404media.co/ai-companies-are-buying-tons-of-old-...
You know that pleasant used bookstore smell? It's paper slowly decomposing.
Modern acid-free paper might last 1000 years; 500 is more typical. Acid paper breaks down in less than a century.
Parchment was so expensive it was often scraped and reused; old texts can sometimes be recovered after being overwritten (palimpsests).
They’re privately putting info into their LLMs, changing whatever they want, sorry, “sanitizing” then destroying the original.
And the only real source for anything will be an LLM response.
https://booksale.org/
Books are information delivery vehicles. We mostly shouldn't care about them any more than we care about a particular set of bits on a disk.
Publishers also pulp large numbers of books themselves. This is a consequence of the Supreme Court's Thor Power Tools ruling, which clarified tax rules in the US so that keeping large inventories of unsold books was less economical.
It's a weird practice to destroy a book when you digitize it, but it's at least an understandable legal strategy to ensure that the digital copy "replaces" the physical one. However, this is only going to apply to books that have active copyrights.
This author suggests that a rare 18th-century botanical text could fall victim to the same fate, but I am somehow doubtful that this is the case. Non-destructive scanning is trivial, and these kinds of books are likely being processed in a quantity that would allow for it without backing up the pipeline.
I really like 404 media, but it doesn't really seem like the evidence points to the conclusion here. Yes, AI companies are shredding books that they digitize, and yes, AI companies are digitizing old, rare books. But the rationale for the book shredding doesn't exist for the old, rare books, so I would need more evidence than just "putting two-and-two together".
At the end of the day, old books with no resale value, including rare books, end up destroyed with some regularity by libraries and bookstores. While this may be an excessively generous take, at least this way the books are getting digitized before they become pulp. The real tragedy will be if the old, rare books that were digitized are never shared with the rest of us, because they were ONLY digitized to train AI, and not to actually preserve anything.
Just one more book, bro, and we'll solve AGI forever. Trust us, bro, just one more book. Come on, let me have those words and we'll solve AGI forever.
Really? That sure wasn't a thing when one startup got sued for streaming from its wall od dvd-players, and they adhered to 1 disc = maximum 1 stream at same time.
* their values are aligned with the best interests of humanity
* their values are not aligned with the best interests of humanity.
Now go ahead and read my mind!
But that's a lie. They're not destroying the last copy of books.
Oh, I have a reputation here. Thanks for letting me know.
But that would help competitors with training data, which I assume is why they don’t do this.
Luckily I don't live in NYC where book hoarders are being evicted because of "fire hazard". Just like in Fahrenheit 451.
Of course even disregarding this fact, this is utterly disgusting attack on humanity heritage, just as much as any group out there destroying what we should all cherish be it for the historical artifact they represent. Whatever how US judge name it, they don’t worth more than their same-behavior consorts that is terrorists and totalitarian governments.
https://en.wikipedia.org/wiki/Palimpsest
https://organiser.org/2025/07/24/304285/world/china-wages-wa...
https://www.historyexpose.com/things/demolition-afghanistans...
https://link.springer.com/chapter/10.1007/978-3-031-96432-9_...
Like, Library of Alexandria or Council of Nicaea bad.
We may never be able to recover the information if, say, one of these AI companies copied or translated it wrong then destroyed the source material.
Maybe it's from bad OCR, or maybe from a bad actor - but there are a lot of ways history and information could change in this game-of-telephone like transfer of knowledge.
What is the point of destroying the source material? I don't buy the copyright thing.
It is the copyright thing.
Despite what people say about scanning, the fact is, non-destructive scanning machines have been built and perfected long time ago. This was preferred in the past, back before some major kerfuffle with the publishers during COVID, but that incidentally happened to be before LLMs became a thing, so AI companies never had that option available.
It is the cheapest way to get them scanned.
It is the fastest way to get them scanned.
It doesn't need to be safely archived for another century until it is resold to someone that has not yet been born.
Is that possible at all or nah?
The main reason is the machines they use to scan books at scale destroy the books in the process.
If we use 'shredding' to mean A book is laid flat, its cover is removed, and then a paper cutter cuts through the binding to create a flat stack of sheets, which are then fed to a sheet feeder, then I could maybe imagine this is better than a page-turning scanner. But, sheet feeding old paper sucks shit, bro. It's not fun.
Upshot, I think we'd like to hear from an anonymous frontier lab employee here to see what's going on -- there are a lot of books in Anna's archive available at considerably less difficulty.
These companies are regressive book burners.
Maybe the Library of Alexandria didn't burn, it was digested.
The market has solved the what do with excess knowledge problem.
Indeed, I think there's a high chance that this process increases preservation of the most relevant part of the media - the actual content!
There are tons of old rolls of film slowly rotting away in warehouses that were never digitised. Even for beloved media, the BBC occasionally tracks down an old lost episode of Dr Who.
For now these books are in corpuses of training data, but eventually I trust they will make their way to the rest of us.
What makes you think they will? What would be the incentives for these companies to do so?
1. At some level of critical information-withholding mass, a leak or disclosure similar to SciHub is inevitable because of the commonly held opposition to hiding knowledge.
2. Availability via Google Books or similar.
3. Availability via AI model reference.
4. Failing any of the above, better AI models that are more capable of doing more things, at the expense of books that were likely to go unread (revealed preference, rare books are often rare for a reason). This will obviously be a nonstarter if you don't want this to happen, but I think it would be good for the world if it did.
I think category of old books that were going to be read or otherwise become important parts of human knowledge that have not yet been digitised and now will never become so because they are instead being shredded and will never make their way into the light because of AI company data hoarding is a small category.
I am however OK with destroying one of the remaining 50 children's books of which only 300 copies were ever printed in a small town in Ohio in the 1970s as a test run for a failed book which was subsequently never commercialised.
What if the perception of those books change over time and are considered masterpieces later on?
Moby Dick was out of print when Melville died 1891 and not a huge success until it got a revival in the 1920s
This is the kind of rationalization hoarders use. The inability to get rid of things because it could turn out to be something we want in the future for reasons that we cannot currently describe.
It's a loss avertive instinct that I think is misplaced. Treating every printed book as priceless is intractable. It's not how we treat these books at the moment. Apparently, today we don't even care enough to spend a few hours per book nondestructively scanning them in.
Let's say we, instead of destructively scanning this books, nondestructively scanned them. What would you propose doing with the copies afterwards? Sell them? To who? They're valueless individually for the overwhelming part. Warehouse them? Why? For who? Do you want to go and look through them? Why haven't you done so already? Have you ever shown interest in consuming an undigitised book a single time in your life to date?
It feels like the anti-AI crowd here have to tie themselves in knots here to make the loss minimization work.
That’s the whole point because the book shredding is already declared legal.
https://en.wikipedia.org/wiki/Visual_Artists_Rights_Act