Before my lifetime, copyright law all but choked and killed the public domain. And now everything is stale and the same.
This particular battle was lost with Google Books and the attempt to make the world's largest library. The modern library of Alexandria.
But then copyright lawyers got involved to get their pound of flesh. And here we are.
I am upset about the destruction of knowledge. Paper is a superior storage medium to any hard-drive any day. We're recovering words from paper from over a thousand years ago. I think it's a mistake to not work with a non-profit, use cheap COTS non-destructive scanning, and write off the costs of rebinding them and rehousing them.
I want copyright to get enforced against these corporations, because then they would have a reason to use their considerable lobbying power to change copyright law, hopefully in a way that benefits everyone (including things like an open digital library).
> Paper is a superior storage medium to any hard-drive any day. We're recovering words from paper from over a thousand years ago
Most paper books don’t make it a thousand years, let alone a hundred. Your local library is probably purging piles of books every year and nobody sheds a tear.
> I think it's a mistake to not work with a non-profit, use cheap COTS non-destructive scanning, and write off the costs of rebinding them and rehousing them
It would be cheaper to buy a second copy of the book and put it in a different warehouse where it can continue to go unread as it did before the AI companies bought it. The cost of non-destructive scanning and rebinding books is crazy high.
It is also difficult to see what the legal basis would be for allowing people to read books if Amazon's use isn't fair. If I read a textbook and come up with an algorithm based on it that gets used at scale, is that supposed to be a copyright violation?
The basic point of a book is for people to (a) read it and (b) synthesise it into a comprehensive world model. And people can burn their own books if they want to, they own the thing. The only possible complaint here seems to be the scale and it seems like a big challenge to say that is a problem given that knowledge from books is allowed to be used at scale.
Incredibly. Even if we throw ethics aside and look at the strategy it is still myopic. The idea is destroy so others can't get the data too. But this assumes you've done a perfect scan and you have backups. There's no good reason to destroy the books other than myopia.
I'm more than willing to bet that in 5-10 years they'll wish they didn't destroy the books, specifically the rare ones.
To be clear, I'm not trying to condone their behavior. I hate it. But I'm trying to show that even if you had their same ethics it's still dumb. I hope workers are secretly stashing the books away. If you're one of the people in charge of destroying the books, you have a cultural duty to preserve them. I'm willing to bet people will go to great lengths to help you do it secretly so you can continue to keep them safe and keep your job. If you're at Amazon, or any company where this is happening, you have a duty to the world to make efforts to preserve the books. Steal the PDFs and lock them away. Create backups in your company. Tell your bosses you don't think this is right. Don't sit silently while this happens. Silence unfortunately is enabling. Unfortunately silence isn't a passive action
This has been less true every year since Usenet and Geocities. Much of the stuff on their was pirated, but that was the "stale and same" stuff you're complaining about anyway. The rest of the stuff was all kinds of original, good and bad. There's never been more total "content" and more total variety. You just have to look around for it (but that's also never been easier).
Letting AI companies try to profit off of all that creativity forever while choking the ability of the creators to make future revenue off of it is exactly what would actually lead to a "everything is stale and the same" situation. Preserving variety and novelty of new creation would look like putting in new restrictions to reduce this not-envisioned-when-the-laws-were-made sort of read-once-slop-forever usage.
It's not necessarily doctrinally that crazy in copyright law. Some courts have considered AI training on copyrighted works a "transformative" use (traditionally more protected in fair use analysis), while providing human beings access to the works a "consumptive" use (traditionally less protected).
There's also a fair use consideration that favors noncommercial use compared to commercial use, but that's not the only question, and the statute doesn't say clearly how to combine the fair use factors.
But it's possible under the Copyright Act that some commercial uses of copyrighted works could be considered fair uses while some noncommercial uses could simultaneously not be considered fair uses.
Wasn't copyright law supposed to support creation and propagation of works of art ? Instead, apparently it is A-OK or even encouraged by the same law to destroy books.
That does not make any sense! Really might be high time to scrap it all.
The ruling in the Anthropic case was that their scanning of the books was only fair use because the original physical copy was destroyed. The copyright lobby wanted scarcity so they could monetize things, and now folks are worried about knowledge being scarce. It's a problem of our own making.
Issue really is first sale doctrine. Meaning that after first sale there is very much leeway for the product owner. Even to scan and destroy it.
Copyright ways. Form change probably should be compensated. Even if it means that you wouldn't be able to format change your own media. Or that such activities wouldn't be allowed without compensation over certain threshold.
This is in tension with Sony vs. Betamax, which held that time shifting is fair use, but that also requires format-shifting (broadcast to recorded media).
Courts found that having a central library of 7 million pirated books is against the law, and assessed a large penalty (small for these robber barons) so Anthropic is destroying them now to conform to the law. The law is bs and Anthropic is run by villains.
No, it doesn't. Copyright doesn't care about copies, it cares about copying. If it did care about copies - i.e. the total number of copies in circulation - then ReDigi and the Internet Archive's Controlled Digital Lending (CDL) program would have both been legal. Destroying a copy does not give you permission to create a replacement copy.
The only relevant case law for AI training in the US is the rulings in the Anthropic lawsuit presided over by Judge Alsup. That lawsuit ruled that it's infringement to build a shadow library from pirated books; but NOT to train AI on those pirated books. The only point where destructive book scanning even comes into play is that Anthropic also had a book scanning program alongside their piracy, Judge Alsup said that program was not infringing, and Anthropic happened to be destroying books. At no point did Alsup say that leaving the books whole would have infringed copyright - it was never even considered as it was outside the scope of the lawsuit.
Now, if Anthropic were to non-destructively scan books, store them in a library, and sell the books on, that could be infringing. All the case law about format shifting presumes the owner retains the original. So Anthropic would likely have to hold onto books, at least the ones they wanted to train on, until they were done training on that book[0]. But they do not have to destroy them permanently. They are destroying these books specifically because it is cheaper to do so than to use, say, the Internet Archive's own custom-built nondestructive scanners.
[0] I am absolutely furious about how much this sounds like "fair use is just an extra license you get when you buy a book", and I would much rather have had Judge Alsup just say AI training is not fair use instead.
The destruction of the books is mentioned many times, and the overall analysis says this:
> The copies used to convert purchased print library copies into digital library copies were justified, too, though for a different fair use. The first factor strongly favors this result, and the third favors it, too. The fourth is neutral. Only the second slightly disfavors it. On balance, as the purchased print copy was destroyed and its digital replacement not redistributed, this was a fair use.
I do not think the ruling would have gone this way if the books were not destroyed, as many would assert that Anthropic would retain them to sell later, otherwise. The fact that destruction is mentioned so pervasively in the decision suggests it is an important fact to consider.
Lars has seemed to me like such a basket case this entire time.
He's lead Metallica through the "Copy our tapes! Let everyone hear our music!" early days, through the "They're stealing our master recordings!" Napster and Senate hearings era, and come all the way back 'round to "I’m just happy that fucking anybody cares about what we’re doing and shows up to see us play and still stream or buy or steal our records or whatever.” [https://consequence.net/2023/09/metallica-lars-ulrich-stream...]
Well, this is what happens when they weren't allowed to use digital only copies, as they now have much more of an ability to get around copyright with physical items. The right thing would've been to allow digital items to also have the first sale doctrine, but we don't, so AI training companies must resort to destroying physical books in the process of scanning. The Cobra Effect strikes again.
I wonder how much of AI is about owning all knowledge, deskilling the masses, reducing access, etc.? The end goal might be actual legit surfdom. Without the means to elevate the people's knowledge (and the companies have already stated they Benevolently Protect Users From Harmful Data™), it quickly becomes a situation where they control all narratives. It took less than a generation for social media to make the spectacular mess of discourse we now have. How long before all knowledge is mediated and moderated?
Original full title (too long): It is a sign of the times that Amazon gets to call this fair use while huge corporations try to sue the Internet Archive out of business.
I remember reading a scene where this exact thing was going on in Vernor Vinge's Rainbow's End: there, the books were being shredded and scanned as tiny bits and stitched together in software similar to shotgun genome sequencing.
At any rate, I ought to pick it up again and finish reading it...
this is a complicated way of counting how many of each word is in the book.
clearly fair use.
does the author think cutting up a book is a copyright concern? theyre buying the books, its up to them what to do with their copy. if you want the books preserved, maybe fund your libraries to get a copy or two?
It's possible for someone to do something socially reprehensible and entirely distasteful and to be perfectly within the letter of the law while doing so. The article doesn't emphasize copyright issues or fair use, because that isn't the point. The post is trying to argue that the act is not sustainable and needlessly destructive, in pursuit of something that many people believe has limited value or unproven value, which is true.
Not the case here, but it's also important to keep in mind that laws too, can be unjust and immoral. Much of humanity's progress has involved the abolishment of immoral laws.
These are books that no one else wants and would get pulped otherwise. No one has space to save the 1754 edition of "whale blubber cookery on the open seas".
Yeah but I unironically would love to have a copy of that on my bookshelf. It likely has fantastic illustrations!
This is the whole problem. A lot of people assuming they can speak unilaterally and with authority on what other people value or should find valuable.
It's disgraceful. What happens to the old hacker ethos? I'd never thought I'd see the day when, on a site called hacker news the majority of posters are defending the right of a large corporation to destroy public goods permanently just to be able to charge people more for "intelligence as a service".
The full headline tries to equate this to The Internet Archive’s recent legal troubles.
For a refresher: The Internet Archive scanned books then shared them online, trying to claim that converting them into digital copies qualified as a derivative work. You don’t have to be a lawyer to see how that claim doesn’t hold up to any scrutiny.
The LLM companies are not redistributing the works. They are using them for training. This blog post calls it IP theft, but it has actually been litigated in court already. Using books for training does not qualify as theft or redistribution, even though some people have different opinions about the moral angles.
One of those same lawsuits also extracted a huge settlement from Anthropic for using digital downloads from pirate sites. The conclusion was that the only acceptable way for them to use the books is to buy them and scan them. The courts forced it to be this way.
The current hand-wringing about the destruction of books is based on claims that they’re doing it to rare books that are also valuable. So far nobody has been able to actually provide an example of one of these books that is supposedly ultra-rare, but also valuable, but also only available at one of these book sellers that sells these things in bulk. We’re supposed to assume that one of these books might actually be super valuable but also super rare and also only available at these places they’re buying from.
It's a win-win-win for Amazon if they destroy the source in the process of scanning it. They get the AI training data, and it's cheaper for them to destroy the book in the process. Destroying the book ensures that it will be more difficult for competitors to scan the same content, thereby increasing the value of the data they've scanned.
Lastly, there's a misguided belief among some that it's somehow less of a copyright violation if the source is destroyed, but it's a violation either way unless the entity doing the scanning has permission from the copyright holder to make the copy.
Is it? They keep the scans (somewhere). These books would have rotted in bookstores and then been trashed once those bookstores went out of business.
There’s no giant “used” section at Barnes and Noble. Libraries only have so much space and they have to cycle through books as the years go by. This is easily the best fate these books could have received.
Bullshit...they aren't sharing them with the internet archive, and at least the inventory of the defunct bookstore might have ended up with a collector or something. These slop factories have a less sound business model than any given bookstore, and you think they'll preserve those scans and share them with the world? Corporations regularly delete media for utterly capricious reasons.
> As mentioned elsewhere, this is also IP theft on a massive scale.
Scanning a legitimately purchased book is IP theft? How can he hold such a copyright-maximalist view, and at the same time defend the Internet Archive?
Your granpa once wrote an obscure book about this amazing way he'd found to cure diabetes.
Corporation A buys all existing copies of the book, scans them, destroys the originals, and sets up a commercial business offering a monthly subscription to alleviate diabetes pains with this new method they claim they discovered.
Person B borrowed the book from a municipal library, Xerox'd it, and keeps a free ledger, open to all who want to read old books, as a way to safeguard free access to the world's knowledge.
Do you think there could exist any possible logic by which some people would defend person B and try to stop Corporation A?
In this hypothetical, my grandpa didn't patent this way to cure diabetes (otherwise the company would owe him royalties, assuming the patent hadn't expired).
So it reduces to a corporation selling services based on public-domain knowledge (I am ignoring the part where a miraculous advance is confined to a single unknown book). Lots of corporations do this, and there's nothing wrong with it.
I'm sure people would prefer if Amazon also offered library-like access to the book, but how is it theft? I'm not asking about "any possible logic" - he called it theft.
Words are subjective. The person is saying that it is theft and if the law doesn't define it as such then the law should change. The implication is that taking from the public domain is theft if you then make it unavailable by destructive processes which render the public unable to access that which was previously available.
> The implication is that taking from the public domain is theft if you then make it unavailable by destructive processes which render the public unable to access that which was previously available.
That is also my understanding of the claim, and I would intuitively agree with it.
> So it reduces to a corporation selling services based on public-domain knowledge (I am ignoring the part where a miraculous advance is confined to a single unknown book). Lots of corporations do this, and there's nothing wrong with it.
Ok, but it's a bit different because they are not selling "the book" they are selling a fundamentally different thing for which the book is consumable input. Let's consider a batch of 5000 books. For kicks let's assume the known copies are N=1 for those books. Today, though only a few people can do this at a time, a person can purchase one of those books once read exactly the text in that book, and keep doing so to their heart's content. Or maybe a library buys it and now a rotating legion of people can do that for free. Now let's say amazon buys and scans them, destroying them in the process and refusing to release scans to avoid giving competitors a training edge. Now:
- You can never get the information as written again. At best you'll get an LLM output approximation/mutation of it. If this was the expression of a real human beings lived experience, that's kind of sad and goes against one of the spirited aspects of human existence, to leave a legacy, doesn't it?
- Amazon is going to charge you per token every time you want to access that information.
- Price demand for the information is now tied up with general demand for LLMs rather than the actual book, either reducing or greatly inflating the cost to you in addition to the now recurring charges.
So, now you (a) can't actually ever access that book as it was written (b) need to pay continually to access an approximation of its contents (c) and possibly more than the book is worth since now it's "value" in a price sense has been absorbed into general llm inference costs.
Not to mention you may have eradicated the last extant copy of a person's memoirs, but I guess it's pretty clear people in this industry don't care at this point. I hope one day your entire life story is ground up and consumed in some data farming operation and you are all summarily forgotten.
They didn't say that the Internet Archive wasn't "IP theft" - nor did they say that IP theft was inherently wrong.
"AI training is infringement" is not exactly a copyright-maximalist view. The explicit training task used for pre-training is reproducing the content of the trained-on books; and models trained on such books are able to reproduce significant infringing chunks of them[0] unless specifically post-trained to refuse to do so.
Additionally, they might have thought that Controlled Digital Lending was OK (it wasn't, but that's a different issue to AI training). As I've mentioned elsewhere in this thread, there's a common misconception that copyright is concerned with the number of copies in circulation as opposed to individual acts of copying.
Or they don't care about any of that and just wanted to highlight the hypocrisy.
[0] Which, under the "compression is intelligence" point of view, is entirely expected and not surprising in the slightest.
> As the bookseller who sold them told me, there are not many people in the world who would care about them in the same way people might care about the first edition of Oliver Twist, but that doesn’t mean they’re not valuable.
It does though. A lot of people seem to think "it's a book!!!" means it takes on some mystical intrinsic value. That's bullshit. There's an absolute mountain of worthless and near-worthless books. I bet the vast vast majority of these books are those books - after all that's what Amazon wants. A ton of human written text as easily as possible. They obviously aren't buying first editions of Oliver Twist, or even second editions of Harry Potter.
> We’re not revealing the titles of the books included in the shipment we tracked
Yeah... because then it would reveal how unimportant they are. This sort of outrage inflation is counter-productive. People see through it, lose trust, and then when something bad actually happens they won't believe you.
Who gets to decide what's "special"? Written works might be personal stories or expression of feeling.
I've read a lot of unknown old literature that has little to recommend it as far as literary history goes, but it all has a certain charm, and at the end of the day it was all the work of a human being who walked this earth just like we did and captured their thought on the page.
What if their family has an interest in the preservation of that text? What if they simply don't yet know of it, or couldn't afford it?
Yes all hypotheticals, but I think a lot of people are underestimating the potentially permanent damage being done to the public good just because the actions are not illegal.
Everyone should be uncomfortable with a single company unilaterally pillaging the public good for its own ends. Period.
IDK about you geniuses on hacker news, but I'd much prefer to pay a few hundred dollars for a rare book once to be able to read it than pay amazon token costs monthly to get the LLM's chopped and screwed regurgitation, without any recourse to even knowing where the information comes from.
Then why does Amazon want them? Why train an AI on worthless books?
Individually you might be right. Some cookbook from 20 years ago probably is worthless to most people. But then there'll be one person who spends an afternoon hunting it down, bookstore to bookstore, because their aunt used to have a copy and they want to check a recipe, or some such.
I've been looking for a recipe book that my mother used to own called "Taste of the Tropics". Every few years I search online, and while I find a listing from time to time, it's always been "Currently unavailable".
I just searched again now and found 3 used copies on AbeBooks for about US$11 incl. shipping. Absolutely made my day.
I don't know if adult me will enjoy it beyond the nostalgia, but there's a coconut cake recipe in there that I will never forget.
Like you said, just about nobody on planet earth would care if the billionaires ingested and then destroyed those final copies, but I'd have been gutted if they'd beat me to it.
>Some cookbook from 20 years ago probably is worthless to most people. But then there'll be one person who spends an afternoon hunting it down, bookstore to bookstore, because their aunt used to have a copy and they want to check a recipe, or some such.
ask your aunt for the recipe, look it up online with a search engine, ask an LLM by describing how it tastes and the ingredients you remember, this is really not that complicated
>This is cultural violence on a massive scale.
could you be any more dramatic? you're acting like we are actually losing anything when in reality we're gaining from this, you'll just be able to ask the AI for the recipe instead of wasting hours and hours hunting down some old book and ordering it online for an exorbitant amount
Just like with most things, LLM's are terrible at recipes if you are actually a good cook.
Your example is very flawed because if the LLM gave you the exact recipe then it would be infringement, but many recipes need to be exact. LLMs aggregate information so wanting a very specific recipe (to the point of hunting down an old cookbook) is a non-starter because it the LLM is really only guessing.
I agree with you on your point, but I just wanted to let you know that recipes can't be copyrighted.
If you write: "Smash one egg into that bowl and feel the power. Sprinkle salt onto it salt bae style." That's copyrighted and I can't copy that word for word.
But I can publish my own recipe which is an exact copy of yours but written like so:
Ingredients:
1 egg
1 teaspoon salt
Method:
Combine egg and salt.
Or write that in any way I want that isn't a word for word copy of your recipe.
What are we gaining other than your silly little hypothetical?
The whole problem here is:
- Not everyone agrees on the utility / value of AI.
- A small handful of companies control frontier models and get to shape how these inputs are ultimately put to use.
- As Sam Altman put it, they want to charge you, in perpetuity, for "intelligence" as a service. Contrarily, anyone can access a book usually for a one time purchase fee and read it forever. You'll be paying per token to look up auntie's recipe every time instead of just, you know, reading the book.
Not to mention you are erasing the voices of human beings who had lives and may have related their experiences and emotions in those texts. That anyone would be comfortable with a corporation they know nothing about destroying potentially unique information of the lives and stories of real human beings, locking it away forever as one granular component of a stochastic content generator, the human expression in its actual form never to be accessed again, is absolutely appalling to me.
I think you should reflect more on who will really benefit, who has the control, and why humanity as a collective is potentially losing as collateral, without having much say about it.
What we have is not capitalism. There is way too much gate-keeping regulation for it to be called capitalism.
Just like China does not have communism. They have too much free-market activity and private ownership for it to classify as such.
If you look at China and think "That's not communism", just be aware that they may also look at the west and think "That's not capitalism." Both are right. In fact both systems have converged.
China took the better aspects of communism and capitalism. We took some of the worst aspects of both... Except maybe free speech; but people are working hard on removing that one!
It makes sense; China came out of an era of extreme government control which has been reduced over time (compared to what it was before). In the west, government control has been on the way up and we can't even imagine how bad it can get.
Most people agree it isn't normal. The problem is the people at the levers of power think otherwise. Capitalism as it exists in the US today, with effectively no regulation and deeply intertwined with governance, but in the wrong direction (corporations steering government) is basically the perfect exemplar of what they call a vicious circle in systems thinking. Couple that with the isolationist and individualist principles at the heart of American thinking and you have a perfect storm of runaway variable maximization (aka instability) as increasingly smaller groups of people amass increasing amounts of resources and power.
Wage labor is then the perfect structure to ensure that most resistance from the bottom is suppressed before it can even really start. When you depend on the corporation to succeed for your subsistence (since there is no adequate social safety net to fall back on) taking a principled stance and refusing to destroy great works of literature becomes a highly costly endeavor for most people.
Before my lifetime, copyright law all but choked and killed the public domain. And now everything is stale and the same.
This particular battle was lost with Google Books and the attempt to make the world's largest library. The modern library of Alexandria.
But then copyright lawyers got involved to get their pound of flesh. And here we are.
I am upset about the destruction of knowledge. Paper is a superior storage medium to any hard-drive any day. We're recovering words from paper from over a thousand years ago. I think it's a mistake to not work with a non-profit, use cheap COTS non-destructive scanning, and write off the costs of rebinding them and rehousing them.
Everyone is impressively short sighted.
Most paper books don’t make it a thousand years, let alone a hundred. Your local library is probably purging piles of books every year and nobody sheds a tear.
> I think it's a mistake to not work with a non-profit, use cheap COTS non-destructive scanning, and write off the costs of rebinding them and rehousing them
It would be cheaper to buy a second copy of the book and put it in a different warehouse where it can continue to go unread as it did before the AI companies bought it. The cost of non-destructive scanning and rebinding books is crazy high.
The basic point of a book is for people to (a) read it and (b) synthesise it into a comprehensive world model. And people can burn their own books if they want to, they own the thing. The only possible complaint here seems to be the scale and it seems like a big challenge to say that is a problem given that knowledge from books is allowed to be used at scale.
AI content is not exactly helping everything feeling like Extruded Product.
I'm more than willing to bet that in 5-10 years they'll wish they didn't destroy the books, specifically the rare ones.
To be clear, I'm not trying to condone their behavior. I hate it. But I'm trying to show that even if you had their same ethics it's still dumb. I hope workers are secretly stashing the books away. If you're one of the people in charge of destroying the books, you have a cultural duty to preserve them. I'm willing to bet people will go to great lengths to help you do it secretly so you can continue to keep them safe and keep your job. If you're at Amazon, or any company where this is happening, you have a duty to the world to make efforts to preserve the books. Steal the PDFs and lock them away. Create backups in your company. Tell your bosses you don't think this is right. Don't sit silently while this happens. Silence unfortunately is enabling. Unfortunately silence isn't a passive action
This has been less true every year since Usenet and Geocities. Much of the stuff on their was pirated, but that was the "stale and same" stuff you're complaining about anyway. The rest of the stuff was all kinds of original, good and bad. There's never been more total "content" and more total variety. You just have to look around for it (but that's also never been easier).
Letting AI companies try to profit off of all that creativity forever while choking the ability of the creators to make future revenue off of it is exactly what would actually lead to a "everything is stale and the same" situation. Preserving variety and novelty of new creation would look like putting in new restrictions to reduce this not-envisioned-when-the-laws-were-made sort of read-once-slop-forever usage.
There's also a fair use consideration that favors noncommercial use compared to commercial use, but that's not the only question, and the statute doesn't say clearly how to combine the fair use factors.
But it's possible under the Copyright Act that some commercial uses of copyrighted works could be considered fair uses while some noncommercial uses could simultaneously not be considered fair uses.
That does not make any sense! Really might be high time to scrap it all.
On the other hand, absolutely yes we should abolish all IP law.
Build a lending library, or build a bonfire, or do anything else with them that you choose.
They're your books. Have at it.
Copyright ways. Form change probably should be compensated. Even if it means that you wouldn't be able to format change your own media. Or that such activities wouldn't be allowed without compensation over certain threshold.
And in times of substantial inequality, they lag yet further.
The only relevant case law for AI training in the US is the rulings in the Anthropic lawsuit presided over by Judge Alsup. That lawsuit ruled that it's infringement to build a shadow library from pirated books; but NOT to train AI on those pirated books. The only point where destructive book scanning even comes into play is that Anthropic also had a book scanning program alongside their piracy, Judge Alsup said that program was not infringing, and Anthropic happened to be destroying books. At no point did Alsup say that leaving the books whole would have infringed copyright - it was never even considered as it was outside the scope of the lawsuit.
Now, if Anthropic were to non-destructively scan books, store them in a library, and sell the books on, that could be infringing. All the case law about format shifting presumes the owner retains the original. So Anthropic would likely have to hold onto books, at least the ones they wanted to train on, until they were done training on that book[0]. But they do not have to destroy them permanently. They are destroying these books specifically because it is cheaper to do so than to use, say, the Internet Archive's own custom-built nondestructive scanners.
[0] I am absolutely furious about how much this sounds like "fair use is just an extra license you get when you buy a book", and I would much rather have had Judge Alsup just say AI training is not fair use instead.
> The copies used to convert purchased print library copies into digital library copies were justified, too, though for a different fair use. The first factor strongly favors this result, and the third favors it, too. The fourth is neutral. Only the second slightly disfavors it. On balance, as the purchased print copy was destroyed and its digital replacement not redistributed, this was a fair use.
I do not think the ruling would have gone this way if the books were not destroyed, as many would assert that Anthropic would retain them to sell later, otherwise. The fact that destruction is mentioned so pervasively in the decision suggests it is an important fact to consider.
https://storage.courtlistener.com/recap/gov.uscourts.cand.43...
“We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility” (404media.co)
164 points | 3 days ago | 321 comments
The whole "copyright infringement is theft" meme seems to have stuck.
20 years ago it was all "information wants to be free" and "infringement is not theft; being deprived of speculative profit is not a loss."
How the turn tables.
He's lead Metallica through the "Copy our tapes! Let everyone hear our music!" early days, through the "They're stealing our master recordings!" Napster and Senate hearings era, and come all the way back 'round to "I’m just happy that fucking anybody cares about what we’re doing and shows up to see us play and still stream or buy or steal our records or whatever.” [https://consequence.net/2023/09/metallica-lars-ulrich-stream...]
2 out of 3 ain't bad, I guess.
"Information wants to be free as long as I benefit from not paying."
I "just" compared prior experience and current experience specifically.
Who's to say? I'll check back in 10 years :)
At any rate, I ought to pick it up again and finish reading it...
this is a complicated way of counting how many of each word is in the book.
clearly fair use.
does the author think cutting up a book is a copyright concern? theyre buying the books, its up to them what to do with their copy. if you want the books preserved, maybe fund your libraries to get a copy or two?
Not the case here, but it's also important to keep in mind that laws too, can be unjust and immoral. Much of humanity's progress has involved the abolishment of immoral laws.
This is the whole problem. A lot of people assuming they can speak unilaterally and with authority on what other people value or should find valuable.
It's disgraceful. What happens to the old hacker ethos? I'd never thought I'd see the day when, on a site called hacker news the majority of posters are defending the right of a large corporation to destroy public goods permanently just to be able to charge people more for "intelligence as a service".
So you are saying to abolish copyright?
Then setup a fund to compensate copyright holders.
I guess publishers do have a right to pull certain works from print though.
For a refresher: The Internet Archive scanned books then shared them online, trying to claim that converting them into digital copies qualified as a derivative work. You don’t have to be a lawyer to see how that claim doesn’t hold up to any scrutiny.
The LLM companies are not redistributing the works. They are using them for training. This blog post calls it IP theft, but it has actually been litigated in court already. Using books for training does not qualify as theft or redistribution, even though some people have different opinions about the moral angles.
One of those same lawsuits also extracted a huge settlement from Anthropic for using digital downloads from pirate sites. The conclusion was that the only acceptable way for them to use the books is to buy them and scan them. The courts forced it to be this way.
The current hand-wringing about the destruction of books is based on claims that they’re doing it to rare books that are also valuable. So far nobody has been able to actually provide an example of one of these books that is supposedly ultra-rare, but also valuable, but also only available at one of these book sellers that sells these things in bulk. We’re supposed to assume that one of these books might actually be super valuable but also super rare and also only available at these places they’re buying from.
Lastly, there's a misguided belief among some that it's somehow less of a copyright violation if the source is destroyed, but it's a violation either way unless the entity doing the scanning has permission from the copyright holder to make the copy.
For the record, i don't much have issue with llms and such.
But... the wholesale destruction of printed books is a bridge too far.
There’s no giant “used” section at Barnes and Noble. Libraries only have so much space and they have to cycle through books as the years go by. This is easily the best fate these books could have received.
Google Books tried. They lost. If you want this fixed, talk to your congressman, not Amazon.
Scanning a legitimately purchased book is IP theft? How can he hold such a copyright-maximalist view, and at the same time defend the Internet Archive?
Your granpa once wrote an obscure book about this amazing way he'd found to cure diabetes.
Corporation A buys all existing copies of the book, scans them, destroys the originals, and sets up a commercial business offering a monthly subscription to alleviate diabetes pains with this new method they claim they discovered.
Person B borrowed the book from a municipal library, Xerox'd it, and keeps a free ledger, open to all who want to read old books, as a way to safeguard free access to the world's knowledge.
Do you think there could exist any possible logic by which some people would defend person B and try to stop Corporation A?
They're intended to be analogous of reality.
Which real-world company is said to be buying all existing copies of a book, here?
So it reduces to a corporation selling services based on public-domain knowledge (I am ignoring the part where a miraculous advance is confined to a single unknown book). Lots of corporations do this, and there's nothing wrong with it.
I'm sure people would prefer if Amazon also offered library-like access to the book, but how is it theft? I'm not asking about "any possible logic" - he called it theft.
That is also my understanding of the claim, and I would intuitively agree with it.
Ok, but it's a bit different because they are not selling "the book" they are selling a fundamentally different thing for which the book is consumable input. Let's consider a batch of 5000 books. For kicks let's assume the known copies are N=1 for those books. Today, though only a few people can do this at a time, a person can purchase one of those books once read exactly the text in that book, and keep doing so to their heart's content. Or maybe a library buys it and now a rotating legion of people can do that for free. Now let's say amazon buys and scans them, destroying them in the process and refusing to release scans to avoid giving competitors a training edge. Now:
- You can never get the information as written again. At best you'll get an LLM output approximation/mutation of it. If this was the expression of a real human beings lived experience, that's kind of sad and goes against one of the spirited aspects of human existence, to leave a legacy, doesn't it? - Amazon is going to charge you per token every time you want to access that information. - Price demand for the information is now tied up with general demand for LLMs rather than the actual book, either reducing or greatly inflating the cost to you in addition to the now recurring charges.
So, now you (a) can't actually ever access that book as it was written (b) need to pay continually to access an approximation of its contents (c) and possibly more than the book is worth since now it's "value" in a price sense has been absorbed into general llm inference costs.
Not to mention you may have eradicated the last extant copy of a person's memoirs, but I guess it's pretty clear people in this industry don't care at this point. I hope one day your entire life story is ground up and consumed in some data farming operation and you are all summarily forgotten.
"AI training is infringement" is not exactly a copyright-maximalist view. The explicit training task used for pre-training is reproducing the content of the trained-on books; and models trained on such books are able to reproduce significant infringing chunks of them[0] unless specifically post-trained to refuse to do so.
Additionally, they might have thought that Controlled Digital Lending was OK (it wasn't, but that's a different issue to AI training). As I've mentioned elsewhere in this thread, there's a common misconception that copyright is concerned with the number of copies in circulation as opposed to individual acts of copying.
Or they don't care about any of that and just wanted to highlight the hypocrisy.
[0] Which, under the "compression is intelligence" point of view, is entirely expected and not surprising in the slightest.
https://news.ycombinator.com/item?id=49310725
https://news.ycombinator.com/item?id=49068738
It does though. A lot of people seem to think "it's a book!!!" means it takes on some mystical intrinsic value. That's bullshit. There's an absolute mountain of worthless and near-worthless books. I bet the vast vast majority of these books are those books - after all that's what Amazon wants. A ton of human written text as easily as possible. They obviously aren't buying first editions of Oliver Twist, or even second editions of Harry Potter.
> We’re not revealing the titles of the books included in the shipment we tracked
Yeah... because then it would reveal how unimportant they are. This sort of outrage inflation is counter-productive. People see through it, lose trust, and then when something bad actually happens they won't believe you.
https://en.wikipedia.org/wiki/Outrage_industrial_complex
And come on. They're not revealing the titles because that would reveal which bookseller they worked with.
I've read a lot of unknown old literature that has little to recommend it as far as literary history goes, but it all has a certain charm, and at the end of the day it was all the work of a human being who walked this earth just like we did and captured their thought on the page.
What if their family has an interest in the preservation of that text? What if they simply don't yet know of it, or couldn't afford it?
Yes all hypotheticals, but I think a lot of people are underestimating the potentially permanent damage being done to the public good just because the actions are not illegal.
Everyone should be uncomfortable with a single company unilaterally pillaging the public good for its own ends. Period.
IDK about you geniuses on hacker news, but I'd much prefer to pay a few hundred dollars for a rare book once to be able to read it than pay amazon token costs monthly to get the LLM's chopped and screwed regurgitation, without any recourse to even knowing where the information comes from.
Individually you might be right. Some cookbook from 20 years ago probably is worthless to most people. But then there'll be one person who spends an afternoon hunting it down, bookstore to bookstore, because their aunt used to have a copy and they want to check a recipe, or some such.
This is cultural violence on a massive scale.
I just searched again now and found 3 used copies on AbeBooks for about US$11 incl. shipping. Absolutely made my day.
I don't know if adult me will enjoy it beyond the nostalgia, but there's a coconut cake recipe in there that I will never forget.
Like you said, just about nobody on planet earth would care if the billionaires ingested and then destroyed those final copies, but I'd have been gutted if they'd beat me to it.
quantity has a quality of its own
>Some cookbook from 20 years ago probably is worthless to most people. But then there'll be one person who spends an afternoon hunting it down, bookstore to bookstore, because their aunt used to have a copy and they want to check a recipe, or some such.
ask your aunt for the recipe, look it up online with a search engine, ask an LLM by describing how it tastes and the ingredients you remember, this is really not that complicated
>This is cultural violence on a massive scale.
could you be any more dramatic? you're acting like we are actually losing anything when in reality we're gaining from this, you'll just be able to ask the AI for the recipe instead of wasting hours and hours hunting down some old book and ordering it online for an exorbitant amount
Your example is very flawed because if the LLM gave you the exact recipe then it would be infringement, but many recipes need to be exact. LLMs aggregate information so wanting a very specific recipe (to the point of hunting down an old cookbook) is a non-starter because it the LLM is really only guessing.
If you write: "Smash one egg into that bowl and feel the power. Sprinkle salt onto it salt bae style." That's copyrighted and I can't copy that word for word.
But I can publish my own recipe which is an exact copy of yours but written like so:
Ingredients: 1 egg 1 teaspoon salt
Method: Combine egg and salt.
Or write that in any way I want that isn't a word for word copy of your recipe.
The whole problem here is:
- Not everyone agrees on the utility / value of AI. - A small handful of companies control frontier models and get to shape how these inputs are ultimately put to use. - As Sam Altman put it, they want to charge you, in perpetuity, for "intelligence" as a service. Contrarily, anyone can access a book usually for a one time purchase fee and read it forever. You'll be paying per token to look up auntie's recipe every time instead of just, you know, reading the book.
Not to mention you are erasing the voices of human beings who had lives and may have related their experiences and emotions in those texts. That anyone would be comfortable with a corporation they know nothing about destroying potentially unique information of the lives and stories of real human beings, locking it away forever as one granular component of a stochastic content generator, the human expression in its actual form never to be accessed again, is absolutely appalling to me.
I think you should reflect more on who will really benefit, who has the control, and why humanity as a collective is potentially losing as collateral, without having much say about it.
Do you even hear yourself?
These behaviours would be described as obsessive and pathological in any clinical setting, but when they're done for profit it's somehow normal.
Just like China does not have communism. They have too much free-market activity and private ownership for it to classify as such.
If you look at China and think "That's not communism", just be aware that they may also look at the west and think "That's not capitalism." Both are right. In fact both systems have converged.
China took the better aspects of communism and capitalism. We took some of the worst aspects of both... Except maybe free speech; but people are working hard on removing that one!
It makes sense; China came out of an era of extreme government control which has been reduced over time (compared to what it was before). In the west, government control has been on the way up and we can't even imagine how bad it can get.
Wage labor is then the perfect structure to ensure that most resistance from the bottom is suppressed before it can even really start. When you depend on the corporation to succeed for your subsistence (since there is no adequate social safety net to fall back on) taking a principled stance and refusing to destroy great works of literature becomes a highly costly endeavor for most people.