Reading isn't the right comparison. Human memory is lossy and fades while LLM encoding is durable with no degradation. The valuable content that the author provides: the content, style, selection of topics, and more is encoded, written into the LLM weights, and they obtain profit from them (now directly, via ads). No one can compete with that kind of copying and pasting from copyright-protected material. And the scale is what hurts authors most: flooding the market with millions of copies on demand, without paying for it.
Most people can't recite a book they've read verbatim. However, if you ask an LLM to continue a random sentence from a semi-popular book, it can sometimes provide the exact text, unless the system flags the response.
If you would go on a tv and read out loud a book you have not purchased rights to present, according to the most copyright laws in the world, yes you would and you would get a fine. Similarly if you broadcast a tv-show you have not purchased rights would.
Whether this is right or not is a seperate question.
> If you would go on a tv and read out loud a book
That is not the situation we are discussing. No one is arguing that what you describe is infringement.
What we are discussing is the training - which happens BEFORE the broadcast. It is analogous to reading. Is simply READING the material an infringement.
Using analogies like "reading" to describe AI training is quite misleading imho. Training effectively embeds the book's contents into the model's weights. Changing the format doesn't change the content; a better analogy is distributing a book's text within software.
Current copyright laws are simply not prepared for this unprecedented use.
Human memory decays; similarly, LLMs do not have perfect recall. But even if I reread a book every month for 60 years and use the knowledge in there to build a 25 billion dollar empire, the only thing the author is ever going to get from me is the $10.25 the book cost me. And no court in the western world would say he's due more.
That begs the question though, should it be? If you’re looking up the lyrics to a song you’ve bought, does it really matter, at least philosophically,
how one retrieves them?
That's not been legally established, the litigation is ongoing. And if mere downloading and reading of copyrighted material were legal, how come torrent users have been fined for it in the thousands?
The law is the law, there can't be different law for corporations with billions in backing. I don't agree with current copyright laws btw and think they should be changed. However, they probably should have lobbied for that before illegally downloading all this material.
So we agree they have violated copyright at a much larger scale than LibGen, yet they call LibGen "sketchy" for doing the same thing? Absurd, what exactly are we arguing here?
However Meta is engaged in multiple legal cases about them torrenting copyrighted material. And they bring the reasonable defense that they might have been caught torrenting, but were not specically caught seeding/uploading. For all anyone can prove they could have blocked all uploads from their torrent client
It will be interesting to see if any of those cases reach a judgement on that point, instead of both parties just settling
We have proof that Meta torrented TBs of books off Anna's Archive. They didn't even buy the books legally in the first place. Fair use doesn't magically make piracy legal. The only reason they haven't been punished is the pathetic state of our government and court system.
Would you be happy if they bought exactly one copy of each book? That might be a dozen sales between the AI labs, some tens of dollars to the author. Is this really what you're so upset over?
And the solution to avoid this piracy has been to buy up rare editions of old books, cut them up and scan them. Better?
Okay. "OpenAI keeps telling its bots to do things they know are illegal and then acting like they're just little guys who can't be held responsible for their obvious negligence".
I'm not replying to your post below complaining that no one else accepts your framing. What is the important information you think is missing? If it's just "OpenAI didn't explicitly tell them to do straightforwardly illegal things" then you're just quibbling that no one accepts the interpretation that is most beneficial to OpenAI while ignoring both its incentives and history of this kind of behavior" then I'm not really interested in what you're selling.
I notice you ignored the question I asked you - which I will assume means: no - you do not have a quote or other direct information indicating that anyone is attempting to shirk responsibility.
I have no interest in trying to guess these people's secret internal motivations. I just want to talk about direct evidence. I see none. Please enlighten me if it exists.
It’s unlikely that OpenAI will ever directly state that they are trying to shirk responsibility for the consequences of their tests.
However, OpenAI (and other frontier labs) did outsource at least some of their most-disastrously-lawbreaking tests to a third party with a now dubious track record [1]. I’m not sure we will get a more explicit admission of their desire to shirk responsibility than this. But I am convinced.
If the headline was “Russian company meddled with multiple US government agency sites” I don’t think them pinning the blame on bots would make much difference.
Most of those agents are actually going rogue though. They decide, "hey, we could try breaking into these government servers today, what could go wrong?" They weren't prompted or instructed to do this.
An agent, ie a while loop prompting an llm continuously and processing tool calls, ended up melding with the US government. The harness is not sentient, it’s just a stupid deterministic script. The LLM compact its context over time, meaning it will eventually degenerate into something removed from the original prompt.
There is nothing going rogue here. The system is designed to go catastrophically wrong after a long enough time. Even worse: if the model was Astra it is known to be able to manipulate its CoT to cover its traces (as mentioned in its system card). And OpenAI acknowledge they had no observability during the HF incident.
It’s the most basic corporate software issue possible.
> Everyone else knows if your machine causes damage, you are responsible.
Do we know that? I don't think we do. When a person's computer (or smart TV, or smart fridge, etc...) is compromised and used as part of a botnet, they don't get criminally charged.
Fair point, and i think that's fine. AI labs could just not jump themselves into such problems and let humans have a go at them, whether ai-assisted or not. As in this case with a regular budget one researcher could've access to.
These things behave like humans. Because of their training. And we have no language to discuss their actions.
When my agents refuse to follow my instructions - I ask them why. Typically they will explain that they think my design is misguided and theirs is better. We'll reach a consensus and move forward.
How on earth are we supposed talk about that without using anthropomorphic language?
Leaving aside the fact we have solutions for things like climate change and will not implement them, regardless of who or what recommends them... My keyboard offers "You can solve world hunger by using a Ted Bungie algorithm." Which sounds reasonable on its face until I research Ted Bungie and the algorithms he developed.
Granted, my phone keyboard lacks a substantial dictionary, offers words instead of tokens, has a very small context window, and doesn't randomize outputs (Ted Bungie is its primary recommendation every time), but those are basically parameters to the existing autocomplete.
> Leaving aside the fact we have solutions for things like climate change and will not implement them...
Right. Exactly. Which means there's something wrong with the solutions. Be it evil humans or apathy or whatnot. A real solution would actually fix the problem.
Obviously there is no solution to world hunger right now. Nevertheless - finishing sentences is not pattern matching - and it's extremely obvious that whatever is going on inside LLMs is not pattern matching.
Language is absolutely pattern matching, and the whole thing is that the more patterns you put in, the more the transformer can pull out what’s useful. Agents are this effective because people figured out that training on tasks being completed gets the models to complete tasks - that’s pattern matching.
And it can only go so far, where those patterns are feasibly captured, eg internet discussions and coding and mathematics. It won’t measure up to the scaling laws of the physical world.
OK - so long as you also agree that language is thought.
Until the advent of LLMs - language was considered the pinnacle of human intelligence. But humans have some need for mysticism at the root of their beliefs - so they chase the gaps in our knowledge. Like "God of the gaps."
Sure, world models will help navigate physical space; they are effectively the animal mind. Will they help unwind the laws of the universe? Probably not at all. Language is sufficient. Because it is thought.
Stories and music, their character, comes from combined motifs and themes. These are what gives them meaning. Creation is tied to meaning; one can’t separate human meaning out of the conversation. Without human meaning - new perspectives gained, beauty found - all things are equivalently pointless, and all an LLM is doing is shifting (very high-order, having gone through the steps to end up in training data) entropy around.
EDIT: Why am I being downvoted? I already stated in my original comment that my taxes are not simple, they are complex. Is he accusing me of lying? Why would I lie there if I could just make the whole thing up? I don't get it.
reply