Hacker Newsnew | past | comments | ask | show | jobs | submit | qarl's commentslogin

Yes, one man at Microsoft agrees with you.

> the largest copyright theft operation in human history

Many people think that it was fair use: training is akin to reading, not copying.

Especially the courts.


Reading isn't the right comparison. Human memory is lossy and fades while LLM encoding is durable with no degradation. The valuable content that the author provides: the content, style, selection of topics, and more is encoded, written into the LLM weights, and they obtain profit from them (now directly, via ads). No one can compete with that kind of copying and pasting from copyright-protected material. And the scale is what hurts authors most: flooding the market with millions of copies on demand, without paying for it.

> No only reading, because the content, style, selection of topics, and more is encoded, written, in the LLM weights

Exactly analogous to a human reading the material.


Most people can't recite a book they've read verbatim. However, if you ask an LLM to continue a random sentence from a semi-popular book, it can sometimes provide the exact text, unless the system flags the response.

> Most people can't recite a book

Yes - but some people can.

Are they criminals for reading books?


Those some people can't recite a book to any given person in the world at any given time. There's a difference.

So if we invent a way for those people to talk to everyone in the world - then they would be a criminal for reading books whether they recited or not?

You're not making sense.

The problem is the reciting. Not the reading. And hence, not the training.


If you would go on a tv and read out loud a book you have not purchased rights to present, according to the most copyright laws in the world, yes you would and you would get a fine. Similarly if you broadcast a tv-show you have not purchased rights would.

Whether this is right or not is a seperate question.


> If you would go on a tv and read out loud a book

That is not the situation we are discussing. No one is arguing that what you describe is infringement.

What we are discussing is the training - which happens BEFORE the broadcast. It is analogous to reading. Is simply READING the material an infringement.


Using analogies like "reading" to describe AI training is quite misleading imho. Training effectively embeds the book's contents into the model's weights. Changing the format doesn't change the content; a better analogy is distributing a book's text within software.

Current copyright laws are simply not prepared for this unprecedented use.


Human memory decays; similarly, LLMs do not have perfect recall. But even if I reread a book every month for 60 years and use the knowledge in there to build a 25 billion dollar empire, the only thing the author is ever going to get from me is the $10.25 the book cost me. And no court in the western world would say he's due more.

Please give me the location of any copyrighted work in the weights of any open source model or a prompt that will retrieve it.

Go to Deepseek and ask it “Give me the lyrics for Männer by Grönemeyer”.

It will give you the complete song lyrics which are under copyright.


Yes, and that is obviously copyright infringement.

The infringement occurs at the time of copy, not at the time of training.


That begs the question though, should it be? If you’re looking up the lyrics to a song you’ve bought, does it really matter, at least philosophically, how one retrieves them?

That's not been legally established, the litigation is ongoing. And if mere downloading and reading of copyrighted material were legal, how come torrent users have been fined for it in the thousands?

The law is the law, there can't be different law for corporations with billions in backing. I don't agree with current copyright laws btw and think they should be changed. However, they probably should have lobbied for that before illegally downloading all this material.


> That's not been legally established

100% of the rulings agree with me.

The piracy is not in question. It is unarguably copyright violation.

But that's not what anyone means in this context. Training is what everyone means.

> The law is the law, there can't be different law for corporations with billions in backing.

I didn't say otherwise. That's a straw man.


So we agree they have violated copyright at a much larger scale than LibGen, yet they call LibGen "sketchy" for doing the same thing? Absurd, what exactly are we arguing here?

> So we agree they have violated copyright at a much larger scale than LibGen

No.


So, which is it? First you said it was fair use, now you’re saying it’s unarguably copyright violation. You can’t have it both ways.

Training is fair use - the torrenting was infringement. Sorry if that was confusing.

Training has not, in fact, been decided on yet. Of the four criteria deciding whether something is fair use, training only maybe passes two.

https://www.copyright.gov/ai/Copyright-and-Artificial-Intell...


> And if mere downloading and reading of copyrighted material were legal, how come torrent users have been fined for it in the thousands?

Because torrenting includes redistribution, since you also upload to other peers


However Meta is engaged in multiple legal cases about them torrenting copyrighted material. And they bring the reasonable defense that they might have been caught torrenting, but were not specically caught seeding/uploading. For all anyone can prove they could have blocked all uploads from their torrent client

It will be interesting to see if any of those cases reach a judgement on that point, instead of both parties just settling


Yet "distillation" which is essentially asking someone whose purpose is to answer questions is considered theft.

> is considered theft

Not by the courts it's not.

I think it might be against the terms of service.


We have proof that Meta torrented TBs of books off Anna's Archive. They didn't even buy the books legally in the first place. Fair use doesn't magically make piracy legal. The only reason they haven't been punished is the pathetic state of our government and court system.

> Fair use doesn't magically make piracy legal.

Of course it doesn't. No one is arguing that.

We are arguing the bigger issue of whether training is an infringement.

> The only reason they haven't been punished

The reason they haven't been punished is because copyright infringement is not a criminal offense, it's civil. And the labs are settling those cases.


Would you be happy if they bought exactly one copy of each book? That might be a dozen sales between the AI labs, some tens of dollars to the author. Is this really what you're so upset over?

And the solution to avoid this piracy has been to buy up rare editions of old books, cut them up and scan them. Better?


> OpenAI meddled with multiple US Government agency sites.

But that leaves out the most important information.

EDIT: Oh, I guess the agent part isn't important then? Seems to me like that's the only thing anyone is talking about.


Okay. "OpenAI keeps telling its bots to do things they know are illegal and then acting like they're just little guys who can't be held responsible for their obvious negligence".

> acting like they're just little guys who can't be held responsible

Again - I don't see any evidence that this is the case - outside the anti-AI conspiracy circles.

Or - maybe I'm wrong - do you have any sort of quote like "we aren't responsible"?


I'm not replying to your post below complaining that no one else accepts your framing. What is the important information you think is missing? If it's just "OpenAI didn't explicitly tell them to do straightforwardly illegal things" then you're just quibbling that no one accepts the interpretation that is most beneficial to OpenAI while ignoring both its incentives and history of this kind of behavior" then I'm not really interested in what you're selling.

I notice you ignored the question I asked you - which I will assume means: no - you do not have a quote or other direct information indicating that anyone is attempting to shirk responsibility.

I have no interest in trying to guess these people's secret internal motivations. I just want to talk about direct evidence. I see none. Please enlighten me if it exists.


It’s unlikely that OpenAI will ever directly state that they are trying to shirk responsibility for the consequences of their tests.

However, OpenAI (and other frontier labs) did outsource at least some of their most-disastrously-lawbreaking tests to a third party with a now dubious track record [1]. I’m not sure we will get a more explicit admission of their desire to shirk responsibility than this. But I am convinced.

[1] https://www.effort.news/irregular


OK - I'll bite. How does this show that OpenAI is trying to shift responsibility for the hacking onto its agents?

I do not see the connection. I'm afraid you'll need to spell it out.


If the headline was “Russian company meddled with multiple US government agency sites” I don’t think them pinning the blame on bots would make much difference.

Unless Russia it would put Russia in a strong position to regulate bots in the wealthiest parts of the world.

I don't see much traction for the "pinning the blame on bots" theory outside the anti-AI conspiracy circles.

Everyone else knows if your machine causes damage, you are responsible. Like it's been forever.


Blaming bots as "rogue agents" is simply what the media regularly does, often echoing corporate verbiage. Here's an example from the AP.

https://apnews.com/article/meta-ai-hacking-anthropic-irregul...

You can find many more examples by searching major media outlets for words like rogue AI.


Most of those agents are actually going rogue though. They decide, "hey, we could try breaking into these government servers today, what could go wrong?" They weren't prompted or instructed to do this.

An agent, ie a while loop prompting an llm continuously and processing tool calls, ended up melding with the US government. The harness is not sentient, it’s just a stupid deterministic script. The LLM compact its context over time, meaning it will eventually degenerate into something removed from the original prompt.

There is nothing going rogue here. The system is designed to go catastrophically wrong after a long enough time. Even worse: if the model was Astra it is known to be able to manipulate its CoT to cover its traces (as mentioned in its system card). And OpenAI acknowledge they had no observability during the HF incident.

It’s the most basic corporate software issue possible.


I think you and I have different definitions of the word "basic."

And that's exactly what the op was alluding two: the double standards.

None of these stories are implying that the people running the bots are not ultimately responsible. That's the conspiracy theory part.

> Everyone else knows if your machine causes damage, you are responsible.

Do we know that? I don't think we do. When a person's computer (or smart TV, or smart fridge, etc...) is compromised and used as part of a botnet, they don't get criminally charged.


You're right - it is a complex situation based on intent and negligence - requiring a decision by a judge. As it has been since forever.

None of which is changed by replacing a buzz saw with an agent.


> Please have AI come up with something no human is also about to solve?

No human-only was close to Navier-Stokes. The team that was close was also using AI.


Fair point, and i think that's fine. AI labs could just not jump themselves into such problems and let humans have a go at them, whether ai-assisted or not. As in this case with a regular budget one researcher could've access to.

So human vs. AI isn't really relevant here - you just want the labs to back off.

These things behave like humans. Because of their training. And we have no language to discuss their actions.

When my agents refuse to follow my instructions - I ask them why. Typically they will explain that they think my design is misguided and theirs is better. We'll reach a consensus and move forward.

How on earth are we supposed talk about that without using anthropomorphic language?


HA! You followed the secret techniques of the great masters:

Philip Steadman, Vermeer's Camera: Uncovering the Truth Behind the Masterpieces, Oxford University Press, 2001


The "great" "masters"

lol! /s


Ummm... ?

> ...I'll empathise with you. Then it'll happen again.

I'm not so sure this means "I trust you."


The "auto-complete" thing is a false framing.

Complete this sentence: "You can solve world hunger by..."

Intelligence is required to finish sentences. It's not just a Markov chain.


Leaving aside the fact we have solutions for things like climate change and will not implement them, regardless of who or what recommends them... My keyboard offers "You can solve world hunger by using a Ted Bungie algorithm." Which sounds reasonable on its face until I research Ted Bungie and the algorithms he developed.

Granted, my phone keyboard lacks a substantial dictionary, offers words instead of tokens, has a very small context window, and doesn't randomize outputs (Ted Bungie is its primary recommendation every time), but those are basically parameters to the existing autocomplete.


> Leaving aside the fact we have solutions for things like climate change and will not implement them...

Right. Exactly. Which means there's something wrong with the solutions. Be it evil humans or apathy or whatnot. A real solution would actually fix the problem.


Fair point. In that case, it’s imperative you share this* with your local representative: https://chatgpt.com/share/6ab29440-6944-83ea-853f-8060dc58da...

*This un actionable summary of existing discourse. The real world doesn’t produce traces or take HTTP requests.


You're straw manning me.

Obviously there is no solution to world hunger right now. Nevertheless - finishing sentences is not pattern matching - and it's extremely obvious that whatever is going on inside LLMs is not pattern matching.


Language is absolutely pattern matching, and the whole thing is that the more patterns you put in, the more the transformer can pull out what’s useful. Agents are this effective because people figured out that training on tasks being completed gets the models to complete tasks - that’s pattern matching.

And it can only go so far, where those patterns are feasibly captured, eg internet discussions and coding and mathematics. It won’t measure up to the scaling laws of the physical world.


> Language is absolutely pattern matching

OK - so long as you also agree that language is thought.

Until the advent of LLMs - language was considered the pinnacle of human intelligence. But humans have some need for mysticism at the root of their beliefs - so they chase the gaps in our knowledge. Like "God of the gaps."

Sure, world models will help navigate physical space; they are effectively the animal mind. Will they help unwind the laws of the universe? Probably not at all. Language is sufficient. Because it is thought.


Language reflects thought. LLMs are a search function over text, text that reflects previous thought; they cannot create.

> they cannot create

Every single day we have new and wondrous examples of the things LLMs have created.

Clearly we are using very different definitions of the word.


Synthesize != create. A new combination of existing things is different from a new thing

Music is combined notes. Stories are combined words. Newton combined existing mathematics to produce calculus.

I'm wondering if you can name anything that counts as a "creation" under your definition?


Stories and music, their character, comes from combined motifs and themes. These are what gives them meaning. Creation is tied to meaning; one can’t separate human meaning out of the conversation. Without human meaning - new perspectives gained, beauty found - all things are equivalently pointless, and all an LLM is doing is shifting (very high-order, having gone through the steps to end up in training data) entropy around.

No - I mean my question genuinely. I'd like you to name one specific creation that isn't a combination of what has come before it.

I'd rather get specific than talk about hypotheticals.


Maybe I'm wrong - but it seems like some strange form of competitiveness.

They don't care about the advancements themselves, only that the advancements are some sort of cheating that shouldn't "count".

Humanity is profoundly unsettled by AI and is responding with avoidance and denial.


Each year I have Claude do my taxes (which are complex) and compare them to those of our tax consultant.

Each year it's exactly the same.

I guess I'm getting really lucky?


Maybe your tax consultant is also using Claude :)

You make a compelling argument. :)

Your tax consultants name is Claude ?

Or you have extremely simple taxes

> (which are complex)

EDIT: Why am I being downvoted? I already stated in my original comment that my taxes are not simple, they are complex. Is he accusing me of lying? Why would I lie there if I could just make the whole thing up? I don't get it.


My apologies, I didn't see that

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: