Hacker Newsnew | past | comments | ask | show | jobs | submit | hexomancer's commentslogin

So you definitely did train on their data, you just think it is unlikely that it impacted the final model significantly?

I have no idea if their data was trained on. For example, if they used ChatGPT, asked a math question, and clicked the thumbs up button, that could have provided a small reward signal. I highly doubt this sort of feedback made a difference to a problem like Navier-Stokes, but it's not something that's feasible for us to prove one way or the other.

Edit: Also, if they opted out of training, then we didn't train on it.


> it's not something that's feasible for us to prove one way or the other.

This kind of question is exactly what a company named _Open_AI and founded as a nonprofit is supposed to be doing; open research on AI that helps inform, rather than obscure.

Anyhow, you do have the data available about the documents in the user's accounts, what they opted into (or were forced into via non-negotiable ToS), and whether they pressed a "thumbs up" button. You can answer whether the data entered the training pipeline or not. Yes, how much influence it had is an open question, and one that would be good to have research on and better tools for exploring, but I'll accept that it can't currently be answered precisely.

But whether the data entered the trianing pipeline can be answered. And how to provide better tools for quantifying and tracing this kind of thing is exactly what should be studied.


I think it should be incredibly easy to verify this. Just look at the training data and see if it contains any of the chats. It should be trivial for a company with tens of thousands of super-genius agents at their disposal.

Two steps would be needed.

(1) We'd have to identify their chats. How would we do this? We'd need them to share their chats with us so we could look for matches.

(2) We'd have to prove those chats changed model behavior. How would we do this? We'd need to retrain many models with those specific chats removed, and ask those models to solve the Navier-Stokes problem many times, and keep doing this until reaching the desired level of statistical significance.

#1 requires their cooperation and a bit of work on our side. #2 is extremely expensive and not really feasible.


> (1) We'd have to identify their chats. How would we do this? We'd need them to share their chats with us so we could look for matches.

According to the statement by Tristan Buckmaster, he was in communication by email and calls several times over the past week with you (OpenAI that is, not you personally), asked about whether his chats were trained on, and was declined an answer (https://cims.nyu.edu/~tristanb/statement.pdf).

However, it seems like there was great pressure to hurry the release to compete with Anthropic's recent release, so he was unable to get an answer in time.

The mealy mouthed statement in the release "We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ." is realy not much. If OpenAI had wanted to be transparent about this, you could have worked with him to identify if his data was used in the training of your new model, and actually made a somewhat more certain statement on that basis. But you have chosen not to; it was more important to scoop Anthropic on this than it was to be transparent about your training data.

> (2) We'd have to prove those chats changed model behavior. How would we do this? We'd need to retrain many models with those specific chats removed, and ask those models to solve the Navier-Stokes problem many times, and keep doing this until reaching the desired level of statistical significance.

Just the information from step (1) would improve transparency. Yes, you still can't prove one way or another how much the effect of the training is. But if it's included in the training data, it provided some effect.


> we'd have to prove that firing the gun caused the murder. how would we do this? we'd need to redo the murder many times, with and without my client firing his pistol. that's extremely expensive and not really feasible. therefore, we must acquit.

#2 (prove those chats changed model behavior) is pretty straightforward if the anonymized data from chats can be actively searched by a model. In fact, it could be very clear if the provenance of context is traced. If anonymized data from chats leak into the context of an actively running model it would clearly influence the answer.

Just because something is in the training data, doesn't mean it is the root of an LLMs output.

Turn off web search and ask a model what a random redditor said about a random topic in 2015. You will only get hallucinations at best, even though that comment is definitely in the training set.


Sure. But it's possible to say: if the document isn't in the training data, it isn't the cause of the output. If it is in the training data, the question gets more complicated.

What they're saying, and I think this was the clear implication of the blog post too, is that the training data definitely would contain these chats and the only question is whether it got encoded into the weights.

Presumably, given that you also operate in the EU, you would have asked for their explicit consent before you did, so you could just check for that?

That’s also what I understand. If true yet another disgusting behavior from the company

> On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved

What's the other one?


I heard Hodge conjecture? Third-hand rumor though...

I assume the rumor is a counterexample? Where do I go to get wind of these rumors?

I think very soon big LLM providers (OpenAI et. al.) will provide a service like this. With the LLMs becoming faster, the agentic task completion bottleneck will soon move to the tool calls (both execution time and round-trip latency), so it makes sense to have a server host the project very close to the actual LLM doing the inference in order to minimize latency.


Also, they can literally train the LLM inside the VM they create so it becomes an expert in whatever tools are there. And since they control the entire sandbox, less permission issues are likely.

I think they will become extremely vertical integrated sooner or later because that will be the most efficient and effective way to do this. I'm guessing anyone building this kind of thing now has only one hope: acquisition.


Anthropic has Claude Code Web -- sorry, looks like the official name is "Claude Code on the web" -- which I think is basically this? I've been using it for a while and I really like it.

Nothing to do with latency, I think, it's all about convenience and safety. I don't have to worry about running an agent on the computer that knows all my admin API tokens and SSH keys, and the agent isn't interrupted if I close my laptop.

The one thing that might be useful is the ability to run on a beefier machine when you need it, e.g. make sure it has direct access to a GPU. But then you're back to thinking about the host VM, not fully abstracting away from it.


OpenAI already does this with Codex cloud

https://chatgpt.com/codex/cloud

From my experience, Orbs have been more consistent, along with Cursor's agents


This is very cool, I wish there was some way to use it on a bicycle though. For example, when moving into a street it could ask (using voice) if this street is paved, and I could answer it using voice too.


I have also been using it for a couple of months and it is great. Most other software that try to emulate a tiling window manager in windows/linux end up being too buggy and annoying to use full time, but in my experience aerospace is the first one that I have been able to run full time with no issues.


Yeah I got exited thinking this is about traffic lights. I use a bike to commute to work and recently I was thinking if I could adjust my cycling cadence so that I never hit a red light, but unfortunately the timing of the traffic lights in my city is not constant. If there was a publicly accessible API to get the current timing info, I could write an app to do that.


If you're in America, take a look at the strobe on top of school busses. I'm not sure if they still have them (they used to). It would flash at a specific frequency and trip a photovoltaic sensor connected to the traffic light, which would turn it green so the kids aren't late for class. If you had a bright enough strobe which flashed at the same frequency...you get the idea.


Is that actually true? I've heard of ambulances & police cars having such devices, but they were supposed to be infrared.

The last time I saw the strobe on top of a school bus active, it was when I was a passenger in one, driving down the freeway at night, and it wasn't strobing particularly fast. It's possible that our driver just forgot to turn it off, I suppose - he was that kind of guy.


School buses in my state are legally required to run the strobe when passengers are onboard.

No two strobes I have seen strobe at the same frequency. I think this traffic control story is urban legend.


I never heard about this being used on school busses. This was always something for emergency services like firetrucks/ambulances to not have to sit in traffic at a red light, but it was only active if they were actively responding to a call with their lights on. Otherwise, they sit at the lights too.


A newspaper article told of a mayor of some city that had one installed so he could zip along to emergencies.


Emergency vehicles have devices that announce their presence to get traffic lights to change in their favor. “Kids being late to class” is not on the order of importance to create a complex scheme to change traffic lights based on strobe lights from a bus.

Sounds like urban legend.


Bus priority lanes and traffic lights that give priority to busses are definitely a thing. Usually for municipal busses and not school busses, but I'd expect a community that had priority lights for busses would allow school busses onto the system as well.

Not specifically to avoid late arrivals of pupils, but because prioritizing many passenger vehicles is valuable.


We definitely have this system in place in some cities in Canada, primarily for express bus routes.


So as a driver, you want to follow an express route bus when you can?


That's so cool, I was actually curious about history of such programs. Thanks for posting this :)


I doubt this comment was in good faith (you decided to ignore literally all the features I mentioned and focused on just creating files) but I am going to reply anyway:

1. There is no way that `touch newfile` is faster. Using voil, you press a keybind, enter `newfile`, save and you are done. Using touch you have to first, use some keybinding to switch to terminal, then type `touch ` (6 letter overhead) then type the name of the file and then switch back to vscode. I am not saying voil is meaningfully faster, but you saying that `touch newfile` is faster is wild to me.

2. If I am editing a comlpex file name I like having access to all the text editing features that I have in vscode as opposed to the barebones text editing features in the terminal.

3. There is also all the other moving/copying/renaming with visual feedback that you decided to completely ignore.

4. If touch was faster then oil.nvim would not have been such a popular extension. I am sure most vim users know how to use `touch`.


If you need complex file manipulation, all of that can be achieved by writing a shell script. That's what I've been doing. You also automatically get access to flow control statements and tools like sed/awk/find.

> all the text editing features that I have in vscode as opposed to the barebones text editing features in the terminal.

VSCode is a very primitive text editor compared to vim, emacs or helix. You don't need to edit the command line right there in the shell prompt, nor do you need to create any files — press Ctrl+X + Ctrl+E and hack away. Save and close the file (ZZ in vim, for example), and it gets executed by the shell.

> then oil.nvim would not have been such a popular extension

Popularity is a bad metric, most people don't bother to learn the tools they're using.


> If you need complex file manipulation, all of that can be achieved by writing a shell script. That's what I've been doing. You also automatically get access to flow control statements and tools like sed/awk/find.

Well yes, of course they all "can" be done by writing a shell script, the same way any text editing with vim "can" also be done using ed.

> VSCode is a very primitive text editor compared to vim, emacs or helix. You don't need to edit the command line right there in the shell prompt, nor do you need to create any files — press Ctrl+X + Ctrl+E and hack away. Save and close the file (ZZ in vim, for example), and it gets executed by the shell.

I actually use vscode with the vim extension. You seem to be assuming I am unfamiliar with vim and emacs, I can assure you I know them well enough (at least vim, I also am familiar with the overall features of emacs, though I lack the muscle memory to use it efficiently).

Here is an example: Let's say you have a file named `feature_experimental.cpp` now you want to remove the `_experimental.cpp` from all the files in the current directory which have `_experimental`. I assure you that I can do it faster using voil than you can with vanilla vscode.


1. There is an inbuilt terminal in VS code. Its almost always active for me and even if it isnt focusing it/bringing it up is the same distance as firing up voil. The benefit here is that it doesnt occupy your editor 2. What complex file names do you need text editing features for? 3. fzf and zoxide covers most of it

I dont want to return the favor of speculate on intent of comment as yours would be petulant and stubborn without focusing on meaningful rebuttal. Im placing this in my comment as based on your other responses there does seem to be a pattern.


You can view the source code and package the extension yourself if you are worried about that. It is only ~2000 LOC.

It is not easy to get verified in vscode marketplace, even major publishers like Qt organization are not verified much less so a solo open source developer like myself.


I’m Iranian too and our names get people a lot more concerned.

If your name sounded English the implicit bias would make you sound more trust worthy.


I have high 2 digits of extensions in my VS Code, and yours is the only one that wouldn't have a verified publisher. And I certainly have more than one from solo developers.

Qt organization (because you mentioned it) also has verification. It displays a different message (because I haven't installed anything from them):

> The extension Qt Core is published by Qt Group. This is the first extension you're installing from this publisher.

> Qt Group has verified ownership of qt.io.

> Visual Studio Code has no control over the behavior of third-party extensions, including how they manage your personal data. Proceed only if you trust the publisher.

EDIT: I'm sure there are other extensions that are also by unverified publishers. It was the first time I was hit with that message though.


The burden isn't just when I install it, I need to validate every time it's updated as well. But let's be realistic, the fact that I intrinsically trust extensions published by Microsoft isn't any better.


> view the source code and package the extension yourself

The problem is that nobody will do that. Even if it were 500 LOC.

And this is why supply chain attacks are on the rise.


What are you proposing? Should I not be allowed to develop and publish an extension that I think is useful?

> nobody will do that

"nobody" is a strong word. Yes, most people don't do that, but if a single person reads the source code and finds something nefarious they can report it or leave a review disclosing that and my reputation would be ruined.


IMO you should avoid installing editor extensions generally. It's better to try to get them merged into the editor itself.

I don't think it's good to constrain people in some way from doing that, you should just have a personal policy of avoiding extensions you're not involved in the development of.


I thought the entire point of vscode was to be an extensible "lightweight" barebones code editor, as opposed to eg jetbrains stuff; what about vim/emacs then?


I did not by any means want to discourage you from developing things and sharing them, if anything I thank you for that.

My intention was to highlight that the SW supply chain nowadays is an insecure mess.

Regarding your last point, for the vast majority of open source SW releases, we can never be sure if the release we get is produced from the same code we see. I do not know if that is the case with VScode addons, but you get my point


> Regarding your last point, for the vast majority of open source SW releases, we can never be sure if the release we get is produced from the same code we see. I do not know if that is the case with VScode addons, but you get my point

You actually can depackage vscode's .vsix files (it is just a zip file) and compare the package contents to the repository.


Yes but realistically, who is going to do that ?

Again, I am not questioning your integrity or your plugin.


>The problem is that nobody will do that. Even if it were 500 LOC.

I do it with the code I download to extend Emacs.


There was some discussion about dired here: https://news.ycombinator.com/item?id=44568404


yup ! thanks i read it all. have been using Emacs for longer than i care to admit.

just like fvwm, there is nothing better than :o) !


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: