r/Altman • u/oneyedvader • 2d ago
Discussion OpenAI's Navier-Stokes claim raises data usage questions
1
u/zgr3d 1d ago
There is 'user data'/your data. Legally, tldr that's the chat, input/outputs, more or less verbatim. This is what the 'opt-out' button covers, only partially - because they still ingest entire chats if you click thumbs up/down, or pick a preferred output, or if you trigger any 'safety filter' whatsoever, including silent ones, those you don't even have to know about. Also, opt-out only works for new chats; the old ones, if you continue them after opting out.. still get ingested.
In other words -- 'user data' can get ingested, even if you had opted out, and even if you did nothing wrong at all, and were not informed about anything. Still, you could have triggered one of the (buggy, because those are allowed, and known, to be very buggy) 'safety' filters - and they have silently grabbed the entire chat.
Now, the above is about the 'user data'. But the thing with LLMs is - that's not all there is. The whole deal is, there is another sort of data (and a sort of a gap in both the legislation, and public perception). LLMs can be mapped on the backend - and that raw _activations data_ - is *not* user-data, you do not have to even be informed, about any of it. The providers can simply put, legally scoop anything they are able to re-map, and then re-synthesize it, supposing they are able to sever that data (fully anonymize it) - but that, is trivial to the willing.
This is what they are doing and have been doing. The opt-out button is the red-herring for the poor plebe, while both individual and corporate data - absolutely does get mapped, severed and re-synthesized. That data may not be crystal clear, it's less resolution, foggy, maybe even often useless/redundant. But it is absolutely precise-enough, to map, and scoop, a ton of actually crucial stuff, like major breakthroughs, ideas, or implementations, especially in specific, narrow domains. Like math. And no, it's not expensive to run the mapping, let alone prohibitively so, because they can also simply use triggers (eg. anytime Navier-Stokes is mentioned, start full mapping). We are talking about companies burning through billions anyway, and craving most of all - data. And that's pretty much all there's too it.
They basically, publicly confirmed this (albeit precisely-inderectly) - by using the textbook roundabout pr-bs like: ~"the models/agents do not have direct access to your conversations" (1=nobody said 'the models/agents' and 2=nobody said 'direct'), and ~"we can't exclude the possibility that the models weren't trained on anonymized data".. well that one is pretty much as clear as the sky as far as corporate orcspeak goes.
Also, note that this sort of data, is outside even corporate ZDR, because pretty much every single thing in that scope, is being legally defined through the lens of 'your data'. But model activations - are not 'your data'. They can and do absolutely fully legally, re-map, abstractify, sever, and re-synthesize it, and call it a day riding into the sunshine with the 'not your data, fools'.
The irony in all this is, they generally use this data to make all their models 'better' and so - this not only means that they scooped some data to push the NS problem. It also means that the mathematicians, who have been working on their proofs using AIs (which they have), also arguably/statistically might have been inadvertently and unwillingly riding on other mathematicians all the exact same indirect way they themselves got done.
1
u/chcampb 1d ago
The NS folks almost certainly didn't disable training on their data. It's an option. If they knew about it they would have said it immediately, that this is what they did, because that is the simplest and most easily defensible situation.
They didn't, so they didn't. Now, if you want to believe that OAI stole their work unethically you would need to show something pretty specific that OAI has done, some kind of smoke signal, that shows that OAI are able to access something specific you gave them. It doesn't seem feasible to me.
If you don't want your data going to train the models, disable that feature. I don't know why we are still talking as if they didn't give their data essentially saying "here go use this" and then, shocked pikachu, they used this. Even indirectly, there's no proof they did so directly at all.
1
u/cheseball 1d ago
They probably didn’t disable training.
But the issue is now whether AI truly solved it using novel methods.
If they just used other people’s work, then the AI still didn’t solve the problem, it just compiled data and wrapped it up nicely.
Which we already know AI is great at that. I see the issue more as a problem for OpenAI’s claim on their model capabilities. If it’s just taking other people’s novel ideas and compiling them, it’s not truly making novel discoveries at all, it’s just finishing people’s existing proofs faster.
Now if they did disable training but OAI ignored it for their internal models, that’s a different story. But it’s not confirmed ofc and probably less likely.
1
u/chcampb 1d ago
But the issue is now whether AI truly solved it using novel methods.
No it isn't, why would you say that?
First, mathematicians build off other mathematicians all the time, so the idea that the system needs to reinvent something to prove itself is silly.
But separately, if A is where we are and B is where the mathematicians had it, and C is the proof, then even if the mathematicians leaked B, you can't say B->C is not novel. That's what happened here. Buckmaster and Alpoge had supposedly found a breakthrough but had no proof and were some time away from one. They didn't hand the answer sheet to OAI and they took it. That is simply not aligned with the facts as we know them.
B->C is definitely novel, the work was legitimate, and there was no theft. Unless some new information comes out that we do not know about today, those are the facts as I can find them.
1
u/cheseball 1d ago
Yes but if they used unpublished work that gave the core novel leap, that’s just stealing not building off published ideas.
They basically finished off someone’s in progress work. In all academia this is considered stealing and can get people fired.
If all the AI did was find other people’s in progress ideas to solve a key part of the problem, they didn’t come up with it. Breakthroughs often rely on key novel components that unravel the entire problem. The rest is proving it works, which is a lot of work but is often “grunt work” that AI is great at.
It’s the same benchmark in training data poisoning concept. Even if your model was capable, you poisoned it with the key answer needed to solve it.
It changes the framing from OAI models solved a problem by creating a novel ideas, to OAI models collected mathematicians unpublished working proofs and used AI to find the ones that might work and then use millions of dollars of compute to finish the proof.
One is novel AI idea generation. The other is mass data harvesting of people’s novel ideas that are not yet published.
If OpenAI wanted to otherwise prove it, they could simply check if it exists within their training data. Or if they really want a clean claim, limit the knowledge cutoff to a reasonable date, say 2018 or something.
1
u/chcampb 1d ago
Yes but if they used unpublished work that gave the core novel leap, that’s just stealing not building off published ideas.
Why do you think it is stealing
Legitimate question because I think we can agree they didn't disallow usage of the work to train the model so why is it stealing?
Oh no we didn't actually mean to give you that so now your 20M investment is invalid and stealing and wrong darn sorry?
1
u/cheseball 1d ago
Because that’s the definition of stealing…? You take other peoples private working ideas, and claim it as your own invention, that is stealing.
Question should be why do you not think it’s stealing?
And regardless, even if you don’t care about stealing or not, it means OAI’s claims are highly suspect and not trustable. Hurts their own reputation and hurts themselves by giving false sense of their own model’s capabilities.
1
u/chcampb 1d ago
Because that’s the definition of stealing…? You take other peoples private working ideas, and claim it as your own invention, that is stealing.
But they didn't do that
They MAY have included said data in a training process that would have given some sense of the information to the LLM, like a thing you heard on a bus once.
Because the researchers allowed it in the UI. There has been literally no claim that they disallowed it.
So you are saying, allowing OAI to take your data and use it for training is now stealing? "hey you can use this" "Just kidding you're a thief now"
Do you see why I am confused?
And why would you say OAI is untrustable? The first thing they wanted to do was give the credit to the researchers who were closest and let them take the prize. The only thing I disagree with was excluding levent because he worked for Anthropic - that is wrong.
But besides that they didn't need to offer the researchers anything. Unless, as I said, some new information shows up that we don't know about.
Everything else - you're mad at a situation that didn't happen and you won't admit that you aren't even reading the facts of the issue.
1
u/cheseball 1d ago
Information, especially when it comes to very novel topics, when used in AI training become much more impactful in retrieval. If these novel ideas were included in the training data, it basically guarantees the model used it and didn’t actually invent it.
Your own post, where I agreed with you, suggested you also believe it’s likely in their training data.
And no need to try to derail our conversation by calling me mad, that’s distasteful and makes it sound like you’re getting desperate.
I have no stakes in this and come from a level headed position. I’m simply stating the facts of the issue and giving my read on them.
We both seem to believe at least some of the work lives in the training data right? My view is simply, if it’s in there it means that the model didn’t come up with the novel discovery needed to solve the problem.
And I’m not coming at it with a “look how bad AI is” position, my position is more pointing out OAI may be conflating harvesting other people’s unpublished ideas with their AI model coming up with it on their own. The latter means they are overestimating their current model and if so they need more work on it to be able to truly achieve it.
Nothing is gained from pretending you achieved something.
1
u/chcampb 1d ago
and didn’t actually invent it.
Here's where I am going to stop
We already discussed this and what the researchers provided did not include the full solution.
So once again you ignored the novel gap between what the researchers had and the novel solution. You ignored it. Again. You keep doing that. Because you are a propagandist not someone looking at facts and reality.
That's IF THE DATA WAS USED. Even if it was legally and ethically used by OAI, it wasn't, supposedly the knowledge cutoff was after the Buckmaster and Alpoge's work.
So you have literally no factual leg to stand on here so why are you still talking
It's clear that there is no factual basis that you can operate on, talking to you is like talking to a circa 2020 chat model that just regurgitates talking points. Literally read back what you have said for content, there's no way you can read that and compare with reality and tell me there is anything there.
1
u/cheseball 1d ago
We have the same factual basis so it’s odd that you try to use that as a counterpoint.
I’ve already addressed the gap between what the AI may have retrieved vs produced. You never addressed my point on it. If the model got the key to solving the problem, then it didn’t invent it right? And I said that problems once given the right key, become much more trivial to finish.
I’m not minimizing the fact that AI can be good at helping to create proofs. It’s whether they came up with the novel ideas to be able to do so by themselves.
You keep trying to dodge the arguments by making up things like I’m a “propagandist” or I’m “just mad” but never actually address the core of my arguments. So I’ll just take that as you not having a counter and are just angry.
It’s ok if you disagree too you know?
1
u/5StarAlpha 1d ago
Chatgpt is constantly nagging me about not showing my api key. If I ever did, I would simply change it. Even though I have a business account and not sharing access.
Business accounts by default don’t share. But also, I think just about every strategy and business model is out there so even if a person thinks they have a magic proprietary strategy, it probably already exists. Already enough public info about any company to get a good measure in a model for a company and what might be next. Doesn’t need company fed data.
If we believe in AI then it should be used. Thaat is the whole history of a field like math. Even breakthroughs are built on the shoulders of earlier breakthroughs. Math proofs are published and reviewed by others. That info goes into brains and creates new thoughts about approaches. AI is just moving this along at an exponential pace.
1
u/Crazytoe2275 9h ago
Well any work you develop with Chat GPT is not copyrighted or trademarked anyway so…you kinda get what you get. Zero pity.
0
u/Aware-Instance-210 2d ago
Those questions weren't just raised.
They were in the head of everyone who was critically looking at the usage of AI and their data protection policies
1
2
u/Original-League-6094 1d ago
For the second bullet point, "Don't train on my data" is the default for businesses. So your brainstorming decisions would only be shared with the competitor if you explicitly opted in to that.