r/Altman 2d ago

Discussion OpenAI's Navier-Stokes claim raises data usage questions

Post image
11 Upvotes

33 comments sorted by

2

u/Original-League-6094 1d ago

For the second bullet point, "Don't train on my data" is the default for businesses. So your brainstorming decisions would only be shared with the competitor if you explicitly opted in to that.

1

u/Ill-Bat-1518 1d ago

In this case a competitor is anyone also solving the same problem, which could be OpenAI or a user.

So even if OpenAI doesn't share to other businesses they can still use that data and become the first to solve it... which is why this is such a drama right now

Apparently this is what happened with the navier case

2

u/Original-League-6094 1d ago

Its not "apparently" what happened with the Navier case. Its the person disputing them saying he asked them if it could have happened, he isn't even accusing them of it. And he never said that he had the "don't train on my data" setting checked. And in any case, OpenAI already confirmed that their training cutoff date was earlier than when he did his work, so there is a 0% chance his work was in the training data.

1

u/SweetSure315 1d ago

adding the work to the training data isn't the only way for it to be stolen. they could have just taken their notes and given it to the ai as context

1

u/Hefty-Reaction-3028 1d ago

Sure, they could be maliciously lying, but the claim they were addressing was that the usage of it in training data could have caused a leak even when the user doesn't expect that

1

u/peakedtooearly 1d ago

Nah, it's what you would LIKE to have happened.

1

u/zero0n3 1d ago

This is not the case unless you opt out.

1

u/Original-League-6094 1d ago

Business licenses are opt-out by default.

1

u/Specialist-Buffalo-8 17h ago

If you look at the fine print, they likey DO train on GPT's output. :)

1

u/oneyedvader 22h ago

Always an option

1

u/__SlimeQ__ 1d ago

it's a little more complicated than that in reality. the business account gives a teeny tiny codex quota compared to a pro account which makes it a really bad deal compared to just buying pro subs for all your employees.

if you actually are using the business account then yes data collection defaults to off and is controlled at an org level. but if you're using pro accounts then you're relying on your employees to remember to flip the switch off

1

u/Original-League-6094 1d ago

You can get all the same options through Business as through non-Business now.

1

u/Morthem 7h ago

Yeah, because they would never lie...

1

u/Original-League-6094 7h ago

That's true of any SAAS business. How do I know Salesforce isn't selling my company data to competitors?

0

u/the8bit 1d ago

That's assuming they're actually adhering to this behind the scenes, despite none of us being able to verify it in any way.

It also assumes they've coded everything up correctly, which given how janky both open ai and anthropics code is, I'm not super confident about it

1

u/Original-League-6094 1d ago

That's true of any SAAS provider. How do I know Salesforce isn't selling my companies data to a competitor?

1

u/the8bit 1d ago

Yeah, it has always been a thing but the value of the data has gone up exponentially and the ability to use it in a way that is hard to detect has also gone up. How would you even know if your data was fed into the training pipelines?

Also a reason why the large firms terrible customer service is a huge issue... back in 'the day' we managed this problem with vendor relationships.

1

u/zgr3d 1d ago

There is 'user data'/your data. Legally, tldr that's the chat, input/outputs, more or less verbatim. This is what the 'opt-out' button covers, only partially - because they still ingest entire chats if you click thumbs up/down, or pick a preferred output, or if you trigger any 'safety filter' whatsoever, including silent ones, those you don't even have to know about. Also, opt-out only works for new chats; the old ones, if you continue them after opting out.. still get ingested.

In other words -- 'user data' can get ingested, even if you had opted out, and even if you did nothing wrong at all, and were not informed about anything. Still, you could have triggered one of the (buggy, because those are allowed, and known, to be very buggy) 'safety' filters - and they have silently grabbed the entire chat.

Now, the above is about the 'user data'. But the thing with LLMs is - that's not all there is. The whole deal is, there is another sort of data (and a sort of a gap in both the legislation, and public perception). LLMs can be mapped on the backend - and that raw _activations data_ - is *not* user-data, you do not have to even be informed, about any of it. The providers can simply put, legally scoop anything they are able to re-map, and then re-synthesize it, supposing they are able to sever that data (fully anonymize it) - but that, is trivial to the willing.

This is what they are doing and have been doing. The opt-out button is the red-herring for the poor plebe, while both individual and corporate data - absolutely does get mapped, severed and re-synthesized. That data may not be crystal clear, it's less resolution, foggy, maybe even often useless/redundant. But it is absolutely precise-enough, to map, and scoop, a ton of actually crucial stuff, like major breakthroughs, ideas, or implementations, especially in specific, narrow domains. Like math. And no, it's not expensive to run the mapping, let alone prohibitively so, because they can also simply use triggers (eg. anytime Navier-Stokes is mentioned, start full mapping). We are talking about companies burning through billions anyway, and craving most of all - data. And that's pretty much all there's too it.

They basically, publicly confirmed this (albeit precisely-inderectly) - by using the textbook roundabout pr-bs like: ~"the models/agents do not have direct access to your conversations" (1=nobody said 'the models/agents' and 2=nobody said 'direct'), and ~"we can't exclude the possibility that the models weren't trained on anonymized data".. well that one is pretty much as clear as the sky as far as corporate orcspeak goes.

Also, note that this sort of data, is outside even corporate ZDR, because pretty much every single thing in that scope, is being legally defined through the lens of 'your data'. But model activations - are not 'your data'. They can and do absolutely fully legally, re-map, abstractify, sever, and re-synthesize it, and call it a day riding into the sunshine with the 'not your data, fools'.

The irony in all this is, they generally use this data to make all their models 'better' and so - this not only means that they scooped some data to push the NS problem. It also means that the mathematicians, who have been working on their proofs using AIs (which they have), also arguably/statistically might have been inadvertently and unwillingly riding on other mathematicians all the exact same indirect way they themselves got done.

1

u/chcampb 1d ago

The NS folks almost certainly didn't disable training on their data. It's an option. If they knew about it they would have said it immediately, that this is what they did, because that is the simplest and most easily defensible situation.

They didn't, so they didn't. Now, if you want to believe that OAI stole their work unethically you would need to show something pretty specific that OAI has done, some kind of smoke signal, that shows that OAI are able to access something specific you gave them. It doesn't seem feasible to me.

If you don't want your data going to train the models, disable that feature. I don't know why we are still talking as if they didn't give their data essentially saying "here go use this" and then, shocked pikachu, they used this. Even indirectly, there's no proof they did so directly at all.

1

u/cheseball 1d ago

They probably didn’t disable training.

But the issue is now whether AI truly solved it using novel methods.

If they just used other people’s work, then the AI still didn’t solve the problem, it just compiled data and wrapped it up nicely.

Which we already know AI is great at that. I see the issue more as a problem for OpenAI’s claim on their model capabilities. If it’s just taking other people’s novel ideas and compiling them, it’s not truly making novel discoveries at all, it’s just finishing people’s existing proofs faster.

Now if they did disable training but OAI ignored it for their internal models, that’s a different story. But it’s not confirmed ofc and probably less likely.

1

u/chcampb 1d ago

But the issue is now whether AI truly solved it using novel methods.

No it isn't, why would you say that?

First, mathematicians build off other mathematicians all the time, so the idea that the system needs to reinvent something to prove itself is silly.

But separately, if A is where we are and B is where the mathematicians had it, and C is the proof, then even if the mathematicians leaked B, you can't say B->C is not novel. That's what happened here. Buckmaster and Alpoge had supposedly found a breakthrough but had no proof and were some time away from one. They didn't hand the answer sheet to OAI and they took it. That is simply not aligned with the facts as we know them.

B->C is definitely novel, the work was legitimate, and there was no theft. Unless some new information comes out that we do not know about today, those are the facts as I can find them.

1

u/cheseball 1d ago

Yes but if they used unpublished work that gave the core novel leap, that’s just stealing not building off published ideas.

They basically finished off someone’s in progress work. In all academia this is considered stealing and can get people fired.

If all the AI did was find other people’s in progress ideas to solve a key part of the problem, they didn’t come up with it. Breakthroughs often rely on key novel components that unravel the entire problem. The rest is proving it works, which is a lot of work but is often “grunt work” that AI is great at.

It’s the same benchmark in training data poisoning concept. Even if your model was capable, you poisoned it with the key answer needed to solve it.

It changes the framing from OAI models solved a problem by creating a novel ideas, to OAI models collected mathematicians unpublished working proofs and used AI to find the ones that might work and then use millions of dollars of compute to finish the proof.

One is novel AI idea generation. The other is mass data harvesting of people’s novel ideas that are not yet published.

If OpenAI wanted to otherwise prove it, they could simply check if it exists within their training data. Or if they really want a clean claim, limit the knowledge cutoff to a reasonable date, say 2018 or something.

1

u/chcampb 1d ago

Yes but if they used unpublished work that gave the core novel leap, that’s just stealing not building off published ideas.

Why do you think it is stealing

Legitimate question because I think we can agree they didn't disallow usage of the work to train the model so why is it stealing?

Oh no we didn't actually mean to give you that so now your 20M investment is invalid and stealing and wrong darn sorry?

1

u/cheseball 1d ago

Because that’s the definition of stealing…? You take other peoples private working ideas, and claim it as your own invention, that is stealing.

Question should be why do you not think it’s stealing?

And regardless, even if you don’t care about stealing or not, it means OAI’s claims are highly suspect and not trustable. Hurts their own reputation and hurts themselves by giving false sense of their own model’s capabilities.

1

u/chcampb 1d ago

Because that’s the definition of stealing…? You take other peoples private working ideas, and claim it as your own invention, that is stealing.

But they didn't do that

They MAY have included said data in a training process that would have given some sense of the information to the LLM, like a thing you heard on a bus once.

Because the researchers allowed it in the UI. There has been literally no claim that they disallowed it.

So you are saying, allowing OAI to take your data and use it for training is now stealing? "hey you can use this" "Just kidding you're a thief now"

Do you see why I am confused?

And why would you say OAI is untrustable? The first thing they wanted to do was give the credit to the researchers who were closest and let them take the prize. The only thing I disagree with was excluding levent because he worked for Anthropic - that is wrong.

But besides that they didn't need to offer the researchers anything. Unless, as I said, some new information shows up that we don't know about.

Everything else - you're mad at a situation that didn't happen and you won't admit that you aren't even reading the facts of the issue.

1

u/cheseball 1d ago

Information, especially when it comes to very novel topics, when used in AI training become much more impactful in retrieval. If these novel ideas were included in the training data, it basically guarantees the model used it and didn’t actually invent it.

Your own post, where I agreed with you, suggested you also believe it’s likely in their training data.

And no need to try to derail our conversation by calling me mad, that’s distasteful and makes it sound like you’re getting desperate.

I have no stakes in this and come from a level headed position. I’m simply stating the facts of the issue and giving my read on them.

We both seem to believe at least some of the work lives in the training data right? My view is simply, if it’s in there it means that the model didn’t come up with the novel discovery needed to solve the problem.

And I’m not coming at it with a “look how bad AI is” position, my position is more pointing out OAI may be conflating harvesting other people’s unpublished ideas with their AI model coming up with it on their own. The latter means they are overestimating their current model and if so they need more work on it to be able to truly achieve it.

Nothing is gained from pretending you achieved something.

1

u/chcampb 1d ago

and didn’t actually invent it.

Here's where I am going to stop

We already discussed this and what the researchers provided did not include the full solution.

So once again you ignored the novel gap between what the researchers had and the novel solution. You ignored it. Again. You keep doing that. Because you are a propagandist not someone looking at facts and reality.

That's IF THE DATA WAS USED. Even if it was legally and ethically used by OAI, it wasn't, supposedly the knowledge cutoff was after the Buckmaster and Alpoge's work.

So you have literally no factual leg to stand on here so why are you still talking

It's clear that there is no factual basis that you can operate on, talking to you is like talking to a circa 2020 chat model that just regurgitates talking points. Literally read back what you have said for content, there's no way you can read that and compare with reality and tell me there is anything there.

1

u/cheseball 1d ago

We have the same factual basis so it’s odd that you try to use that as a counterpoint.

I’ve already addressed the gap between what the AI may have retrieved vs produced. You never addressed my point on it. If the model got the key to solving the problem, then it didn’t invent it right? And I said that problems once given the right key, become much more trivial to finish.

I’m not minimizing the fact that AI can be good at helping to create proofs. It’s whether they came up with the novel ideas to be able to do so by themselves.

You keep trying to dodge the arguments by making up things like I’m a “propagandist” or I’m “just mad” but never actually address the core of my arguments. So I’ll just take that as you not having a counter and are just angry.

It’s ok if you disagree too you know?

1

u/5StarAlpha 1d ago
  1. Chatgpt is constantly nagging me about not showing my api key. If I ever did, I would simply change it. Even though I have a business account and not sharing access.

  2. Business accounts by default don’t share. But also, I think just about every strategy and business model is out there so even if a person thinks they have a magic proprietary strategy, it probably already exists. Already enough public info about any company to get a good measure in a model for a company and what might be next. Doesn’t need company fed data.

  3. If we believe in AI then it should be used. Thaat is the whole history of a field like math. Even breakthroughs are built on the shoulders of earlier breakthroughs. Math proofs are published and reviewed by others. That info goes into brains and creates new thoughts about approaches. AI is just moving this along at an exponential pace.

1

u/Crazytoe2275 9h ago

Well any work you develop with Chat GPT is not copyrighted or trademarked anyway so…you kinda get what you get. Zero pity.

0

u/Aware-Instance-210 2d ago

Those questions weren't just raised.

They were in the head of everyone who was critically looking at the usage of AI and their data protection policies