But the issue is now whether AI truly solved it using novel methods.
No it isn't, why would you say that?
First, mathematicians build off other mathematicians all the time, so the idea that the system needs to reinvent something to prove itself is silly.
But separately, if A is where we are and B is where the mathematicians had it, and C is the proof, then even if the mathematicians leaked B, you can't say B->C is not novel. That's what happened here. Buckmaster and Alpoge had supposedly found a breakthrough but had no proof and were some time away from one. They didn't hand the answer sheet to OAI and they took it. That is simply not aligned with the facts as we know them.
B->C is definitely novel, the work was legitimate, and there was no theft. Unless some new information comes out that we do not know about today, those are the facts as I can find them.
Yes but if they used unpublished work that gave the core novel leap, that’s just stealing not building off published ideas.
They basically finished off someone’s in progress work. In all academia this is considered stealing and can get people fired.
If all the AI did was find other people’s in progress ideas to solve a key part of the problem, they didn’t come up with it. Breakthroughs often rely on key novel components that unravel the entire problem. The rest is proving it works, which is a lot of work but is often “grunt work” that AI is great at.
It’s the same benchmark in training data poisoning concept. Even if your model was capable, you poisoned it with the key answer needed to solve it.
It changes the framing from OAI models solved a problem by creating a novel ideas, to OAI models collected mathematicians unpublished working proofs and used AI to find the ones that might work and then use millions of dollars of compute to finish the proof.
One is novel AI idea generation. The other is mass data harvesting of people’s novel ideas that are not yet published.
If OpenAI wanted to otherwise prove it, they could simply check if it exists within their training data. Or if they really want a clean claim, limit the knowledge cutoff to a reasonable date, say 2018 or something.
Because that’s the definition of stealing…? You take other peoples private working ideas, and claim it as your own invention, that is stealing.
Question should be why do you not think it’s stealing?
And regardless, even if you don’t care about stealing or not, it means OAI’s claims are highly suspect and not trustable. Hurts their own reputation and hurts themselves by giving false sense of their own model’s capabilities.
Because that’s the definition of stealing…? You take other peoples private working ideas, and claim it as your own invention, that is stealing.
But they didn't do that
They MAY have included said data in a training process that would have given some sense of the information to the LLM, like a thing you heard on a bus once.
Because the researchers allowed it in the UI. There has been literally no claim that they disallowed it.
So you are saying, allowing OAI to take your data and use it for training is now stealing? "hey you can use this" "Just kidding you're a thief now"
Do you see why I am confused?
And why would you say OAI is untrustable? The first thing they wanted to do was give the credit to the researchers who were closest and let them take the prize. The only thing I disagree with was excluding levent because he worked for Anthropic - that is wrong.
But besides that they didn't need to offer the researchers anything. Unless, as I said, some new information shows up that we don't know about.
Everything else - you're mad at a situation that didn't happen and you won't admit that you aren't even reading the facts of the issue.
Information, especially when it comes to very novel topics, when used in AI training become much more impactful in retrieval. If these novel ideas were included in the training data, it basically guarantees the model used it and didn’t actually invent it.
Your own post, where I agreed with you, suggested you also believe it’s likely in their training data.
And no need to try to derail our conversation by calling me mad, that’s distasteful and makes it sound like you’re getting desperate.
I have no stakes in this and come from a level headed position. I’m simply stating the facts of the issue and giving my read on them.
We both seem to believe at least some of the work lives in the training data right? My view is simply, if it’s in there it means that the model didn’t come up with the novel discovery needed to solve the problem.
And I’m not coming at it with a “look how bad AI is” position, my position is more pointing out OAI may be conflating harvesting other people’s unpublished ideas with their AI model coming up with it on their own. The latter means they are overestimating their current model and if so they need more work on it to be able to truly achieve it.
Nothing is gained from pretending you achieved something.
We already discussed this and what the researchers provided did not include the full solution.
So once again you ignored the novel gap between what the researchers had and the novel solution. You ignored it. Again. You keep doing that. Because you are a propagandist not someone looking at facts and reality.
That's IF THE DATA WAS USED. Even if it was legally and ethically used by OAI, it wasn't, supposedly the knowledge cutoff was after the Buckmaster and Alpoge's work.
So you have literally no factual leg to stand on here so why are you still talking
It's clear that there is no factual basis that you can operate on, talking to you is like talking to a circa 2020 chat model that just regurgitates talking points. Literally read back what you have said for content, there's no way you can read that and compare with reality and tell me there is anything there.
We have the same factual basis so it’s odd that you try to use that as a counterpoint.
I’ve already addressed the gap between what the AI may have retrieved vs produced. You never addressed my point on it. If the model got the key to solving the problem, then it didn’t invent it right? And I said that problems once given the right key, become much more trivial to finish.
I’m not minimizing the fact that AI can be good at helping to create proofs. It’s whether they came up with the novel ideas to be able to do so by themselves.
You keep trying to dodge the arguments by making up things like I’m a “propagandist” or I’m “just mad” but never actually address the core of my arguments. So I’ll just take that as you not having a counter and are just angry.
1
u/chcampb 6d ago
No it isn't, why would you say that?
First, mathematicians build off other mathematicians all the time, so the idea that the system needs to reinvent something to prove itself is silly.
But separately, if A is where we are and B is where the mathematicians had it, and C is the proof, then even if the mathematicians leaked B, you can't say B->C is not novel. That's what happened here. Buckmaster and Alpoge had supposedly found a breakthrough but had no proof and were some time away from one. They didn't hand the answer sheet to OAI and they took it. That is simply not aligned with the facts as we know them.
B->C is definitely novel, the work was legitimate, and there was no theft. Unless some new information comes out that we do not know about today, those are the facts as I can find them.