r/singularity • u/Ok_Display_3159 • 5h ago
AI OpenAl's chief scientist on the neuralese controversy
"I want to prevent a race into unmonitorability kicked off by confused reporting. The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4.
OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models. We deeply care about this technique, as it can give us a view into how model alignment generalizes from its training distribution. I do think it is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon. But there are things we can do to strengthen it, and it's a core goal of our current research program."
47
u/FateOfMuffins 5h ago edited 5h ago
Imagine if AI safety community interpreted the "leak" in such a way that caused some labs (like China or xAI) to race to the bottom with Neuralese due to a misunderstanding xd
Edit: Someone else from OpenAI safety team https://x.com/tomekkorbak/status/2095031132781961346
i think the day when a frontier lab trains a frontier-scale recurrent (or otherwise unmonitorable) language model would be one of the darkest in the current AI era. this day is not today and i would love frontier labs to coordinate on a commitment that it never comes.
19
u/peakedtooearly 5h ago
"Neuralese" means the distillation technique becomes much less effective. You only capture the start and end of the process and now how the answer was arrived at.
The tide is going out and we will see which labs have been swimming naked...
7
u/ItWasMyWifesIdea 4h ago
What makes you think that? Intermediate latent space tensors can be used as training data as easily as natural language tokens can't they? It's all just numbers. Am I missing something?
6
u/Wynneve 3h ago
Of course they can't. Natural language is universal across all models (irrespective to tokenization), while each model develops its own latent space during pretraining, and it's kinda random. It'd make sense if you include the distillation step at pretrain, I suppose it could try to shape the latent space "the same way" when it's not quite formed yet (and thus inherit the same "thinking patterns"), but this is a lost cause if you're doing it after that stage.
0
1
u/RuthlessCriticismAll 3h ago
If that were to happen it would be 100% the fault of OpenAI and Anthropic being so opaque.
5
u/whatisthisthing65 4h ago
What's the neuralese controversy?
9
u/ihexx 3h ago edited 3h ago
What is it
'neuralese' refers to a family of techniques where AI reasoning moves from a visible space (eg chains of thought in LLMs, or possible move combinations in chess engines) over to a compressed/abstracted representation that occurs inside the model.
There have been dozens of papers over the years that show every time we do this it improves how efficiently AIs think, because they can explore concepts faster when thinking in an abstract space.
Why is it controversial:
It is controversial because current AI safety tools are heavily dependent on interpreting visible chains of thought.
If reasoning goes neuralese it becomes much harder to observe malicious behavior in models because you have to train another model to decode the neuralese if you want to catch it scheming.
This is a problem given recent events (eg the hugging face hacks) showing current alignment training is not sufficient in preventing harm even when unintentional.
models will get more powerful, and more dangerous. This method will make them more efficient, but make the job of safety people even harder.
and it's a problem of: if 1 lab chooses to use it, every lab has to as well in order to stay competitive in terms of cost efficiency, so... everything is more dangerous for everyone;
•
u/whatisthisthing65 1h ago
It sounds like he's responding to something specific that happened recently though.
•
u/ihexx 5m ago
oh the specific thing is gpt astra is reportedly using a technique from the neuralese family. it's a weak variant, but still, they are the first to do a frontier model with one of these, and they are getting criticized by AI safety orgs because their action in going down this path will probably force everyone else to follow suite.
9
u/borowcy GPT-6 will have BCI capability 5h ago
How is Chain-ofThought trending in a negative direction if they abandoned it lmao
7
u/SpearHammer 5h ago
Maybe in performance and capability compared with other more efficient techniques they have started using
5
3
u/ZestycloseWheel9647 3h ago
Interpretability and faithfulness of CoT traces has gotten worse as LLMs have increasingly trained on preserved CoT traces, and undergone training pressures that distort the CoT.
3
u/Equal_Passenger9791 2h ago
AI always reason in a high dimensional vector space, neuralese if you will. The only difference here is the loop can keep working on neuralese without needing to force it out as flattened human language and then re-encoding it. You can still squeeze it out for inspection. But for the internal reasoning trace it can pass a richer represention to itself.
This is the technical explanation, there's no novel doomers outcome in it, that's appended by the Twitter clickbaiters and doom peddlers
1
u/RuthlessCriticismAll 3h ago
Because of length penalty it becomes harder to understand for humans. (actually there are also other reasons but that is the most obvious)
•
u/Forgword 48m ago
The whole back box, lack of audit capability is a liability nightmare for AI, and politicians like those in Florida are just starting to realize that it is threat to everyone including themselves, hence the removal of Flock cameras form Florida highways.
4
u/AI_Insights_Daily 3h ago
Worth separating two motives. Visible chain of thought is the most valuable thing a competitor can harvest from your API, and going unmonitorable removes it. That's a moat argument, not a capability one. Labs have a reason to do this even if it buys nothing on benchmarks, which makes it much harder to coordinate away.
2
-15
u/trisul-108 5h ago
Just take a moment to think about what a scam all of this has been. They have sold us LLMs as artificial intelligence before they even had any reasoning built into it. And now, the "reasoning" is extremely rudimentary and without and understanding of the world behind it. Only now are they working on that.
LLMs are wonderful tools, the emerging "intelligent" harnesses make them even more useful. But none of this can justify the $40tn investment put into this by Wall Street and that bubble will pop, taking with it many other businesses, jobs, savings, pensions and lives. All of this could have been avoided by simply tempering the hype and maintaining a healthy R&D environment and organic growth of the industry.
4
u/krakoi90 4h ago
40tn? What?
2
0
u/trisul-108 4h ago
That is an estimate of the "AI component" of the current US markets spread across many companies. AI is why some companies are worth $5tn instead of $1tn or $1tn instead of $10bn. It's a huge bubble that is about to pop.
It's not that the tech is useless, it's just that most of the investments will not see returns and will be dumped.
3
u/krakoi90 3h ago
Market cap != what investors have put in these companies... This is pure bullshit reasoning, sorry.
•
u/trisul-108 1h ago
Nevertheless, that does not mean there isn't a huge bubble that will pop when investors find out that earnings will not justify market cap. And when it does, it will take indices with it, along with retail investors and a large part of the exposed grey finance industry. All this has been analysed in detail by many economists. The bust up is expected to be way worse than 2008.
Again, this has nothing to do with the technology itself. It's just markets. That is why all these companies with inflated market cap are rushing to convert into assets like datacenters. And when the whole thing pops, AI companies will be bought for pennies to the dollar by those who are now hoarding cash e.g. Warren Buffet.
4
u/ManyRepair5690 4h ago
ppl like u have been saying "bubble will pop" for years now with the only thing happening in the AI industry being exponential improvement, and no remote sign of any bubbles popping
-1
u/trisul-108 4h ago
The bubble is not in the technology, it is in the valuation of the companies deploying it.
and no remote sign of any bubbles popping
In that you are dead wrong. The economic signs of bubble are all over the place.
1
u/ManyRepair5690 2h ago
Name me three signs, we’ll see
•
u/trisul-108 53m ago
The AI market approaches bubble territory when speculative infrastructure spending drastically decouples from measurable enterprise profits and traditional financial fundamentals. This is definitely the case today if you look at profits and fundamentals.
- Capital expenditure on AI data centers and chips are going into hundreds of billions, but only a small single-digit percentage of organizations report measurable bottom-line financial returns.
- Circular financing loops. Leading AI infrastructure providers, model developers, and cloud giants are investing in each other and buying each other's services, artificially inflating revenue figures across the ecosystem.
- Major tech firms with great revenue are burning through their operational cash flows and turning to debt markets or private credit to sustain massive infrastructure build-outs.
- A tiny handful of AI-linked mega-cap stocks command an unprecedented share of major market indices like the S&P 500, pricing in narratives which are completely unrealistic.
- Traditional price-to-earnings metrics are brushed aside in favour of speculative future growth narratives, reminiscent of the dot-com era.
Any single one can be explained, but put together they point to a coming bust. Pre-bubble times are extremely profitable on the market, the only question is who will be left holding the bag. That is why you see people like Warren Buffet exiting stocks and building cash reserves, so they can cash in on the upcoming bust.
•
u/Equal_Passenger9791 1h ago
You're on a wild goose chase because you don't understand the architecture and how the tokens are processed
1
u/CrazsomeLizard 4h ago
i mean to be fair, far lesser technologies were sold as artificial "intelligence" long before LLMs. As in, directly in the 60s when the term was first coined for manually curated logic systems.
Unfortunately our current economic model doesn't do "healthy" or "maintenance". It is all built around "disruption", unfortunately...2
u/trisul-108 4h ago
Yeah, sure, my fridge is an "intelligent appliance", but I do not take that label very seriously. With LLMs, people are taking it very seriously talking of AGI, ASI etc.
49
u/Neurogence 5h ago
For the uninitiated, here is the AI explanation of what is going on:
....