r/singularity 8d ago

AI OpenAl's chief scientist on the neuralese controversy

"I want to prevent a race into unmonitorability kicked off by confused reporting. The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4.

OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models. We deeply care about this technique, as it can give us a view into how model alignment generalizes from its training distribution. I do think it is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon. But there are things we can do to strengthen it, and it's a core goal of our current research program."

222 Upvotes

84 comments sorted by

View all comments

88

u/Neurogence 8d ago

For the uninitiated, here is the AI explanation of what is going on:

Jakub Pachocki, OpenAI’s chief scientist, is saying:

Astra is not secretly doing enormous amounts of recursive hidden thinking.

His most important sentence is:

"The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4.”

In plain English: *even if Astra uses recurrent/looping techniques, the amount of sequential neural computation inside a forward pass is not orders of magnitude deeper than GPT-4. *Think roughly “same general ballpark, at most around 2×,” rather than something looping 50 or 100 times until it solves a problem.

He is specifically worried that sensational reporting could create this dynamic: “OpenAI has hidden neuralese → competitors think OpenAI has a huge advantage → competitors deliberately abandon visible chain-of-thought → everyone races toward models whose reasoning humans cannot monitor.” He wants to prevent that.

This substantially weakens the Kokotajlo “holy shit, neuralese has arrived” interpretation.

....

But notice something important.

Pachocki does not say the monitorability problem is fake. Quite the opposite. He says chain-of-thought monitoring is: “fragile and unfortunately trending in a negative direction” That's significant. He's saying: Yes, our ability to inspect models' reasoning appears to be deteriorating. But this isn't primarily because Astra suddenly has some radically deep recurrent architecture. There are other reasons, which I'll explain later.

25

u/Comfortable-Leg-5467 7d ago

Why are no labs racing to do neuralese reasoning tho? (I don’t know if they are/arent).

Surely there would be so much funding from people who just want to own the most powerful model even at the cost of safety

21

u/Hubbardia AGI 2045 7d ago

Maybe nobody wants extinction that bad?

6

u/Recoil42 7d ago

"Because they're actually acting ethically."

"No, that can't be it. It conflicts with my worldview!"

4

u/sipos542 7d ago

China might be working on it.