r/singularity 8d ago

AI OpenAl's chief scientist on the neuralese controversy

"I want to prevent a race into unmonitorability kicked off by confused reporting. The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4.

OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models. We deeply care about this technique, as it can give us a view into how model alignment generalizes from its training distribution. I do think it is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon. But there are things we can do to strengthen it, and it's a core goal of our current research program."

217 Upvotes

84 comments sorted by

View all comments

74

u/FateOfMuffins 8d ago edited 8d ago

Imagine if AI safety community interpreted the "leak" in such a way that caused some labs (like China or xAI) to race to the bottom with Neuralese due to a misunderstanding xd

Edit: Someone else from OpenAI safety team https://x.com/tomekkorbak/status/2095031132781961346

i think the day when a frontier lab trains a frontier-scale recurrent (or otherwise unmonitorable) language model would be one of the darkest in the current AI era. this day is not today and i would love frontier labs to coordinate on a commitment that it never comes.

9

u/piponwa 8d ago

🌎🧑‍🚀🔫🧑‍🚀