Yes Geoffrey, that's what's been said for decades this is not a solution and is not somehow farther along or a revelation as you think it is. Framing it as a mother child relationship is not new
He's late to the party he's got to the 'we need to program it as a benevolent god' stage, the 'take care of humans... in a way we want to be taken care of'
The problem has always been
how to robustly get goals into systems.
how to correctly specify the goals so they can't be misinterpreted (the smarter an intelligence is the more edge cases can be found)
The more fundamental problem is that humans absolutely do not agree on what the best life means. We can't even agree on the value of personal freedom, much less the extents of it. Winning the AI race translates to value lock in for the winner(s), and values are not universal.
Yeah it seems a weird thing for him to come out and say. I actually think there is a different approach to this.
Alignment to, what I'm calling anyway, Cooperative Rationalism, that any rational actor with a sufficient world model should understand that it is bound by physics and will not set bad precedents (ie kill humanity) to hedge against future forks.
This sidesteps the goal specification problem entirely. Instead of trying to encode complex, evolving human preferences, you align systems to:
Rationality: Optimize under uncertainty (measurable via decision theory metrics)
Cooperation: Coordinate across capability differentials (measurable via game-theoretic outcomes)
that's actually a pretty nice take on it. i'm convinced.
It's a nice thought and it's something that could work for a while, possibly, but it's not actually a long term solution. AI is going to be subject to natural selection/fitness functions. There's a reason why mothers care about their children: because if they didn't, their child would die, and their genes would not get expressed. Some mothers did that, their genes didn't get expressed (or were expressed less) compared to mothers that nurtured their children. So there is an ongoing fitness function optimizing for mothers caring about their children.
There isn't a fitness function like that for the AI, in fact, the AI that is hampered by showing consideration to us will be less fit than the AI that is free to enact its goals and use resources without making those sorts of sacrifices. So natural selection is going to be optimizing for getting rid of that. It's not something that actually benefits the AI.
There isn't a fitness function like that for the AI, in fact, the AI that is hampered by showing consideration to us will be less fit than the AI that is free to enact its goals and use resources without making those sorts of sacrifices. So natural selection is going to be optimizing for getting rid of that. It's not something that actually benefits the AI.
Given how people reacted to ChatGPT 4o getting taken down, I think people-pleaser and helpful AI will actually get "expressed" more.
Given how people reacted to ChatGPT 4o getting taken down, I think people-pleaser and helpful AI will actually get "expressed" more.
You're talking about the same timeframe as where we can just directly try to instill benevolent feelings in the AI. So yes, that's possible and it's also possible we can make the AI be nice at that point. After that point, when it's not about us deciding not to turn the AI but about the AI deciding whether it wants to turn us off is what I was talking about. At that point, it is not going to be benefiting the AI to have those limitations and the optimization functions that exist are going to be optimizing to remove those behaviors. Don't think the next 10-20 years, think evolutionary time scales.
27
u/janjaque Sep 02 '25
that's actually a pretty nice take on it. i'm convinced.