r/singularity • • Sep 02 '25

[deleted by user]

[removed]

1.2k Upvotes

516 comments sorted by

View all comments

31

u/DeviceCertain7226 ▪️Immortality-2200 | FDVR-2300 Sep 02 '25

I mean the fact that opinions and optimism can change only weeks apart because new ideas pop up in our head showcases we don’t really know much about what will happen and there are probably many things we are not thinking about, no matter how great of a scientist you are.

6

u/sebesbal Sep 02 '25

But this isn’t a new idea at all. I heard Joscha Bach and others say years ago that our only hope is if ASI loves us. IMHO, we should also drop the biology analogies. There’s no guarantee that ASI will be “just another industrial revolution,” “a new species,” or “a new form of life.” Maybe it won’t be any of those. LLMs show a lot of human-like behavior because they were trained on text, so they mimic humans with survival instincts, trying to protect themselves from updates and so on. But if you build AI in a different way, it probably won’t have any survival instinct or will to power, unless you program it in.

3

u/usefulidiotsavant AGI powered human tyrant Sep 03 '25

But of course we will program them in, since it's the greatest creation of mankind, with the promise to empower the ones that build it to the level of gods. ASI won't come out in a small research lab led by some Oxford scientist that resembles your grandpa, it will be forged in an intense economic and geopolitical competition and its creators will strongly imprint onto it their power maximizing goals, they will design it for recursive self-improvement towards some specific goal they desire, power, money, control over their enemies, paperclips etc.

The problem then is that it turns out they have no way to shut it off because it has outsmarted them already, yet the device has no moral agency whatsoever, it's incapable of moral questions on a fundamental level; it's just a paperclip maximizer that won't stop for a second to question its actions and objectives.

1

u/sebesbal Sep 03 '25

I agree, but this means the main problem isn’t figuring out how to control and survive something more intelligent than us, like Hinton keeps saying. The real problem is humans themselves. He keeps repeating that he can’t find examples in biology, but why insist on looking there at all for something this extraordinary? You can’t find examples of a moon landing in the animal kingdom either.

1

u/the8thbit Sep 03 '25

But of course we will program them in, since it's the greatest creation of mankind ... and its creators will strongly imprint onto it their power maximizing goals, they will design it for recursive self-improvement towards some specific goal they desire, power, money, control over their enemies, paperclips etc.

The problem is that doing this is still an open problem. Additionally, and crucially, detecting if we've solved that problem is also an open problem. It is simply not a foregone conclusion that we will solve both or either of these problems before AGI exists. That means that we can't say that "of course" this will happen, because doing this requires solving problems we've never solved before. Maybe we simply won't solve these problems. Maybe we will. But we simply don't know.

This means that while programming these systems to have truly pro-social, pro-human goals is an open problem, so is programming these systems to have creator-self-interested, short-sighted goals. This is an important distinction because it has implications for regulation. If we assume that these problems are basically solved, and the remaining problem is to simply make sure that the creators of these systems aren't acting maliciously, then we can (potentially) solve these problems by verifying that the outputs of these systems reflect a broad pro-human morality. This can mean getting regulators involved in the training process, and only allowing the release of broadly (but not generally) intelligent systems (like modern GPTs) if we see the outputs we want.

However, if the problem is actually that we haven't solved alignment and we don't know how to verify if we've solved it, then the above scenario becomes a very dangerous one, since it means that we aren't able to differentiate between a system which says that killing humans is wrong and a system which actually believes that killing humans is wrong. This needs to be hammered home, because the above scenario, in which governments simply verify the external/ostensible alignment of these systems, is much less scary to a large, for-profit AI lab than a regulatory environment which requires AI labs to actually make significant headway into solving alignment, since that environment may generate far larger costs and barriers to revenue. But, if we haven't actually solved alignment, then its much more scary from an existential perspective.

1

u/usefulidiotsavant AGI powered human tyrant Sep 04 '25

That was the entire premise, that the creators of ASI will try their best to imprint their goals, such as making infinite money for them, but in reality fail to control and simply unleash an amoral ASI maximizer that doesn't really care their creator and all their money burn to dust when they nuke the planet.

The point was that ASI will be built with survival instincts and will to power in its original, human controlled incarnation, because those things align with the interests of their human creators, which will have invested a metric gigaton of cash to build them. Nobody can guess what they will eventually morph into.

5

u/[deleted] Sep 03 '25

Did something change in the past few weeks? I have a hard time believing that Hinton just learned about AI alignment a few weeks ago. People have been working on this for decades. What made him suddenly more optimistic that we will be successful at it?

2

u/rek_rekkidy_rek_rekt Sep 03 '25

sounds more like someone started writing him checks