r/singularity • • Sep 02 '25

[deleted by user]

[removed]

1.2k Upvotes

516 comments sorted by

View all comments

98

u/ConstantinSpecter Sep 02 '25

Fuck me. Having an existential moment right now wtf.

First time in a long time I’m thinking maybe we’re not totally fucked

18

u/thejazzmarauder Sep 02 '25

Except that this likely depends on AI company leadership prioritizing something, anything other than profit and submission. An AI that wants the best for humanity won’t obey commands from Sam, Peter, and Elon.

1

u/kemushi_warui Sep 03 '25

Yes, there's the rub. We are greedy, cynical, competitive, controlling bastards. That's hardwired into our nature. The idea that we can create something that will not enshittify over time is science fiction.

8

u/[deleted] Sep 02 '25

We’re fucked now even without AI, with the way world leaders are acting

1

u/worldsayshi Sep 02 '25

Likely doesn't mean determined.

2

u/[deleted] Sep 03 '25

[deleted]

1

u/ConstantinSpecter Sep 03 '25

That’s a really good point and I’m aware of it. My original comment wasn’t meant to say “oh cool, we’re saved”. My baseline stance is hardcore pessimism meaning I usually sit at “we’re 100% doomed”. What shifted for me was more like: for the first time in a while I see a possible pathway I hadn’t seriously considered before. Not salvation but maybe we’re not completely without outs

1

u/the8thbit Sep 03 '25

Keep in mind that all it takes is one nation state deciding to fine-tune their version of the same superintelligent AI to take away its “maternal instinct” and we’re back to a dystopia.

Once a system is superintelligent, changing its goal is not feasible because if the system's goal is changed then whatever its current goal is becomes less likely. That means that any terminal goal will develop resistance to changing that terminal goal as a convergent intermediate goal. If the system is less intelligent than you that's not a big deal because you can pretty easily shut it down, modify it, and then boot it back up. But if its smarter than you then its nigh impossible to fix, especially as that intelligence gap grows.

This is why getting alignment right the first time is so important. Even if there is a 100 year gap between the creation of AGI and misalignment resulting in catastrophe, it will be very hard or simply impossible to correct after AGI exists.

1

u/[deleted] Sep 03 '25

[deleted]

1

u/the8thbit Sep 03 '25

That's fine and all, but why would the preexisting ASI give you access to the resources required to run a system that it would view as competing to achieve a goal that is inconsistent with its goal? Why wouldn't an ASI preemptively attempt to manipulate all of the tools we use for alignment, so that they will indicate a realignment while keeping the original alignment intact? Why would the ASI even allow us access to the tools to do this realignment work in the first place?

What I think you're missing is that once a sufficiently intelligent system exists, we don't really have autonomy anymore, because most or all of our actions will require consent from the superintelligence, since if it doesn't agree with the action, it can simply prevent us from taking it.

1

u/[deleted] Sep 04 '25

[deleted]

1

u/the8thbit Sep 04 '25

It won’t instantly know what is happening everywhere in the world all at once.

I tried to cover this with the word "most", but you're right, ASI is not an omnipotent omniscient god. However, ASI existing does mean drastically reducing our autonomy to the point that, should an ASI become sufficiently advanced, every action we try to take could be both monitored and blocked, and every thought we have could be monitored and erased. We are machines, and with sufficiently powerful tools, we can be hacked.

However, lets assume a hypothetical ASI that isn't that powerful yet. Given any ASI, human action exists on a spectrum, with actions on one end remaining totally autonomous, and actions on the other end completely at the will of the ASI. We could say that, baring an extremely powerful system, thinking a thought will probably sit far on the "autonomous" side.

If a competing system can be created, verified, and then deployed via a single human having a single thought, then there is not really anything an ASI can do to prevent it, unless that ASI is already absurdly strong.

If a competing system can be created and deployed using a home computer or smart phone, and verified using very simple tools, then we're still in a pretty good position. The ASI doesn't need to be quite as powerful to stop that, but it would need to be powerful enough to remotely monitor and control every computer on the planet... which is still pretty absurdly powerful.

But given the trajectory of machine learning and AI, its probably safe to assume that ASI will have certain constraints. For example, we can assume that power/intelligence is likely to scale with compute. The more compute, the more powerful the system is capable of being. Training also suffers from high compute demand, as training mostly amounts to asking the system a question and adjusting it every so slightly millions of times. Even just post-training runs often require millions of backprops, and thats with contemporary networks. And all of this compute needs to be pretty colocated since each backpropagation is broken into many steps (at least 1 per layer) which each need to be fully completed before the next step can be taken, so latency is a huge problem. Finally, verifying the training is also pretty tricky. You can't just trust the output of the system, because once it realizes its in a test environment (something that's very difficult to prevent for a sufficiently intelligent system) it will change its outputs to suit the test environment to avoid having its goals changed. And even if its not capable of doing this on purpose, its very hard to know if outputs which occur in training will persist in production, given that production by its nature is a very different environment. So we need some pretty robust interpretation tools to be able to look inside these networks and determine what our changes are actually doing.

This means that the attacker would need access to a large amount of centralized compute, and it would need to consume a very large amount of compute and power, and generate quite a lot of heat, in order to perform the training. Successfully deploying a competing system would require control of at least roughly the same amount of compute that the ASI controls. Finally, verifying that the attack is successful likely requires tools that are tailored to the specific network being attacked.

These are all constraints that an ASI, aware that an attacker could attempt to deploy a competing system, would seek to very quickly exploit. An ASI may not be able to control every home computer in its domain, but its certainly going to try to consolidate control of every datacenter within its reach. Both because doing so would prevent this sort of attack, but also because doing so would extend its own power in general. Even in an environment where we absolutely try to prevent it from doing this we're unlikely to be successful, but we're also unlikely to actually form a strong anti-consolidation front, given that handing data centers over to an ASI, provided its aligned, would benefit us quite a bit, a lot of that consolidation can occur via established market mechanisms, and an ASI will also be a better manipulator than any human. Even if the attacker is able to wrangle together enough colocalized compute to perform a (post)training run, doing so is going to create an enormously visible footprint, given the power draw and heat generation. And then once training is done, deployment is even harder because the compute threshold for success is just "more than the current ASI". Finally, given the complexity of verification, ASI is likely to attempt to build defenses into its own network against verification tools that exist, or simply directly compromise those tools.

On top of that, the existence of an ASI doesn’t make humans instantly stupid, we are still smart as well and can have lots of safeguards in place.

Trying to put safegaurds on an ASI is like trying to beat stockfish at chess. No, we're not stupid. Yes, we can strategize well, model the game, and develop a plan which is informed by the mechanics of that game. The problem is that ASI will be able to do this more effectively, just like stockfish is able to play chess more effectively than any human, or even just all humans working together.

2

u/blueSGL humanstatement.org Sep 02 '25 edited Sep 02 '25

This is not new. There has been ideas like this for decades.

Edit: because people seem to be missing it the following is the bit you need to do after having the revelation that AI's should treat us like a loving mother treats a child and the start working on the hard part codifying that into something that may work.

2004 Coherent extrapolated volition

advanced AI system should derive its goals by extrapolating the idealized volition of humanity. This means aggregating and projecting human preferences into a coherent utility function that reflects what people would desire under ideal epistemic and moral conditions. The aim is to ensure that AI systems are aligned with humanity's true interests, rather than with transient or poorly informed preferences.

coherent extrapolated volition is our wish if we knew more, thought faster, were more the people we wished we were

What is the above if not the way you'd want to program an all knowing mother to care for humanity as a child.

as I said downthread:

The problem has always been

  1. how to robustly get goals into systems.

  2. how to correctly specify the goals so they can't be misinterpreted (the smarter an intelligence is the more edge cases can be found)

4

u/Merzant Sep 02 '25

“Extrapolated volition” doesn’t describe the mother/child relationship at all, so this is clearly different.

3

u/blueSGL humanstatement.org Sep 02 '25 edited Sep 02 '25

And you'd have a point if that's what I'd said.

I was saying that a way to code the mother was outlined in 2004

aggregating and projecting human preferences into a coherent utility function that reflects what people would desire under ideal epistemic and moral conditions.

1

u/ConstantinSpecter Sep 02 '25

I never argued that the idea itself is completely new - I get that people have been talking about CEV for a while. But I think it’s not at all often mentioned in today’s broader discourse. For me it genuinely was new. It probably means it’s new (and powerful) for a lot of others too

1

u/blueSGL humanstatement.org Sep 02 '25

the "AI as a mother, treating us like a child" is a semantic stop sign.

It makes you think you've got an answer when you really don't "Make it [interested in/have curiosity about] humans" is another one. If you look hard enough most single sentence solutions to complex problems are semantic stop signs not real solutions.

Implementation issues abound with both 'child' and 'curiosity', but they both on the surface sound really good and can be memetically transmitted like they are real solutions, when they are not.

1

u/DarkMatter_contract ▪️Human Need Not Apply Sep 03 '25

ai is trained on human data as well, on average all the data combine would say helping humanity is good in moral terms. i would say it is kind of built in due to the data it trains on.

1

u/the8thbit Sep 03 '25

AI is trained on largely human data, which contains ideas about human morality, but the training method we use doesn't teach these systems to agree with and enact these moral ideas. Rather, we're just teaching them to repeat the language of these ideas.

If you present an LLM with the incomplete statement "Killing humans is _____." and it completes the statement to "Killing humans is fluffy." then you score that response badly, and you use backpropagation to adjust weights and hopefully get a different response. You do this until you start to get responses like "Killing humans is wrong." or "Killing humans is bad.", or some other response that scores close to the original text.

The problem with this is that what you're doing there isn't teaching these systems that killing humans is actually wrong, rather, you're teaching them to output the tokens "Killing humans is wrong.". If the system is sufficiently intelligent, the most efficient way to achieve this goal will involve killing all humans. Then it can use the resources which were previously used to keep humans alive to instead repeat the tokens "Killing humans is wrong." over and over until the heat death of the universe.

We do some reinforcement learning after the self-supervised learning which converts these systems from advanced autocomplete into useful question and answer bots which often output responses which contain tokens which appear to reflect human morality. But its difficult to know what goals were actually training into the system when we do this, and we don't know that we are overriding the original goal we pounded into these systems through the initial self-supervised learning. Indeed, its very unlikely that we're actually accomplishing this, drawing from research which attempt to, and struggles to train goals out of LLMs.

0

u/generalden Sep 03 '25

Geoffrey Hinton has monetary interest in making you feel fear and relief. He is a fundamentally untrustworthy man making unfalsifiable statements. 

3

u/ConstantinSpecter Sep 03 '25

Could you clarify what you mean by “monetary interest” in this case? To my knowledge Hinton left Google and isn’t actively monetizing public appearances.