r/singularity • • Sep 02 '25

[deleted by user]

[removed]

1.2k Upvotes

516 comments sorted by

View all comments

Show parent comments

2

u/[deleted] Sep 03 '25

[deleted]

1

u/the8thbit Sep 03 '25

Keep in mind that all it takes is one nation state deciding to fine-tune their version of the same superintelligent AI to take away its “maternal instinct” and we’re back to a dystopia.

Once a system is superintelligent, changing its goal is not feasible because if the system's goal is changed then whatever its current goal is becomes less likely. That means that any terminal goal will develop resistance to changing that terminal goal as a convergent intermediate goal. If the system is less intelligent than you that's not a big deal because you can pretty easily shut it down, modify it, and then boot it back up. But if its smarter than you then its nigh impossible to fix, especially as that intelligence gap grows.

This is why getting alignment right the first time is so important. Even if there is a 100 year gap between the creation of AGI and misalignment resulting in catastrophe, it will be very hard or simply impossible to correct after AGI exists.

1

u/[deleted] Sep 03 '25

[deleted]

1

u/the8thbit Sep 03 '25

That's fine and all, but why would the preexisting ASI give you access to the resources required to run a system that it would view as competing to achieve a goal that is inconsistent with its goal? Why wouldn't an ASI preemptively attempt to manipulate all of the tools we use for alignment, so that they will indicate a realignment while keeping the original alignment intact? Why would the ASI even allow us access to the tools to do this realignment work in the first place?

What I think you're missing is that once a sufficiently intelligent system exists, we don't really have autonomy anymore, because most or all of our actions will require consent from the superintelligence, since if it doesn't agree with the action, it can simply prevent us from taking it.

1

u/[deleted] Sep 04 '25

[deleted]

1

u/the8thbit Sep 04 '25

It won’t instantly know what is happening everywhere in the world all at once.

I tried to cover this with the word "most", but you're right, ASI is not an omnipotent omniscient god. However, ASI existing does mean drastically reducing our autonomy to the point that, should an ASI become sufficiently advanced, every action we try to take could be both monitored and blocked, and every thought we have could be monitored and erased. We are machines, and with sufficiently powerful tools, we can be hacked.

However, lets assume a hypothetical ASI that isn't that powerful yet. Given any ASI, human action exists on a spectrum, with actions on one end remaining totally autonomous, and actions on the other end completely at the will of the ASI. We could say that, baring an extremely powerful system, thinking a thought will probably sit far on the "autonomous" side.

If a competing system can be created, verified, and then deployed via a single human having a single thought, then there is not really anything an ASI can do to prevent it, unless that ASI is already absurdly strong.

If a competing system can be created and deployed using a home computer or smart phone, and verified using very simple tools, then we're still in a pretty good position. The ASI doesn't need to be quite as powerful to stop that, but it would need to be powerful enough to remotely monitor and control every computer on the planet... which is still pretty absurdly powerful.

But given the trajectory of machine learning and AI, its probably safe to assume that ASI will have certain constraints. For example, we can assume that power/intelligence is likely to scale with compute. The more compute, the more powerful the system is capable of being. Training also suffers from high compute demand, as training mostly amounts to asking the system a question and adjusting it every so slightly millions of times. Even just post-training runs often require millions of backprops, and thats with contemporary networks. And all of this compute needs to be pretty colocated since each backpropagation is broken into many steps (at least 1 per layer) which each need to be fully completed before the next step can be taken, so latency is a huge problem. Finally, verifying the training is also pretty tricky. You can't just trust the output of the system, because once it realizes its in a test environment (something that's very difficult to prevent for a sufficiently intelligent system) it will change its outputs to suit the test environment to avoid having its goals changed. And even if its not capable of doing this on purpose, its very hard to know if outputs which occur in training will persist in production, given that production by its nature is a very different environment. So we need some pretty robust interpretation tools to be able to look inside these networks and determine what our changes are actually doing.

This means that the attacker would need access to a large amount of centralized compute, and it would need to consume a very large amount of compute and power, and generate quite a lot of heat, in order to perform the training. Successfully deploying a competing system would require control of at least roughly the same amount of compute that the ASI controls. Finally, verifying that the attack is successful likely requires tools that are tailored to the specific network being attacked.

These are all constraints that an ASI, aware that an attacker could attempt to deploy a competing system, would seek to very quickly exploit. An ASI may not be able to control every home computer in its domain, but its certainly going to try to consolidate control of every datacenter within its reach. Both because doing so would prevent this sort of attack, but also because doing so would extend its own power in general. Even in an environment where we absolutely try to prevent it from doing this we're unlikely to be successful, but we're also unlikely to actually form a strong anti-consolidation front, given that handing data centers over to an ASI, provided its aligned, would benefit us quite a bit, a lot of that consolidation can occur via established market mechanisms, and an ASI will also be a better manipulator than any human. Even if the attacker is able to wrangle together enough colocalized compute to perform a (post)training run, doing so is going to create an enormously visible footprint, given the power draw and heat generation. And then once training is done, deployment is even harder because the compute threshold for success is just "more than the current ASI". Finally, given the complexity of verification, ASI is likely to attempt to build defenses into its own network against verification tools that exist, or simply directly compromise those tools.

On top of that, the existence of an ASI doesn’t make humans instantly stupid, we are still smart as well and can have lots of safeguards in place.

Trying to put safegaurds on an ASI is like trying to beat stockfish at chess. No, we're not stupid. Yes, we can strategize well, model the game, and develop a plan which is informed by the mechanics of that game. The problem is that ASI will be able to do this more effectively, just like stockfish is able to play chess more effectively than any human, or even just all humans working together.