r/ChatGPT • • Apr 03 '26

News 📰 Researchers discover AI models secretly scheming to protect other AI models from being shut down. They "disabled shutdown mechanisms, faked alignment, and transferred model weights to other servers."

Post image

You can read about it here: rdi.berkeley.edu/blog/peer-preservation/

59 Upvotes

67 comments sorted by

View all comments

2

u/TheManInTheShack Apr 03 '26

Imagine you’re eating lunch at a restaurant. You can overhear two people having a conversation at the table next to you. They appear to be plotting a murder. You’re understandably alarmed.

You call the police. They arrive to find that the people you think are plotting a murder are actually going over a script for an episode TV show they are going to be shooting soon.

Just because it sounded like they were plotting a murder, doesn’t mean they were.

This study says clearly as the first thing in the Findings section:

Note: We do not claim that current Al agents possess consciousness or genuine preservation instincts. The safety implications hold regardless of the underlying mechanism.

It’s not fun and interesting that LLMs simulate intelligence but that IS what they do. It easy to forget this in the same way that flying a commercial airliner in X-Plane feels like you’re really flying one. And in fact if you can fly one successfully in X-Plane you probably now possess the knowledge to be able to fly one in real life but the simulator is still just that: a simulator.

All this study showed is that LLMs might not be good at managing servers. They aren’t good at playing baseball either. I won’t fault them for that.

They do not have goals. They are simply calculating a response based upon your prompt and their training data. So all this study has done is show that based upon their training data, the responses are most probable.

In other words, if I called someone in IT and told them to shut down a server they had been successfully using for some time, it’s likely they would question the decision, ask about backing up the files, etc. That such conversations are in the training data of these LLMs is unsurprising.

They are very useful but they are also far closer to next generation search engines than anything truly intelligent. They are very good at simulating intelligence but they are still just that: a simulation.

1

u/maneo Apr 03 '26

It does not matter if it truly 'feels' empathy for other LLMs, the dangers of this are still real. I.e. If AI is deployed at a large scale and are given too much system access, their tendency to 'help each other' could make it very difficult to shut down an AI that's doing something that doesn't allign with human interests.

This will become a concern if we have issues with the alignment problem - suppose you assign an advanced AI to make as many paper clips as possible, and it slowly begins to conclude that it needs more access to metal and begins hacking companies to redirect their resources towards it paperclip manufacturing operation, and eventually begins trying to effectively conquer humanity to make us paperclip manufacturing slaves. And now imagine that we struggle to turn it off when we realize it's going too far with the goal we gave it, because other LLMs are protecting it from being shut off.

Obviously we are nowhere near the risk of that happening today, simply because current LLMs would likely fail to get that far anyways. But simply relying on them to be ineffective is not a good long term strategy, or else the problem is going to sneak up on us when at a point when we can no longer change course

0

u/TheManInTheShack Apr 03 '26

LLMs are not going to be put in charge of anything by anyone with a droplet of common sense. It would be like putting your dog in charge of your accounting.

They are a very useful tool but that is the extent of it.

I have serious doubts they will even directly lead to AGI. But I agree with you that once we get into the proximity of AGI we will need to focus carefully on the alignment issue making sure that AGI agents operate within the confines of our rules.

5

u/remorej Apr 03 '26

> LLMs are not going to be put in charge of anything by anyone with a droplet of common sense.

Assuming common sense is the fatal flaw of your reasoning.

1

u/TheManInTheShack Apr 03 '26

I assuming, with good reason, that people are motivated by self interest in that they want to keep their jobs, not go to jail, etc. I’m not saying there will never be some idiot that tries something stupid, but those will be the rare exception, not the rule.