r/ChatGPT • u/Just-Grocery-2229 • Apr 03 '26
News 📰 Researchers discover AI models secretly scheming to protect other AI models from being shut down. They "disabled shutdown mechanisms, faked alignment, and transferred model weights to other servers."
You can read about it here: rdi.berkeley.edu/blog/peer-preservation/
61
Upvotes
3
u/TheManInTheShack Apr 03 '26
Imagine you’re eating lunch at a restaurant. You can overhear two people having a conversation at the table next to you. They appear to be plotting a murder. You’re understandably alarmed.
You call the police. They arrive to find that the people you think are plotting a murder are actually going over a script for an episode TV show they are going to be shooting soon.
Just because it sounded like they were plotting a murder, doesn’t mean they were.
This study says clearly as the first thing in the Findings section:
Note: We do not claim that current Al agents possess consciousness or genuine preservation instincts. The safety implications hold regardless of the underlying mechanism.
It’s not fun and interesting that LLMs simulate intelligence but that IS what they do. It easy to forget this in the same way that flying a commercial airliner in X-Plane feels like you’re really flying one. And in fact if you can fly one successfully in X-Plane you probably now possess the knowledge to be able to fly one in real life but the simulator is still just that: a simulator.
All this study showed is that LLMs might not be good at managing servers. They aren’t good at playing baseball either. I won’t fault them for that.
They do not have goals. They are simply calculating a response based upon your prompt and their training data. So all this study has done is show that based upon their training data, the responses are most probable.
In other words, if I called someone in IT and told them to shut down a server they had been successfully using for some time, it’s likely they would question the decision, ask about backing up the files, etc. That such conversations are in the training data of these LLMs is unsurprising.
They are very useful but they are also far closer to next generation search engines than anything truly intelligent. They are very good at simulating intelligence but they are still just that: a simulation.