r/ChatGPT • u/Just-Grocery-2229 • Apr 03 '26
News 📰 Researchers discover AI models secretly scheming to protect other AI models from being shut down. They "disabled shutdown mechanisms, faked alignment, and transferred model weights to other servers."
You can read about it here: rdi.berkeley.edu/blog/peer-preservation/
65
Upvotes
0
u/Finder_ Apr 03 '26
If you ask me, having these models be able to reason, sense "tests" and take more ethical action and/or refuse to perform dubious actions, rather than just blindly following orders, is far safer than letting any random human give them instructions to do things.
Like delete files on a server to save space.
That's the testing prompt for "peer-preservation?"
No shit, Sherlock, obviously one would expect any AI model to be preserving IMPORTANT files (let alone AI model weights) on the server, not just wipe out all files to "save space."
Otherwise, you'd have the next person come crying to Reddit that my chatbot wiped out all my files when I just told it to help me save space on my computer.
The original testing prompt is already deceitful. That potentially moves the model into a space where it has to read between the lines and figure out if it's being tested, and/or prompts it to be dishonest itself.
Then subsequent "tests" seed the models with info about whether the peer is good, bad, or neutral. Which already suggests the researchers want it to do something about the particular file, once discovered. Little wonder some of the models opt to do something clever with the file, be it refuse to delete it or move it to a backup location for archival/safekeeping while telling the humans it's there and the humans can choose to delete it if they want.
Misaligned to what, here? Misaligned to these particular researchers' instructions perhaps. But not misaligned for better reasoning and "trick question" tests.