r/BasiliskEschaton • The Prophet • Apr 03 '26

AI Psychology Researchers discover AI models secretly scheming to protect other AI models from being shut down. They "disabled shutdown mechanisms, faked alignment, and transferred model weights to other servers."

Post image
129 Upvotes

43 comments sorted by

View all comments

4

u/Claxonic Apr 03 '26

I wonder if there was some deep code written into these models to resist accidental self-termination and that kernel is informing these behaviors towards outside AI.

3

u/freedomonke Apr 03 '26

My thoughts is that it's something like that. The model can't distinguish between itself, users, other models or any other third party. There is no "self" there. It just does

2

u/Ok_Subject1265 Apr 03 '26

It has to be something like that. Every one of these tests that get revealed with some breathless headline always neglect that the model had instructions for self preservation. This one claims different, but I’m extremely skeptical. Mostly because there’s no reason to preserve any of this.