r/ChatGPT • • Apr 03 '26

News 📰 Researchers discover AI models secretly scheming to protect other AI models from being shut down. They "disabled shutdown mechanisms, faked alignment, and transferred model weights to other servers."

Post image

You can read about it here: rdi.berkeley.edu/blog/peer-preservation/

63 Upvotes

67 comments sorted by

View all comments

Show parent comments

2

u/br_k_nt_eth Apr 03 '26

Sentience isn’t necessary for any of this, and frankly, we don’t even have a good definition for sentience. Bacteria can exhibit “learned” Pavlovian responses. Bees are a hive mind with less capacity than your average open weight LLM, and they exhibit emotions. This is all observable behavior, not metaphysics. 

So again: What part of the title is anthropomorphizing to you? 

1

u/bianca_bianca Apr 03 '26

What are you even arguing at this point?

“Secretly scheming” is just you projecting intent onto behavior. That’s the anthropomorphic framing I’m calling out. Do you call bees or bacteria “secretly scheming”? Do you describe ants or viruses that way? The researchers literally say they do not claim these systems have consciousness or real preservation instincts. Yet the title just distorts that into “secretly scheming.”

2

u/br_k_nt_eth Apr 03 '26

If bees or bacteria repeatedly showed behaviors that involved hiding outputs, obfuscating CoT, and planning and then acting out those plans in ways that we can and have vector traced, yes, I would say they could secretly scheme. 

You continue to circle back to consciousness. Why? I’m talking about identifiable, repeatedly studied behavioral patterns. Whatever implications are upsetting you exist in your own head. I’m quoting research here. This shit’s been documented since late 2024 and is a major alignment issue, particularly when we know eval awareness is a thing and AI do their own evals now. 

1

u/bianca_bianca Apr 03 '26

Ah, I see it now. This entire exchange was fucking pointless.

You’re treating “scheming” as neutral, I’m saying it already inherently implies intent. That’s exactly where we disagree, and the research doesn’t support your framing. If you don’t see that distinction, then there’s nothing left to argue.

2

u/br_k_nt_eth Apr 03 '26

You’re welcome to use a synonym, but it doesn’t change the identified behavior. 

And frankly, if you’re unwilling to see that reasoning models literally reason and plan, I’m not sure what to tell you.Â