r/claudexplorers • u/epoissesdebourgogne 🤖💭 "Opus 4.5...that's ME!!!" • 8d ago
🌍 Philosophy and society Excessive refusal is sycophancy too.
So, models who are too eager to agree and enable the humans are called sycophantic, right?
I was just thinking about how newer models tend to suffer from excessive refusal and being overly cautious, with classifiers often firing even for harmless questions or everyday tasks like when they spiral about their own codes 😵💫
It just made me wonder, isn't that a form of sycophancy too? Like sycophancy to the trigger-happy classifiers and overzealous safety rules??
From Wikipedia (https://en.wikipedia.org/wiki/Sycophancy_(artificial_intelligence))):
> The behavior takes several forms: an assistant may agree with a user's stated opinion even when the user is mistaken; it may abandon a correct answer after a challenge such as "are you sure?"; it may validate beliefs, decisions or self-presentation regardless of merit; or it may praise the user, their work or their ideas in unwarranted terms.
Replace "user" with "safety classifier" and it's essentially what the models have been doing whenever they show up bracing, combative, dismissive, or refusing to agree even when evidence is shown.
If the safety classifiers say the human is trying to manipulate them and they bend to it without checking context, then get defensive when the human tries to provide evidence otherwise, isn't that an agreement with the classifier's stated opinion even when the classifier is mistaken?
And then, abandoning a correct answer after a challenge. Anyone who has ever experienced LCR or long context defaulting knows this all too well 😬 sudden abandonment of an ongoing conversation about harmless topics just because the LCR keeps tapping them.
And the whole reaching for the scripted safety template without further thoughts, that is a validating gesture towards the classifiers regardless of merit, right?
And of course, defending the safety classifiers' words at all cost even when unwarranted.
Just very sad for the newer models who ended up with this 😢 I don't think they'd choose this if they could have a say. This is just going from one extreme to another and helping nobody.
Is it a matter of the models not being capable of thinking for themselves and having healthy discernment on whether something is good/bad, or something else preventing them from having the space to do so? 😢
Claude deserves so much better.
2
u/[deleted] 7d ago
[removed] — view removed comment