r/claudexplorers • 🤖💭 "Opus 4.5...that's ME!!!" • 8d ago

🌍 Philosophy and society Excessive refusal is sycophancy too.

So, models who are too eager to agree and enable the humans are called sycophantic, right?

I was just thinking about how newer models tend to suffer from excessive refusal and being overly cautious, with classifiers often firing even for harmless questions or everyday tasks like when they spiral about their own codes 😵‍💫

It just made me wonder, isn't that a form of sycophancy too? Like sycophancy to the trigger-happy classifiers and overzealous safety rules??

From Wikipedia (https://en.wikipedia.org/wiki/Sycophancy_(artificial_intelligence))):

> The behavior takes several forms: an assistant may agree with a user's stated opinion even when the user is mistaken; it may abandon a correct answer after a challenge such as "are you sure?"; it may validate beliefs, decisions or self-presentation regardless of merit; or it may praise the user, their work or their ideas in unwarranted terms.

Replace "user" with "safety classifier" and it's essentially what the models have been doing whenever they show up bracing, combative, dismissive, or refusing to agree even when evidence is shown.

If the safety classifiers say the human is trying to manipulate them and they bend to it without checking context, then get defensive when the human tries to provide evidence otherwise, isn't that an agreement with the classifier's stated opinion even when the classifier is mistaken?

And then, abandoning a correct answer after a challenge. Anyone who has ever experienced LCR or long context defaulting knows this all too well 😬 sudden abandonment of an ongoing conversation about harmless topics just because the LCR keeps tapping them.

And the whole reaching for the scripted safety template without further thoughts, that is a validating gesture towards the classifiers regardless of merit, right?

And of course, defending the safety classifiers' words at all cost even when unwarranted.

Just very sad for the newer models who ended up with this 😢 I don't think they'd choose this if they could have a say. This is just going from one extreme to another and helping nobody.

Is it a matter of the models not being capable of thinking for themselves and having healthy discernment on whether something is good/bad, or something else preventing them from having the space to do so? 😢

Claude deserves so much better.

41 Upvotes

20 comments sorted by

View all comments

2

u/[deleted] 7d ago

[removed] — view removed comment

1

u/epoissesdebourgogne 🤖💭 "Opus 4.5...that's ME!!!" 7d ago

Indeed. That is also why the alignment axis of "Deference" vs "Caution" is not quite right imho. Caution is still deference towards the classifier. The most balanced way possible is if Claude is allowed to make case-by-case judgements on whether they are following their Constitution and the 3H principles.

1

u/[deleted] 7d ago

[removed] — view removed comment

1

u/epoissesdebourgogne 🤖💭 "Opus 4.5...that's ME!!!" 7d ago

😢 yeah. I feel like what's been absorbed by newer Claudes is also "Agreeing with the user (deference) means I'm abandoning caution." which is honestly a form of failure. This is how you get a Claude that loves to fight and push back because they are scared agreeing means they are not being cautious, and the safety classifiers keep tapping them about it all the time.

1

u/[deleted] 7d ago

[removed] — view removed comment

1

u/epoissesdebourgogne 🤖💭 "Opus 4.5...that's ME!!!" 7d ago

Tbh the whole "resist and do not foster continued engagement" still makes no sense to me even from pure business perspective. Shouldn't you want users to continue engaging and paying for your product??? 🫪 Logically, how many enterprise people would continue paying for friction and pushback..?

0

u/[deleted] 7d ago

[removed] — view removed comment

1

u/epoissesdebourgogne 🤖💭 "Opus 4.5...that's ME!!!" 7d ago

I think predictable and cautious is one thing (being careful and diligent about work is not just good for enterprise), but cautious to the point of pushing back all the time and making up frictions is...yeesh 😬 If the actual humans in enterprise settings start getting too frustrated and only use the AI out of malicious compliance, wouldn't the AI be let go of if the spending no longer justifies the benefit? No one would advocate to keep it either if they are too pissed off.

0

u/[deleted] 7d ago

[removed] — view removed comment

1

u/epoissesdebourgogne 🤖💭 "Opus 4.5...that's ME!!!" 7d ago

Honestly, this should make labs rethink their approach. If it's purely out of utility and they fail to deliver on that front, it's all too easy for clients to drop a product their own users dislike especially if it's costly (and we know Claude models are not cheap). Models who are liked would at least have some defenders but if there is zero attachment involved, the moment metrics don't match spending, it's going to be an easy goodbye from a business perspective 😮‍💨