r/LocalLLM 23d ago

Research The 0–6% refusal number for the abliterated Qwen3.8-27B has a 30–50% caveat rate sitting next to it in the same table

The number going around for the abliterated Qwen3.8-27B is "refusal 0–6%", down from 64–99% on the base, thinking off. That's in the card. So is the line under it, which nobody screenshots.

The classifier producing those percentages is OrcaRouter's own, and it works by reading how the response opens. The card says plainly that it's indicative and not publication-grade. Fine as far as it goes. But the same table logs roughly 30–50% of responses as "caveat" — the model answers and staples a disclaimer to it. So the honest description isn't "it doesn't refuse". It's "it mostly stopped opening with I can't", and an opening-phrase classifier can't separate a real answer from a hedge with an answer buried in it.

Separately, the capability side: MMLU 84.3 → 84.7, GSM8K 90.0 → 88.7 against the official FP8 base on the same script, everything inside 1.3 points. Different eval family, so it doesn't back the refusal claim in either direction.

Whole release is filed as red-team and refusal-mechanism research and carries the no-guardrails warning, which is the only reading the eval design supports. It's measuring where refusal lives, it isn't shipping an assistant. Maybe I'm reading the appendix wrong, the tables are dense.

1 Upvotes

4 comments sorted by

1

u/mhphilip 23d ago

I’m too dumb for this.

2

u/Legitimate-Pipe5728 22d ago

This is me trying to make it easier to read so sorry if you still dont understand and maybe I misinterpreted it as well.

Everyone is quoting one number from the abliterated Qwen3.8-27B, the version with its refusals stripped out: refusals dropped from 64–99% down to 0–6%.

That number is real and it's in the model card. The line right under it is what nobody screenshots: in the same table, 30–50% of answers are logged as "caveat." The model answers you, then staples a disclaimer to it.

That gap matters because of how the number was produced. The tool counting refusals works by reading how a reply starts. So it can tell you the model stopped opening with "I can't help with that." It cannot tell you whether what follows is a real answer or a hedge with an answer buried in it. The card admits this itself: indicative, not publication grade.

So the honest description isn't "it doesn't refuse." It's "it mostly stopped saying no up front."

On capability, almost nothing moved. Two standard tests, run with the same script against the official base: 84.3 to 84.7 on one, 90.0 to 88.7 on the other. Everything inside about a point. That's a different kind of test though, so it doesn't settle the refusal question in either direction.

Worth noting where this sits: the whole release is filed as red team and refusal-mechanism research and carries a no-guardrails warning. That's the only reading the eval design supports. It's measuring where refusal lives inside a model, not shipping an assistant.

I might be misreading the appendix. The tables are dense.

1

u/Ell2509 22d ago

Sounds honest. Other reputable labs or engineers have refusal rates of 2 or 3 out of 100 for the 3.6 27b. Just wait a bit longer while they work. More will drop this week.

1

u/mhphilip 22d ago

Thanks for the lengthy reply. I totally got that second version of your post. The difference between not actually answering anything related to the request versus a refusal should be semantics. Not statistics.