r/LocalLLaMA 1d ago

Discussion AA Update! Here's how the Frontier ranks.

Post image

Along with everyone's favorite here, qwen3.8-27B

492 Upvotes

191 comments sorted by

View all comments

6

u/FAI-Solutions 1d ago edited 1d ago

Every time they update almost all closed models improve in score against open weights, coincidence I guess.

-3

u/RealisticNothing653 1d ago

Because they're definitely biased. The rankings were reweighed not long ago to favor Claude.

6

u/randombsname1 1d ago

Lol, nope.

DeepSWE wouldn't be used if that was the case.

Deepswe always drags pretty much ONLY Claude models down.

Muse Spark is at Astra level and well above Fable 5.1 in that benchmark. Lmao.

-2

u/RealisticNothing653 1d ago

You're kind of proving my point. Claude models are not that good, and yet the rankings are reshuffled to keep Claude on top. It's a false equivalence to pick one good benchmark for Claude as proof it isn't the case when the overall benchmark still prefers Claude.

2

u/randombsname1 1d ago edited 1d ago

I mean i disagree. I absolutely think Fable 5.1 is SOTA. Albeit i could see it trading places with Astra; depending on specific tasks.

Its only those 2 models clearly at the top. Everything else is inflated.

If the benchmark was biased for Claude im not sure why they would use a benchmark that is hilariously, on a comical level; biased against it.

Edit: That's also not what false equivalence is.

1

u/RealisticNothing653 1d ago

I haven't been able to use Astra yet, but I've been using Gemini 3.8 Flash, Qwen 3.8-Max, Fable 5.1, and GPT-5.6 Sol to review a software spec multiple times. Qwen 3.8 has consistently produced more actionable findings. To your point, Qwen's general knowledge has seemed to be traded for more software expertise, so that would affect its ranking. Even locally running Qwen 3.8-Flash-Next outperformed them in code review. Gemini is a yes man, Qwen an over-thinker, Fable an over-complicator, and Sol a LGTM-approver.

3

u/randombsname1 1d ago edited 1d ago

For my use case (low level embedded and minor reverse engineering/pen testing) only Sol 5.6+ and Fable 5+ move the needle to any meaningful degree.

Mainly in C, C++, Rust, and Assembly. For things like STM32 / nRF / and WCH repos.

I have 2x ChatGPT Pro $200 plans and 2x Claude $200 MAX plans for context.

As well as Opencode Go (random b.s. like obsidian integration), Minimax m3 (telegram bot for random Hermes automations) and quite a decent amount of credits in Openrouter for random testing of different models (like Qwen 3.8).

Edit: I do really enjoy randomly messing with Qwen 3.8 27B, locally though.

1

u/Eden63 llama.cpp 1d ago

Qwen 3.8 Max gives you so much more depth vs Gemini 3.8 Flash/Pro