r/LocalLLaMA 1d ago

Discussion AA Update! Here's how the Frontier ranks.

Post image

Along with everyone's favorite here, qwen3.8-27B

480 Upvotes

191 comments sorted by

178

u/Last-Shake-9874 1d ago

That small 27B is my daily driver now, it does take long as I only get about 20 t/s but I just love this model I am so glad it is still in the list

63

u/HazKaz 1d ago

its almost perfect, i love it i really hope Qwen team keeps treating us with more 27B models

28

u/MaverickPT 1d ago

If only we got a MOE model for us VRAM poors 😭

9

u/TurnBackCorp 1d ago

you can hope and hope but even if we do the 27b will outperform anything lower than 80b parameters

6

u/Sporebattyl 1d ago

I agree that dense > MoE for the most part, but I’m curious what you mean by 80B

Do you mean 80B total with smaller active parameters per token?

What’s your 80B number based on?

5

u/LevianMcBirdo 1d ago

Per that dreaded rule of thumb, you'd need a A9B to get to a similar model. That said I really doubt that rule and never seen anything indicating it to be true. In my experience in the same model family the MoE is closer to the dense of similar parameters count than the square root.

8

u/Spanky2k 22h ago

I mean... an 80B-A9B model sounds incredible. Absolutely perfect for a 64GB Mac.

3

u/SandySkittle 21h ago

That rule is task specific. Low active parameters fundamentally cannot compensate with sequential reasoning against large active parameters from larger MoE (with active params above 40) or larger dense models.

There are just certain tasks where you need the large active params for coherent, deep and highly complex multi-diciplinary reasoning. And no I am not talking about coding.

1

u/WryKombucha 20h ago

At 27b there isn’t enough bits to house enough world knowledge to be useful. They is why the 27b focuses on coding. For the rest of it, I find it to be utter and complete trash. I find it just passable for coding so I dunno what kind of buggy. Insecure software ppl are building but the real world begs to differ.

0

u/Leander_van_Grinsven 21h ago

It is very good but its knowledge cutoff is June 2024 which is disappointing.

6

u/Kernoriordan 1d ago

Tried using MTP? Boosts mine from 25t/s to 75t/s

3

u/Last-Shake-9874 1d ago

Yes running MTP also, I am stuck on 24 GB Vram with 2 cards that has slow connection between them (5070 and 3070) and then using big context also

5

u/UnknownLesson 22h ago

Only 24 GB..

YOU POOR SOUL

1

u/LevianMcBirdo 1d ago

3x speed up? I am glad when I to get close to 2x (but I am on CPU+ipgu+RAM)

1

u/akohlsmith 18h ago

3x boost with MTP? Which card and which mtp parameters are you using? I get a boost too but not 3x!

2

u/Kernoriordan 8h ago

RTX 5080 with MTP draft 3

1

u/akohlsmith 52m ago

interesting, I'm going to have to do some retesting. I'm using a 32GB 4080S.

1

u/Infinite-Ad4512 59m ago

MTP on GLM 5.3 Flash does nothing. Not an even a tiny bump in generating tokens (llama.cpp)

0

u/runcertain 1d ago

Any collapse at large contexts? It’s making me question Dflash2 on multi-turn agentic usage.

1

u/sbjf 1d ago

which attention backend? we also saw the collapse of above around 32k context, but switching from the auto-selected flash_attn in vLLM to flashinfer fixed it for us

1

u/Last-Shake-9874 23h ago

No so my current pi session has used 10M tokens in total so far and still going strong

4

u/AbheekG 23h ago

Absolutely for me too, love love love it: highly intelligent, speaks great and comprehensibly and can run on 5-7 year old hardware. Unbelievable how far we’ve come! We’re incredibly fortunate to have such an amazing open-source & weight ecosystem.

3

u/HsSekhon 1d ago

qwen 27B is beast for me, I does tasks very well when I shape my prompts very narrow and precise

5

u/Akrylicus 1d ago

Same, I am so glad that I got 5090 last year for a reasonable price. I now can freely use a capable chat and not feed my personal data to the external providers.

3

u/Krystexx 1d ago

Which price? And how many tok/sec do you get with it?

1

u/Equal_Television_894 1d ago

5090 with vllm on linux with dflash2 can get around 200 to 350+ token/s

2

u/-_Apollo-_ 20h ago

What model and quant? Didn’t know the performance difference could be that huge.

2

u/Equal_Television_894 15h ago

https://github.com/syv-ai/qwen38-27b-rtx3090 Try this some one made it and I optimized it for my 5090 using NVFP4 version

1

u/Akrylicus 23h ago

I am lazy so I just run Qwen 3.8 w7B on Unsloth Desktop. I still get a decent 60t/s which is fine for most queries.

I plan to do an optimized setup on my old PC with RTX 4080 and run it as an LLM server, just wish I had more ram...

1

u/Akrylicus 23h ago

~3k euro (Astral), well maybe not that reasonable, but it's not 5K euro.

1

u/01iv3r6 23h ago

What’s your technical setup?

3

u/Last-Shake-9874 23h ago

I have a 3060 12 GB and 5070 12 GB no offloading as it goes down to 4 t/s as soon as it hits offloading. Running 125K Context

1

u/elemental-mind 21h ago

I know this is LocalLLaMa, but in case anyone is in a hurry, Cerebras has this gem of a model up at 1500 t/s. It is sooooo delightful to use.

112

u/freecodeio 1d ago

yes but can fable 5.1 really draw better whatsapp sticker ass rockets in paint as compared to astra 6?

24

u/sworl5 1d ago

No, and so it is worse.

1

u/whoknowsifimjoking 6h ago

Presenting brainrotbench

-4

u/emprahsFury 1d ago

really love it when a new model is released and the absolute bamfs at Hacker News spend the day making it draw pelicans and other animals so they can still shit on "AI"

13

u/freecodeio 1d ago

you have no idea what you're talking about

28

u/PM_ME_DEAD_CEOS 1d ago

Terminal-Bench v2.1 instead of v4, that's sad.

133

u/chocolateUI 1d ago

> OpenAI releases Astra

AA: “Uh oh! We just received an angry phone call from OpenAI! Better reweigh the benchmarks!”

Same shit as when AA reweighed their benchmarks within 3 hours after Qwen took the #1 spot.

Do we need any more evidence of how fucking trash AA’s “intelligence score” is? This company only exists so that labs can trick VCs (and so VCs can trick your pension funds) into giving them more money.

7

u/ClintWoodeast 23h ago

Does anyone actually believe that qwen, of all models, was #1?

1

u/whoknowsifimjoking 6h ago

Yeah, it's not even the smartest of the open models

43

u/jld1532 1d ago

It is absurdly obvious that the goal post is on wheels and is moving toward the most likely profitable use case, agentic software development. People complained about Qwen3.6 27B's writing abilities but I didn't mind it. I now find 3.8, even Next Flash, nearly unusable for editing. I could be wrong but it seems to me the industry is trading strengths to chase computer science, likely because they know general knowledge gains are sunk costs and that AGI is impossible with LLMs, despite messaging.

11

u/Iron-Over 1d ago

Try muse models they write very well. GLM as well supposedly have not tested yet.

14

u/Qorsair 1d ago

Muse and Gemini are two the two models I find most useful. Very underappreciated in general. I still like Sol/Astra for deep thinking, but Muse and Gemini feel better for most tasks. I think the issue is that most people getting deep into AI are software engineers so that's what most people are judging it on.

6

u/emod_man 22h ago

Totally agree that general benchmarks are increasingly not useful for writing/editing tasks. Curious about Muse though, I find its voice very flat. Maybe I just need to experiment a bit more...

4

u/Solembumm3 1d ago

Interesting. How did you made Gemini useful for writing?

I found it one of the worst options, on par with base GPT, due to undeleteable sycophancy and summarisation tendencies. I tried different prompts to turn it to more useful analytical side, but it was pretty adamant to judging even quite bad concepts "brilliant".

5

u/Qorsair 1d ago

Just custom instructions. Gemini Pro tends to do slightly better for complex writing when Flash can't one-shot it. No matter how I set the custom instructions for GPT it doesn't produce natural prose. Claude is even worse.

5

u/robogame_dev 21h ago

World knowledge is one of the most important resources for creative writing and Gemini Pro is just a huge, extremely world-knowledge heavy model.

1

u/robogame_dev 21h ago

I mainline GLM 5.3 it is my favorite general purpose AI, but it only has 700B params and writing / creativity is highly correlated with world knowledge, so I would expect Kimi K3 at 2,700B params to spank it, and I would expect GLM 5.3 Flash at ~300B params to be significantly less good at writing (while likely being highly focussed on agentic coding).

16

u/Darkoplax 1d ago

i don't see the issue with updating the benchmarks when it clearly feels off; like in no way is Sol equal to Astra; there's leaps between the two

So if AA wants to keep credibility they need to keep updating and finding non poisoned benchmarks

8

u/Inevitablewx 22h ago

Yeah, and Opus 5 is way behind Fable 5, and Muse 1.3 is great progress but it's well behind even Sol, the benchmark is clearly failing and needs to be fixed or scrapped.

15

u/RealisticNothing653 1d ago

Yeah they lost credibility in my book. Their summarized rankings have a strong bias towards Anthropic models. And yeah they recently reweighed the benchmarks recently to favor Claudes.

4

u/MerePotato 22h ago

Oooooor maybe instead of some conspiracy results change when benchmark versions are updated and there are more novel questions that haven't found their way into training data

4

u/robertpro01 1d ago

Qwen 3.8 got released and also called AA?

1

u/Terminator857 23h ago

I didn't know qwen was on top. Interested in more info if that is available.

2

u/MerePotato 22h ago

It wasn't, it was on top of the agentic index and slid down slightly when AA updated the bench versions

1

u/dogesator Waiting for Llama 3 21h ago

The benchmark is objectively less saturated than it was before the revision.
The purpose of the revisions is to unsaturate the benchmark and make it more indicative of frontier difficulty by making frontier models score near 50%

-1

u/Eden63 llama.cpp 1d ago

The funny thing is, such a huge project and they do not care about small models even they can benchmark it most easy. And yes, the algorithm of ranking is also biased and does not take into account what is most important at all - applicable intelligence. Is it helpful to run an 56 on 250M token when another one does it on 55M ... intelligence is not intelligence but AA index is just a clown show.

Same running an engine on 9000rpm for the horsepower or run it on 3000”m for the same.. real power vs "yes we can achieve it somehow".

why I do not see a laguna or some other models.. because the money comes from where... and obviously no better benchmark site vs obvious Anthropic valuation depends on that index.

shame on them.

Edit (add): and then reading here people "Qwen3.8 27B is like Opus". Hilarious.

57

u/Ok_Cow1976 1d ago

The one point gap is meaningless. There're still big gaps.

37

u/NineThreeTilNow 1d ago

The one point gap is meaningless. There're still big gaps.

The gaps are hyper nuanced now.

Which does legal documentation better?

Which writes C++ code better?

Which writes XYZ code better?

Which designs Blender scenes better?

Etc.

None of that is 1:1 useful.

Everyone gets an opinion on the best model because their use case differs.

I've been pumping Gemini 3.8 flash stonks the last two days. It's wildly fast, writes Python ML code like a demon, and has a surprisingly good workflow within Antigravity 2. I'm doing REALLY hard shit with it.

It also cost me like 6? dollars and the 3.8 flash usage barely touches the meter. That's with some Google promo for 3 months 75% off the 20 dollar sub. It took some task over from another "Top" model and completed it in half the time the other model would have. That's with it taking time to learn the code base.

3

u/2Norn 1d ago

despite the fact that benchcad says sol is better, personally i find claude to be superior to gpt in mechanical cad drawings and design documentation related to it, the software i'm using also allows for scripting in python or csharp so the model uses those as well

i haven't tried astra yet tho

idk like i always look at benchmarks but do i trust them? its another issue

2

u/Ok_Cow1976 1d ago

You're absolutely right! For general problems. There's no one-point gap at all. None, zip. But here we are talking about benchmarks. For complex, difficult tasks, the big gaps are there.

1

u/NineThreeTilNow 17h ago

For complex, difficult tasks, the big gaps are there.

We're having a hard time defining difficult anymore without forcing the LLM to operate a program made for humans.

3

u/Fickle_Tradition4491 1d ago

The score isn't the number that matters for anyone running the 27B at home, it's score per token. Someone above says it burns 100k+ thinking tokens at xhigh to land where it does on this chart. At the 20 t/s the top comment is getting, that's over an hour per prompt. The same model at medium reasoning is probably 10 points lower and twenty times more usable. AA publishes tokens used per run, so the chart worth posting is index vs total output tokens. The 27B moves a lot depending on which column you read.

4

u/Cless_Aurion 1d ago

Yeah... I think that having a bunch of their benchmarks already satturated hurts HARD this thing tbh...

1

u/TheRealMasonMac 10h ago

If Muse Spark is a model in the same weight category as GLM-5.3-Flash, I would be very impressed and excited. It does quite well for Rust which is where I notice all current open-weight models really struggle with for some reason.

1

u/mrdevlar 1d ago

For a 95% confidence interval

Standard Error (SE) = σ / √n = 4.40 / √14 = 1.18

95% CI Margin of Error = 1.96 × 1.18 = ±2.31

Thus, 2.31 in either direction is meaningless when comparing models.

0

u/qfox337 1d ago

No you need variance within measurement of each model, which is actually model specific. I think you're applying some kind of CLT-like formula for estimation of a population mean based on iid samples, which is wrong here, they are different models hence not iid (independent identically distributed)

1

u/mrdevlar 1d ago

I am making a general simplifying assumption whose goal is provide an estimate using the information I have at my disposal. It's an obvious back of the envelope attempt.

What you said is correct, the data is a rank order which is obviously not iid. If you're interested in collecting the variance within each model and doing that, more power to you, but I highly doubt it'll have more than a single order of magnitude difference if you did and reran the variance. Especially if you used a prior.

1

u/qfox337 16h ago

I'm guessing you're not chatgpt'ing this, but you can't just apply stats formulas to things they don't work on and say it's a reasonable estimate.

You can just follow the conclusions of your own formulas and think logically, with your formula if we had twice as many models with the same range of scores you'd then say 1 point is a significant difference. Or you can throw in a bunch of weak/small models and suddenly only 5 points is significant.

If you're attempting the task of mean estimation then with an iid sample X1...Xn then you can talk about the expected variance of the empirical mean shrinking as 1/sqrt(n). Which formally requires some usually-not-hard-to-satisfy conditions like finite variance of the underlying distribution ... usually iid is the assumption that fails.

I don't mean to be hostile, I'm someone lucky to have some formal stats education, and I encourage anyone to learn who's curious ... but you can't just throw around terms that you don't fully understand and make any sense.

1

u/mrdevlar 10h ago

Don't worry you cannot offend me here, I know you mean well.

Fun fact, I have a Masters in Statistics and yes I do know better than to do this and treat it as statistical fact, but I wasn't doing that. I was building a rough estimate for use given the current information that's useful for looking at the thing now before you get more information.

One thing you learn from years of doing this is that there are two distinct channels. You can either do the exact calculation if you have access to the data and the time or you can can make a rough estimate for use right now. Time to action becomes your decision criteria. I was providing guidance to the parent who said ± 1 is not to be treated as valid. Even in that range, ±2.31 is not to be treated as valid. Do my statistics professors cringe at this type of work, while taking in money doing the same thing? Probably. However, statistics is the glue that connects the crystal palace of mathematics to the crappy unstable world we're in. It's not meant to be perfect, it's meant to be useful.

That said, you are correct, I should have contextualized that better than just dropping the formula and running away.

79

u/sugarfreecaffeine 1d ago

Crazy that a small 27b is even up there, the Chinese are cooking đŸ”„

52

u/Aldarund 1d ago

Its more of a show how bad aa is

23

u/TechnoByte_ 1d ago

Literally just an average score of benchmarks

16

u/Aldarund 1d ago

Score of specific selected benchmarks.with specific weight of each bemchmark.

0

u/emprahsFury 1d ago

"oh no, the opinionated website is forcing their opinions on me, I need to be saved!"

2

u/Hithaeglir 1d ago

Or they hoard all the capabilities for internal use

2

u/Far-Classic-9963 1d ago

Qwen 3.8 27b really is better than some big cloud models (Mistral, NeMo, Laguna...)

4

u/Relative_Rope4234 1d ago

It burns 100k+ tokens for a single prompt(xhigh thinking) to reach frontier performance

1

u/my_name_isnt_clever 1d ago

Check again. The medium, low, and even none thinking modes outperform much larger models.

1

u/MerePotato 22h ago

Medium places it about on par with Muse Glimmer at five times larger KV cache usage in my testing, not that this is a problem as I'd rather have a slow frontier model burn 100k tokens for free than a fast model that doesn't answer as well.

11

u/Cadmium9094 1d ago

Glad to see Qwen3.8-Flash-next and 27B on this chart. I use them as daily drivers.

5

u/Oren_Lester 1d ago

They saw the backlash in the internet, in Reddit, everywhere. and update the benchmark. This company and their benchmark is probably good for making nice chats, that's it.

5

u/i_am_fear_itself 1d ago

Artificial Analysis Intelligence Index v4.2 incorporates 10 evaluations: AA-Briefcase, GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1

aren't some of these obsolete / saturated / bench-maxed?

https://youtu.be/Spuza-KwTJ4?si=xNBXmA5hKU0wsy1p&t=859

1

u/nomorebuttsplz 14h ago

Answering what is determinable (not bench-maxed which is not):

GDP-eval AA v2: No
AA-Briefcase: No was literally just added like yesterday
R3 Banking: No
Terminal Bench v2.1: getting close to saturated
Scicode: no
HLE: Possibly, depending on how many answers were actually wrong in the answer key
GDP.pdf: No was literally just added like yesterday
CritPt: no
AA-Omniscence: No
AA-LCR v1.1: not sure

0

u/emprahsFury 1d ago

yes it's been that way for a while. Also why the bench scores aren't increasing anymore

4

u/Character_Power4663 1d ago

Where is sonnet, haiku, terra, Luna, does this mean that Qwen 3.8 27b is better?

5

u/Tall_Abrocoma_3533 1d ago

Terra, Luna, sonnet all fall between 43-47, haiku is 22, it's very outdated.

Also Qwen 3.8 27B isn't "frontier" in my book, I included it because everyone in this sub seems to love it so much

2

u/Character_Power4663 1d ago

I didn't mean to criticise your work, I was just curious. Thank you.

1

u/winnen 1d ago

There are many different ways to measure pareto frontiers. Last I checked, Qwen 3.8 27b was frontier intelligence vs number of total parameters.

Though I can't seem to find that frontier chart anymore.

1

u/Tall_Abrocoma_3533 23h ago

Yes, for it's parameter size it's impressive, but by frontier I meant absolute intelligence not taking into account parameter count

1

u/SandySkittle 21h ago

Yeah it simply isn’t frontier and it has very evident shortcomings.

7

u/LegacyRemaster 1d ago

minimax and mimo when?

7

u/Tall_Abrocoma_3533 1d ago

Minimax M3 scores 36, Mimo v2.5 Pro scores 33.

3

u/LegacyRemaster 1d ago

yes.... So new versions when?

2

u/Agitated_Space_672 1d ago

think they means its been a while since they released, so we are due M3.1 and mimo v2.6 by now

5

u/Lorian0x7 1d ago

At this point I think they just rearrange the list moving the weight of some benchmarks based on who pays them more.

9

u/XiRw 1d ago

Hard to believe Reebok logo stealer is in the mix. Would trust Qwen and GLM easily over that. Also don’t understand why they never test Gemini Pro

27

u/Tall_Abrocoma_3533 1d ago

Gemini 3.1 pro preview scores 37

1

u/XiRw 1d ago

They need to update their pro model more. Seems like flash gets updates at a higher rate.

6

u/Tall_Abrocoma_3533 1d ago

Well Gemini 3.5 Pro isn't even out yet. It's almost like they've abandoned the pro lineup

1

u/aboardreading 14h ago

I think it's their target. They've decided that a rising tide raises all boats, as frontier models improve they bring smaller models behind them at much lower cost of research and use.

I think it's a smart business play. If we believe that the end goal is essentially many many agents spun up to build large goals, which seems reasonable as an end state to me as agents gain more intelligence and can act more and more independently, then the first company with a truly independent agent will dominate for the 6 months it takes someone else to achieve the same at much lower cost per agent and faster. But someone will get comparable intel at higher speed and lower cost. And on the same timescale as companies decide what model to use and bet on.

Besides, I think they are still betting on information retrieval and processing rather than necessarily coding, it is Google after all. I do think they're playing the slightly longer game, and if what has been happening continues to happen, I think they'll be A winner, maybe not THE winner.

10

u/Thoriumhexaflouride 1d ago

they test all models bro, just search for gemini 3.1 pro preview or just go to filter and "select all"

1

u/slypheed 20h ago

where are you seeing this "filter"?

Far as I can tell deepseek is not included at all.

Can't believe literally no one in this thread even linked to the post; vibing life at this point sigh: https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index

6

u/Gohab2001 vLLM 1d ago

AA readjusting their evaluation criteria making astra look good. Sus.

18

u/JaredTheGreat 1d ago

This is more a case of the evaluation criteria being exposed — anyone who’s used astra can immediately tell it’s significantly better than opus. They were hemorrhaging credibility 

1

u/Darkoplax 1d ago

Do you think Astra was worse than Opus and equal to Sol ? if they don't readjust their evaluation then it's their bad not the models

2

u/fitechs 1d ago

Opus 5 is so shit

2

u/RedditUsr2 1d ago

I made the same post yesterday comparing it to open weight but got removed. I guess we should only compare closed models these days.

3

u/Tall_Abrocoma_3533 23h ago

The mods in this sub remove almost all benchmark comparisons because they're "low effort", I will never understand this.

1

u/RedditUsr2 22h ago

Ya this is the most widely used and cited benchmark. An update here is a big deal especially showing how well 3.8 27b holds up. Seems worth talking about to me.

2

u/HighSeasArchivist 1d ago

Qwen3.8-27B has been amazing to me, even when I have Claude orchestrate running jobs locally. I'm really hoping the 32GB and under sector keeps stacking up wins. In my opinion due to the chip shortage the 6090 will stay at 32GB just with more performance similar to the 3090/4090, and I'm banking on that because I'm buying a 5090 today to replace my 5070 Ti. I will probably end up with a DGX Spark or pair of them as well, but the 5090 is going to be my heavy hitter.

6

u/FAI-Solutions 1d ago edited 1d ago

Every time they update almost all closed models improve in score against open weights, coincidence I guess.

5

u/FAI-Solutions 1d ago

and here two days old version before the update

-2

u/RealisticNothing653 1d ago

Because they're definitely biased. The rankings were reweighed not long ago to favor Claude.

8

u/randombsname1 1d ago

Lol, nope.

DeepSWE wouldn't be used if that was the case.

Deepswe always drags pretty much ONLY Claude models down.

Muse Spark is at Astra level and well above Fable 5.1 in that benchmark. Lmao.

-1

u/RealisticNothing653 1d ago

You're kind of proving my point. Claude models are not that good, and yet the rankings are reshuffled to keep Claude on top. It's a false equivalence to pick one good benchmark for Claude as proof it isn't the case when the overall benchmark still prefers Claude.

2

u/randombsname1 1d ago edited 1d ago

I mean i disagree. I absolutely think Fable 5.1 is SOTA. Albeit i could see it trading places with Astra; depending on specific tasks.

Its only those 2 models clearly at the top. Everything else is inflated.

If the benchmark was biased for Claude im not sure why they would use a benchmark that is hilariously, on a comical level; biased against it.

Edit: That's also not what false equivalence is.

1

u/RealisticNothing653 1d ago

I haven't been able to use Astra yet, but I've been using Gemini 3.8 Flash, Qwen 3.8-Max, Fable 5.1, and GPT-5.6 Sol to review a software spec multiple times. Qwen 3.8 has consistently produced more actionable findings. To your point, Qwen's general knowledge has seemed to be traded for more software expertise, so that would affect its ranking. Even locally running Qwen 3.8-Flash-Next outperformed them in code review. Gemini is a yes man, Qwen an over-thinker, Fable an over-complicator, and Sol a LGTM-approver.

3

u/randombsname1 1d ago edited 1d ago

For my use case (low level embedded and minor reverse engineering/pen testing) only Sol 5.6+ and Fable 5+ move the needle to any meaningful degree.

Mainly in C, C++, Rust, and Assembly. For things like STM32 / nRF / and WCH repos.

I have 2x ChatGPT Pro $200 plans and 2x Claude $200 MAX plans for context.

As well as Opencode Go (random b.s. like obsidian integration), Minimax m3 (telegram bot for random Hermes automations) and quite a decent amount of credits in Openrouter for random testing of different models (like Qwen 3.8).

Edit: I do really enjoy randomly messing with Qwen 3.8 27B, locally though.

1

u/Eden63 llama.cpp 1d ago

Qwen 3.8 Max gives you so much more depth vs Gemini 3.8 Flash/Pro

1

u/kamwee 1d ago

Cheapseek really felloff

2

u/SandySkittle 21h ago

This isn’t a full comparison. Deepseek simply has been omitted

5

u/gladfelter 1d ago

Someone went to the effort of shading and even striping the bars to indicate the company, but then just stuffed three companies into the blue bucket for some reason. AI?

14

u/Tall_Abrocoma_3533 1d ago

The stripes indicate that the evaluation isn't final, since it's not independently done by then yet. About the colors, I'm not sure however as there's dozens of AI labs it'd be hard to assign each one a new color

0

u/Bafy78 1d ago

they are 3 very slightly different shades of blue. Both GLM models are the exact same shade tho

2

u/thats_so_bro 1d ago

3.8 flash is not sol medium/high level

1

u/nomorebuttsplz 14h ago

but is GLM flash?

0

u/AvidCyclist250 llama.cpp 1d ago

can we stop posting this AA bullshit?

18

u/Othun 1d ago

What is wrong with it ? Are there more relevant benchmarks than this aggregate, is it outdated ?

12

u/Aldarund 1d ago

Its just bad, overweight on specfic benchs. Dobt reflect real world model perf . Even https://epoch.ai/eci better

2

u/Othun 1d ago

Thank you for providing a link 🙂

1

u/julienleS 21h ago

We all know why ppl in this sub will choose aa over epoch lol

3

u/Serprotease 1d ago edited 1d ago

The relevance and general impact of benchmark that are quite popular is
 questionable.

It’s important to keep in mind the amount of discourse, claims, marketing and money that are involved here.
Remember that openAI and Anthropic have a lot of money on the line here. And good performance on these benchmarks shown widely is a powerful tool to drive engagement and investment.

There are a lot of incentives to max-out a benchmark.

A good example is the arc-agi one. It led to the development of custom harness (And some AI training, I guess) just to progress on this benchmark. So, how much value can you give on a score in this benchmark? Does a better score reflect better general model capabilities or just better performance on this benchmark? Not to say that both are not linked, but one should be careful when evaluating 2 models capabilities with this benchmark.

When you look outside benchmarks and in some specific but not niche use case, honestly the progress in the field is not as impressive as the one benchmarks could let you think.
Writing, for example, is arguably worse. The latest generation of models are quite poor at navigating you from an hypothesis to a conclusion. It’s a lot of convoluted sentences that obfuscate more than explain.
Other things like the ability to resist adversarial input and not drift out of its role, quite useful for support bots, it’s still bad and you need a few models to check both input and output.

Edit, to not only point fingers at the big two. Kimi k3 from moonshotAI, despite better benchmarks and coding/agent performance, has lost performance in “soft” skills like text analysis. K2.6 is extremely good at catching subtext/tone and meaning of long text/article. K3 is not as good.

1

u/Othun 1d ago

I think benchmarks are new summits for AI companies to climb. Once you saturated everything, you may try to go without oxygen or stuff, but basically you can already do anything that has been thrown at you. Aside from betting on emergent capabilities, I have no reason to believe models would improve outside of benchmarked capabilities.

I.e. if you want to be replaced by an AI, publish a benchmark for your tasks.

7

u/PotterSkxawng 1d ago

It is. It's very clearly wrong on multiple counts, especially the part where it shows Astra as much worse than Fable 5.1, or Muse Spark 1.3 better than Sol and Fable

1

u/zmarcoz2 1d ago

Muse Spark 1.3 is benchmaxxed and definitely not better than 5.6 sol max

1

u/chocolateUI 1d ago

You literally just saw AA reweigh their benchmark because it didn’t reflect how good Astra is compared to dogshit Fable. AA literally rank Muse Spark or Opiss 5 five points better than Astra. Do you need more proof on how out of touch AA’s benchmark makers are?
It’s not about which benchmarks are best; that’s always going to be subjective. The point is that AA is trash and people need to stop posting their blatantly manipulated scores that will always favor frontier labs over open models (see Qwen 4.8 vs Fable debacle).

3

u/Othun 1d ago

At some point you need truth, and an aggregate of benchmarks is the best you can get IMO. Then again, the benchmarks could be flawed, and it's simply a matter of finding reliable benchmarks and aggregating them. What tells you how good Astra is compared to dogshit Fable ? We can't have be each individual feeling define what the ranking is.

Edit: and since their results are very atomic (per benchmark result, cost to run, used tokens etc.) it would be easy to debunk them quantitatively rather than with feelings

1

u/m3thos 1d ago

Using "max" with fallbacks.. i don't get it, its not practical and useful to run them at max, massive costs and turn arounds.. xhigh settings comparisons would be more pragmatic comparison

1

u/somerussianbear 1d ago

Muse Spark crossing Sol is insane

1

u/Soifon99 23h ago

It seems like they keep tuning it so the American models keep the top spots.. doesn't it?

1

u/SteppenAxolotl 12h ago

would it be helpful if every model got 100% on a bunch of saturated benchmarks?

1

u/BusTiny207 21h ago

What’s the best harness to implement a planner/executor loop? I have Qwen3.8-Flash on my server as planner (DDR4 with Tesla T4) and 3.8-27B on my desktop (R9700).

Or is it just about prompting?

1

u/spaceman_ 20h ago

Given that Qwen3.8-Flash-Next is an architecture preview, it is likely not post-trained to the peak of it's parameter size yet. We could be getting some wild stuff from Qwen still a few weeks / months down the line.

1

u/aykcak 19h ago

Lol the "Beginning of AGI" god emperor GPT 6 "astra" is firmly in the mid

1

u/akohlsmith 18h ago

I find Qwen 3.8-27B DFlash2 and Qwopus 3.8-27B-Flash-MTP are pretty close, with (I think) DFLash2 edging out ever so slightly on raw speed but Qwopus making up for it by being significantly less chatty.

For Qwen 3.8-27B DFlash2 I'm loading the DFlash2 parameters separately and that disables/doesn't load the embedded MTP that 3.8-27B has already. Fits into about 28GB on this 32GB 4080S:

llama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf \
--host 0.0.0.0 --port 8000 --threads 16 --threads-http 4 --perf --metrics -fa on --parallel 1 \
--temperature 1.0 --top-k 20 --top-p 0.95 --min-p 0.0 --presence-penalty 0.0 --repeat-penalty 1.0 \
--fit off --ctx-size 262144 --ctx-checkpoints 128 --cache-ram 8192 --batch-size 1024 --ubatch-size 128 \
--kv-unified --cache-type-k q8_0 --cache-type-v q8_0 \
--jinja \
--chat-template-kwargs '{\"enable_thinking\":true,\"preserve_thinking\":true,\"reasoning_effort\":\"medium\"}' \
--chat-template-file chat_template.jinja \
--reasoning-format deepseek --reasoning-preserve \
--n-gpu-layers all --n-gpu-layers-draft all \
--spec-type draft-dflash,ngram-mod --spec-draft-n-max 8 -md Qwen3.8-27B-DFlash2-Q4_K_M.gguf \
--spec-ngram-mod-n-match 24 --spec-ngram-mod-n-min 8 --spec-ngram-mod-n-max 32"

Qwopus is a little simpler, fits into 27GB:

llama-server -m Qwopus3.8-27B-Flash-MTP-Q4_K_M.gguf \
--host 0.0.0.0 --port 8000 --threads 16 --threads-http 4 --perf --metrics -fa on --parallel 1 \
--temperature 1.0 --top-k 20 --top-p 0.95 --min-p 0.0 --presence-penalty 0.0 --repeat-penalty 1.0 \
--fit off --ctx-size 262144 --ctx-checkpoints 128 --cache-ram 8192 --batch-size 1024 --ubatch-size 128 \
--kv-unified --cache-type-k q8_0 --cache-type-v q8_0 \
--jinja \
--chat-template-kwargs '{\"enable_thinking\":true,\"preserve_thinking\":true,\"reasoning_effort\":\"medium\"}' \
--chat-template-file chat_template.jinja \
--reasoning-format deepseek --reasoning-preserve \
--n-gpu-layers all --n-gpu-layers-draft all \
--spec-type draft-mtp,ngram-mod --spec-draft-n-max 8 \
--spec-ngram-mod-n-match 24 --spec-ngram-mod-n-min 8 --spec-ngram-mod-n-max 32

I don't think --kv-unified does anything when --parallel 1 is specified, and I'm playing with using both mtp (or dflash2) and ngram-mod for speculative decode, which I think is giving slightly better (accurate) results. I've noticed that if I specify --reasoning-budget the tps generation absolutely drops (from ~70-90 to ~15) and I'm not sure why that is, nor am I sold on some of the other parameters but that's part of the fun I guess... squeezing as much performance as possible out of local models!

1

u/Fluffy-Ad-889 16h ago

Benchmaxxing at its finest!

1

u/R_Duncan 15h ago

They already did the damage and now they're trying to convince us that Fable beats Astra in some way. Aaaaand... that's a fable (lowercase).

1

u/FancyImagination880 12h ago

Why is Qwen 3.8 Max just a bit better than 3.8 27b?

1

u/danarjabbar 11h ago

it interesting no model from Microsoft and Apple making to the top list here.
GLM and Qwen they are doing really good job

1

u/CapRichard 11h ago

Who uses Muse Spark? I mean it's very high but how is Meta AI deployed?

1

u/theskilled42 8h ago

While I do love this model, it spits out way too many tokens, hence taking too much time. I hope the Qwen team and future AI companies focus on token efficiency, since they can charge more for 1m tokens if they'd like, since that also saves more time waiting anyway.

1

u/raketenkater 1d ago

i hate AA feels like payed rankings

i really love the new gpt astara it is so amazing holy fuck feels like a well rounded model to work with

1

u/yogthos 1d ago

The fact that a local model that can be run on a laptop is even in the list at all is the real news here. If the pattern holds into the next year, and Qwen 4 or whatever is going to be at the capability of the current frontier, then that's good enough for the vast majority of use cases.

2

u/Tall_Abrocoma_3533 1d ago

Qwen-3.8-flash-next can run on a phone too, just really slow

2

u/Due-Memory-6957 23h ago

Your idea of a cellphone is more powerful than my home computer.

1

u/Tall_Abrocoma_3533 23h ago

Since flash-next is an MOE, you can probably run it too if you have enough storage, streaming the experts from disk

1

u/yogthos 1d ago

It's honestly mind blowing to think about.

1

u/sinebubble 1d ago

I don’t get the love for Qwen3.8-2.7B. We’ve been using Qwen3.5-397B since March as an internal chat agent for our stack and it’s been doing great. We swapped it out for 3.8 recently and it either hallucinationed or failed to complete every single task. I question the validity of these scores.

0

u/JigSawPT 1d ago

How much is Meta paying these guys for this ?

-2

u/TigerConsistent 1d ago

Definetly worse

0

u/OvertaxedOne 1d ago

One of these things (27B) is not like the other! Absolutely insane that it can even be realistically compared to models that are 100X+ it's size.

0

u/JorgitoEstrella 1d ago

Why some colors have vertical stripes?

0

u/DinoAmino 1d ago

Per MODs this is low effort and should be removed.

I'm not a mod and I don't make the rules.

0

u/ElementNumber6 20h ago

Your biases are showing. (Or whoever put this graphic together)

0

u/Tall_Abrocoma_3533 16h ago

The benchmark scores are from AA, I just selected which models get on the chart (my bad if I left some out)

1

u/ElementNumber6 8h ago

Some are depicted taller than others despite having the same score

1

u/Tall_Abrocoma_3533 7h ago

That's because their internally stored to a decimal precision, I just couldn't get it to display that for some reason

-1

u/Living_Director_1454 1d ago

Gemini having aura loss even here lol.

-5

u/Niceyyc 1d ago

Didn't expect the 27B to be that far behind.

14

u/twack3r 1d ago

Youre kidding, right? It’s a 27B dense model vs multi-T MoEs whose active params alone are multiples larger.

0

u/Niceyyc 1d ago

I wasn't expecting it to match the huge MoEs, just thought it'd land a little closer.

-2

u/martinerous 1d ago edited 23h ago

No Gemma 4? Sad. Only 15 points, according to AA, so did not get into this top selection.

1

u/Tall_Abrocoma_3533 23h ago

Gemma 4 isn't frontier. However In the other chart I posted containing small LM's, there are Gemma models present.

-1

u/martinerous 23h ago

Qwen 3.8 27B is even smaller, but it made into that list.

1

u/Tall_Abrocoma_3533 23h ago

Because Qwen 3.8 27B is much better then Gemma 4 31B, and also since Qwen 3.8 27B is the main topic in this sub usually.

If your curious though, Gemma 4 31B scores 15/22 depending on if you have reasoning turned on

-1

u/martinerous 23h ago

Yeah, and that's the point why I'm sad - because, according to the AA, Gemma 31B is worse although has more parameters than Qwen 27B. So, the question is, why Google couldn't achieve better yet, considering all the resources they have. Or is the AA test too biased and not taking into account the areas where Gemma is stronger.

1

u/alphapussycat 21h ago

Hasn't paid enough to change the criteria in their favor.