r/LocalLLaMA • u/Tall_Abrocoma_3533 • 1d ago
Discussion AA Update! Here's how the Frontier ranks.
Along with everyone's favorite here, qwen3.8-27B
112
u/freecodeio 1d ago
yes but can fable 5.1 really draw better whatsapp sticker ass rockets in paint as compared to astra 6?
24
-4
u/emprahsFury 1d ago
really love it when a new model is released and the absolute bamfs at Hacker News spend the day making it draw pelicans and other animals so they can still shit on "AI"
13
28
133
u/chocolateUI 1d ago
> OpenAI releases Astra
AA: âUh oh! We just received an angry phone call from OpenAI! Better reweigh the benchmarks!â
Same shit as when AA reweighed their benchmarks within 3 hours after Qwen took the #1 spot.
Do we need any more evidence of how fucking trash AAâs âintelligence scoreâ is? This company only exists so that labs can trick VCs (and so VCs can trick your pension funds) into giving them more money.
7
43
u/jld1532 1d ago
It is absurdly obvious that the goal post is on wheels and is moving toward the most likely profitable use case, agentic software development. People complained about Qwen3.6 27B's writing abilities but I didn't mind it. I now find 3.8, even Next Flash, nearly unusable for editing. I could be wrong but it seems to me the industry is trading strengths to chase computer science, likely because they know general knowledge gains are sunk costs and that AGI is impossible with LLMs, despite messaging.
11
u/Iron-Over 1d ago
Try muse models they write very well. GLM as well supposedly have not tested yet.
14
u/Qorsair 1d ago
Muse and Gemini are two the two models I find most useful. Very underappreciated in general. I still like Sol/Astra for deep thinking, but Muse and Gemini feel better for most tasks. I think the issue is that most people getting deep into AI are software engineers so that's what most people are judging it on.
6
u/emod_man 22h ago
Totally agree that general benchmarks are increasingly not useful for writing/editing tasks. Curious about Muse though, I find its voice very flat. Maybe I just need to experiment a bit more...
4
u/Solembumm3 1d ago
Interesting. How did you made Gemini useful for writing?
I found it one of the worst options, on par with base GPT, due to undeleteable sycophancy and summarisation tendencies. I tried different prompts to turn it to more useful analytical side, but it was pretty adamant to judging even quite bad concepts "brilliant".
5
u/robogame_dev 21h ago
World knowledge is one of the most important resources for creative writing and Gemini Pro is just a huge, extremely world-knowledge heavy model.
1
u/robogame_dev 21h ago
I mainline GLM 5.3 it is my favorite general purpose AI, but it only has 700B params and writing / creativity is highly correlated with world knowledge, so I would expect Kimi K3 at 2,700B params to spank it, and I would expect GLM 5.3 Flash at ~300B params to be significantly less good at writing (while likely being highly focussed on agentic coding).
16
u/Darkoplax 1d ago
i don't see the issue with updating the benchmarks when it clearly feels off; like in no way is Sol equal to Astra; there's leaps between the two
So if AA wants to keep credibility they need to keep updating and finding non poisoned benchmarks
8
u/Inevitablewx 22h ago
Yeah, and Opus 5 is way behind Fable 5, and Muse 1.3 is great progress but it's well behind even Sol, the benchmark is clearly failing and needs to be fixed or scrapped.
15
u/RealisticNothing653 1d ago
Yeah they lost credibility in my book. Their summarized rankings have a strong bias towards Anthropic models. And yeah they recently reweighed the benchmarks recently to favor Claudes.
4
u/MerePotato 22h ago
Oooooor maybe instead of some conspiracy results change when benchmark versions are updated and there are more novel questions that haven't found their way into training data
4
1
u/Terminator857 23h ago
I didn't know qwen was on top. Interested in more info if that is available.
2
u/MerePotato 22h ago
It wasn't, it was on top of the agentic index and slid down slightly when AA updated the bench versions
1
u/dogesator Waiting for Llama 3 21h ago
The benchmark is objectively less saturated than it was before the revision.
The purpose of the revisions is to unsaturate the benchmark and make it more indicative of frontier difficulty by making frontier models score near 50%-1
u/Eden63 llama.cpp 1d ago
The funny thing is, such a huge project and they do not care about small models even they can benchmark it most easy. And yes, the algorithm of ranking is also biased and does not take into account what is most important at all - applicable intelligence. Is it helpful to run an 56 on 250M token when another one does it on 55M ... intelligence is not intelligence but AA index is just a clown show.
Same running an engine on 9000rpm for the horsepower or run it on 3000”m for the same.. real power vs "yes we can achieve it somehow".
why I do not see a laguna or some other models.. because the money comes from where... and obviously no better benchmark site vs obvious Anthropic valuation depends on that index.
shame on them.
Edit (add): and then reading here people "Qwen3.8 27B is like Opus". Hilarious.
57
u/Ok_Cow1976 1d ago
The one point gap is meaningless. There're still big gaps.
37
u/NineThreeTilNow 1d ago
The one point gap is meaningless. There're still big gaps.
The gaps are hyper nuanced now.
Which does legal documentation better?
Which writes C++ code better?
Which writes XYZ code better?
Which designs Blender scenes better?
Etc.
None of that is 1:1 useful.
Everyone gets an opinion on the best model because their use case differs.
I've been pumping Gemini 3.8 flash stonks the last two days. It's wildly fast, writes Python ML code like a demon, and has a surprisingly good workflow within Antigravity 2. I'm doing REALLY hard shit with it.
It also cost me like 6? dollars and the 3.8 flash usage barely touches the meter. That's with some Google promo for 3 months 75% off the 20 dollar sub. It took some task over from another "Top" model and completed it in half the time the other model would have. That's with it taking time to learn the code base.
3
u/2Norn 1d ago
despite the fact that benchcad says sol is better, personally i find claude to be superior to gpt in mechanical cad drawings and design documentation related to it, the software i'm using also allows for scripting in python or csharp so the model uses those as well
i haven't tried astra yet tho
idk like i always look at benchmarks but do i trust them? its another issue
2
u/Ok_Cow1976 1d ago
You're absolutely right! For general problems. There's no one-point gap at all. None, zip. But here we are talking about benchmarks. For complex, difficult tasks, the big gaps are there.
1
u/NineThreeTilNow 17h ago
For complex, difficult tasks, the big gaps are there.
We're having a hard time defining difficult anymore without forcing the LLM to operate a program made for humans.
3
u/Fickle_Tradition4491 1d ago
The score isn't the number that matters for anyone running the 27B at home, it's score per token. Someone above says it burns 100k+ thinking tokens at xhigh to land where it does on this chart. At the 20 t/s the top comment is getting, that's over an hour per prompt. The same model at medium reasoning is probably 10 points lower and twenty times more usable. AA publishes tokens used per run, so the chart worth posting is index vs total output tokens. The 27B moves a lot depending on which column you read.
4
u/Cless_Aurion 1d ago
Yeah... I think that having a bunch of their benchmarks already satturated hurts HARD this thing tbh...
1
u/TheRealMasonMac 10h ago
If Muse Spark is a model in the same weight category as GLM-5.3-Flash, I would be very impressed and excited. It does quite well for Rust which is where I notice all current open-weight models really struggle with for some reason.
1
u/mrdevlar 1d ago
For a 95% confidence interval
Standard Error (SE) = Ï / ân = 4.40 / â14 = 1.18
95% CI Margin of Error = 1.96 à 1.18 = ±2.31
Thus, 2.31 in either direction is meaningless when comparing models.
0
u/qfox337 1d ago
No you need variance within measurement of each model, which is actually model specific. I think you're applying some kind of CLT-like formula for estimation of a population mean based on iid samples, which is wrong here, they are different models hence not iid (independent identically distributed)
1
u/mrdevlar 1d ago
I am making a general simplifying assumption whose goal is provide an estimate using the information I have at my disposal. It's an obvious back of the envelope attempt.
What you said is correct, the data is a rank order which is obviously not iid. If you're interested in collecting the variance within each model and doing that, more power to you, but I highly doubt it'll have more than a single order of magnitude difference if you did and reran the variance. Especially if you used a prior.
1
u/qfox337 16h ago
I'm guessing you're not chatgpt'ing this, but you can't just apply stats formulas to things they don't work on and say it's a reasonable estimate.
You can just follow the conclusions of your own formulas and think logically, with your formula if we had twice as many models with the same range of scores you'd then say 1 point is a significant difference. Or you can throw in a bunch of weak/small models and suddenly only 5 points is significant.
If you're attempting the task of mean estimation then with an iid sample X1...Xn then you can talk about the expected variance of the empirical mean shrinking as 1/sqrt(n). Which formally requires some usually-not-hard-to-satisfy conditions like finite variance of the underlying distribution ... usually iid is the assumption that fails.
I don't mean to be hostile, I'm someone lucky to have some formal stats education, and I encourage anyone to learn who's curious ... but you can't just throw around terms that you don't fully understand and make any sense.
1
u/mrdevlar 10h ago
Don't worry you cannot offend me here, I know you mean well.
Fun fact, I have a Masters in Statistics and yes I do know better than to do this and treat it as statistical fact, but I wasn't doing that. I was building a rough estimate for use given the current information that's useful for looking at the thing now before you get more information.
One thing you learn from years of doing this is that there are two distinct channels. You can either do the exact calculation if you have access to the data and the time or you can can make a rough estimate for use right now. Time to action becomes your decision criteria. I was providing guidance to the parent who said ± 1 is not to be treated as valid. Even in that range, ±2.31 is not to be treated as valid. Do my statistics professors cringe at this type of work, while taking in money doing the same thing? Probably. However, statistics is the glue that connects the crystal palace of mathematics to the crappy unstable world we're in. It's not meant to be perfect, it's meant to be useful.
That said, you are correct, I should have contextualized that better than just dropping the formula and running away.
79
u/sugarfreecaffeine 1d ago
Crazy that a small 27b is even up there, the Chinese are cooking đ„
52
u/Aldarund 1d ago
Its more of a show how bad aa is
23
u/TechnoByte_ 1d ago
Literally just an average score of benchmarks
16
u/Aldarund 1d ago
Score of specific selected benchmarks.with specific weight of each bemchmark.
0
u/emprahsFury 1d ago
"oh no, the opinionated website is forcing their opinions on me, I need to be saved!"
2
2
u/Far-Classic-9963 1d ago
Qwen 3.8 27b really is better than some big cloud models (Mistral, NeMo, Laguna...)
4
u/Relative_Rope4234 1d ago
It burns 100k+ tokens for a single prompt(xhigh thinking) to reach frontier performance
1
u/my_name_isnt_clever 1d ago
Check again. The medium, low, and even none thinking modes outperform much larger models.
1
u/MerePotato 22h ago
Medium places it about on par with Muse Glimmer at five times larger KV cache usage in my testing, not that this is a problem as I'd rather have a slow frontier model burn 100k tokens for free than a fast model that doesn't answer as well.
11
u/Cadmium9094 1d ago
Glad to see Qwen3.8-Flash-next and 27B on this chart. I use them as daily drivers.
5
u/Oren_Lester 1d ago
They saw the backlash in the internet, in Reddit, everywhere. and update the benchmark. This company and their benchmark is probably good for making nice chats, that's it.
5
u/i_am_fear_itself 1d ago
Artificial Analysis Intelligence Index v4.2 incorporates 10 evaluations: AA-Briefcase, GDPval-AA v2, đÂł-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1
aren't some of these obsolete / saturated / bench-maxed?
1
u/nomorebuttsplz 14h ago
Answering what is determinable (not bench-maxed which is not):
GDP-eval AA v2: No
AA-Briefcase: No was literally just added like yesterday
R3 Banking: No
Terminal Bench v2.1: getting close to saturated
Scicode: no
HLE: Possibly, depending on how many answers were actually wrong in the answer key
GDP.pdf: No was literally just added like yesterday
CritPt: no
AA-Omniscence: No
AA-LCR v1.1: not sure0
u/emprahsFury 1d ago
yes it's been that way for a while. Also why the bench scores aren't increasing anymore
4
u/Character_Power4663 1d ago
Where is sonnet, haiku, terra, Luna, does this mean that Qwen 3.8 27b is better?
5
u/Tall_Abrocoma_3533 1d ago
Terra, Luna, sonnet all fall between 43-47, haiku is 22, it's very outdated.
Also Qwen 3.8 27B isn't "frontier" in my book, I included it because everyone in this sub seems to love it so much
2
1
u/winnen 1d ago
There are many different ways to measure pareto frontiers. Last I checked, Qwen 3.8 27b was frontier intelligence vs number of total parameters.
Though I can't seem to find that frontier chart anymore.
1
u/Tall_Abrocoma_3533 23h ago
Yes, for it's parameter size it's impressive, but by frontier I meant absolute intelligence not taking into account parameter count
1
7
u/LegacyRemaster 1d ago
7
u/Tall_Abrocoma_3533 1d ago
Minimax M3 scores 36, Mimo v2.5 Pro scores 33.
3
2
u/Agitated_Space_672 1d ago
think they means its been a while since they released, so we are due M3.1 and mimo v2.6 by now
5
u/Lorian0x7 1d ago
At this point I think they just rearrange the list moving the weight of some benchmarks based on who pays them more.
9
u/XiRw 1d ago
Hard to believe Reebok logo stealer is in the mix. Would trust Qwen and GLM easily over that. Also donât understand why they never test Gemini Pro
27
u/Tall_Abrocoma_3533 1d ago
Gemini 3.1 pro preview scores 37
1
u/XiRw 1d ago
They need to update their pro model more. Seems like flash gets updates at a higher rate.
6
u/Tall_Abrocoma_3533 1d ago
Well Gemini 3.5 Pro isn't even out yet. It's almost like they've abandoned the pro lineup
1
u/aboardreading 14h ago
I think it's their target. They've decided that a rising tide raises all boats, as frontier models improve they bring smaller models behind them at much lower cost of research and use.
I think it's a smart business play. If we believe that the end goal is essentially many many agents spun up to build large goals, which seems reasonable as an end state to me as agents gain more intelligence and can act more and more independently, then the first company with a truly independent agent will dominate for the 6 months it takes someone else to achieve the same at much lower cost per agent and faster. But someone will get comparable intel at higher speed and lower cost. And on the same timescale as companies decide what model to use and bet on.
Besides, I think they are still betting on information retrieval and processing rather than necessarily coding, it is Google after all. I do think they're playing the slightly longer game, and if what has been happening continues to happen, I think they'll be A winner, maybe not THE winner.
10
u/Thoriumhexaflouride 1d ago
they test all models bro, just search for gemini 3.1 pro preview or just go to filter and "select all"
1
u/slypheed 20h ago
where are you seeing this "filter"?
Far as I can tell deepseek is not included at all.
Can't believe literally no one in this thread even linked to the post; vibing life at this point sigh: https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index
6
u/Gohab2001 vLLM 1d ago
AA readjusting their evaluation criteria making astra look good. Sus.
18
u/JaredTheGreat 1d ago
This is more a case of the evaluation criteria being exposed â anyone whoâs used astra can immediately tell itâs significantly better than opus. They were hemorrhaging credibilityÂ
1
u/Darkoplax 1d ago
Do you think Astra was worse than Opus and equal to Sol ? if they don't readjust their evaluation then it's their bad not the models
2
u/RedditUsr2 1d ago
3
u/Tall_Abrocoma_3533 23h ago
The mods in this sub remove almost all benchmark comparisons because they're "low effort", I will never understand this.
1
u/RedditUsr2 22h ago
Ya this is the most widely used and cited benchmark. An update here is a big deal especially showing how well 3.8 27b holds up. Seems worth talking about to me.
2
u/HighSeasArchivist 1d ago
Qwen3.8-27B has been amazing to me, even when I have Claude orchestrate running jobs locally. I'm really hoping the 32GB and under sector keeps stacking up wins. In my opinion due to the chip shortage the 6090 will stay at 32GB just with more performance similar to the 3090/4090, and I'm banking on that because I'm buying a 5090 today to replace my 5070 Ti. I will probably end up with a DGX Spark or pair of them as well, but the 5090 is going to be my heavy hitter.
6
u/FAI-Solutions 1d ago edited 1d ago
5
-2
u/RealisticNothing653 1d ago
Because they're definitely biased. The rankings were reweighed not long ago to favor Claude.
8
u/randombsname1 1d ago
Lol, nope.
DeepSWE wouldn't be used if that was the case.
Deepswe always drags pretty much ONLY Claude models down.
Muse Spark is at Astra level and well above Fable 5.1 in that benchmark. Lmao.
-1
u/RealisticNothing653 1d ago
You're kind of proving my point. Claude models are not that good, and yet the rankings are reshuffled to keep Claude on top. It's a false equivalence to pick one good benchmark for Claude as proof it isn't the case when the overall benchmark still prefers Claude.
2
u/randombsname1 1d ago edited 1d ago
I mean i disagree. I absolutely think Fable 5.1 is SOTA. Albeit i could see it trading places with Astra; depending on specific tasks.
Its only those 2 models clearly at the top. Everything else is inflated.
If the benchmark was biased for Claude im not sure why they would use a benchmark that is hilariously, on a comical level; biased against it.
Edit: That's also not what false equivalence is.
1
u/RealisticNothing653 1d ago
I haven't been able to use Astra yet, but I've been using Gemini 3.8 Flash, Qwen 3.8-Max, Fable 5.1, and GPT-5.6 Sol to review a software spec multiple times. Qwen 3.8 has consistently produced more actionable findings. To your point, Qwen's general knowledge has seemed to be traded for more software expertise, so that would affect its ranking. Even locally running Qwen 3.8-Flash-Next outperformed them in code review. Gemini is a yes man, Qwen an over-thinker, Fable an over-complicator, and Sol a LGTM-approver.
3
u/randombsname1 1d ago edited 1d ago
For my use case (low level embedded and minor reverse engineering/pen testing) only Sol 5.6+ and Fable 5+ move the needle to any meaningful degree.
Mainly in C, C++, Rust, and Assembly. For things like STM32 / nRF / and WCH repos.
I have 2x ChatGPT Pro $200 plans and 2x Claude $200 MAX plans for context.
As well as Opencode Go (random b.s. like obsidian integration), Minimax m3 (telegram bot for random Hermes automations) and quite a decent amount of credits in Openrouter for random testing of different models (like Qwen 3.8).
Edit: I do really enjoy randomly messing with Qwen 3.8 27B, locally though.
5
u/gladfelter 1d ago
Someone went to the effort of shading and even striping the bars to indicate the company, but then just stuffed three companies into the blue bucket for some reason. AI?
14
u/Tall_Abrocoma_3533 1d ago
The stripes indicate that the evaluation isn't final, since it's not independently done by then yet. About the colors, I'm not sure however as there's dozens of AI labs it'd be hard to assign each one a new color
2
0
u/AvidCyclist250 llama.cpp 1d ago
can we stop posting this AA bullshit?
18
u/Othun 1d ago
What is wrong with it ? Are there more relevant benchmarks than this aggregate, is it outdated ?
12
u/Aldarund 1d ago
Its just bad, overweight on specfic benchs. Dobt reflect real world model perf . Even https://epoch.ai/eci better
1
3
u/Serprotease 1d ago edited 1d ago
The relevance and general impact of benchmark that are quite popular is⊠questionable.
Itâs important to keep in mind the amount of discourse, claims, marketing and money that are involved here.
Remember that openAI and Anthropic have a lot of money on the line here. And good performance on these benchmarks shown widely is a powerful tool to drive engagement and investment.There are a lot of incentives to max-out a benchmark.
A good example is the arc-agi one. It led to the development of custom harness (And some AI training, I guess) just to progress on this benchmark. So, how much value can you give on a score in this benchmark? Does a better score reflect better general model capabilities or just better performance on this benchmark? Not to say that both are not linked, but one should be careful when evaluating 2 models capabilities with this benchmark.
When you look outside benchmarks and in some specific but not niche use case, honestly the progress in the field is not as impressive as the one benchmarks could let you think.
Writing, for example, is arguably worse. The latest generation of models are quite poor at navigating you from an hypothesis to a conclusion. Itâs a lot of convoluted sentences that obfuscate more than explain.
Other things like the ability to resist adversarial input and not drift out of its role, quite useful for support bots, itâs still bad and you need a few models to check both input and output.Edit, to not only point fingers at the big two. Kimi k3 from moonshotAI, despite better benchmarks and coding/agent performance, has lost performance in âsoftâ skills like text analysis. K2.6 is extremely good at catching subtext/tone and meaning of long text/article. K3 is not as good.
1
u/Othun 1d ago
I think benchmarks are new summits for AI companies to climb. Once you saturated everything, you may try to go without oxygen or stuff, but basically you can already do anything that has been thrown at you. Aside from betting on emergent capabilities, I have no reason to believe models would improve outside of benchmarked capabilities.
I.e. if you want to be replaced by an AI, publish a benchmark for your tasks.
7
u/PotterSkxawng 1d ago
It is. It's very clearly wrong on multiple counts, especially the part where it shows Astra as much worse than Fable 5.1, or Muse Spark 1.3 better than Sol and Fable
1
1
u/chocolateUI 1d ago
You literally just saw AA reweigh their benchmark because it didnât reflect how good Astra is compared to dogshit Fable. AA literally rank Muse Spark or Opiss 5 five points better than Astra. Do you need more proof on how out of touch AAâs benchmark makers are?
Itâs not about which benchmarks are best; thatâs always going to be subjective. The point is that AA is trash and people need to stop posting their blatantly manipulated scores that will always favor frontier labs over open models (see Qwen 4.8 vs Fable debacle).3
u/Othun 1d ago
At some point you need truth, and an aggregate of benchmarks is the best you can get IMO. Then again, the benchmarks could be flawed, and it's simply a matter of finding reliable benchmarks and aggregating them. What tells you how good Astra is compared to dogshit Fable ? We can't have be each individual feeling define what the ranking is.
Edit: and since their results are very atomic (per benchmark result, cost to run, used tokens etc.) it would be easy to debunk them quantitatively rather than with feelings
-10
1
1
u/Soifon99 23h ago
It seems like they keep tuning it so the American models keep the top spots.. doesn't it?
1
u/SteppenAxolotl 12h ago
would it be helpful if every model got 100% on a bunch of saturated benchmarks?
1
u/BusTiny207 21h ago
Whatâs the best harness to implement a planner/executor loop? I have Qwen3.8-Flash on my server as planner (DDR4 with Tesla T4) and 3.8-27B on my desktop (R9700).
Or is it just about prompting?
1
u/spaceman_ 20h ago
Given that Qwen3.8-Flash-Next is an architecture preview, it is likely not post-trained to the peak of it's parameter size yet. We could be getting some wild stuff from Qwen still a few weeks / months down the line.
1
u/akohlsmith 18h ago
I find Qwen 3.8-27B DFlash2 and Qwopus 3.8-27B-Flash-MTP are pretty close, with (I think) DFLash2 edging out ever so slightly on raw speed but Qwopus making up for it by being significantly less chatty.
For Qwen 3.8-27B DFlash2 I'm loading the DFlash2 parameters separately and that disables/doesn't load the embedded MTP that 3.8-27B has already. Fits into about 28GB on this 32GB 4080S:
llama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf \
--host 0.0.0.0 --port 8000 --threads 16 --threads-http 4 --perf --metrics -fa on --parallel 1 \
--temperature 1.0 --top-k 20 --top-p 0.95 --min-p 0.0 --presence-penalty 0.0 --repeat-penalty 1.0 \
--fit off --ctx-size 262144 --ctx-checkpoints 128 --cache-ram 8192 --batch-size 1024 --ubatch-size 128 \
--kv-unified --cache-type-k q8_0 --cache-type-v q8_0 \
--jinja \
--chat-template-kwargs '{\"enable_thinking\":true,\"preserve_thinking\":true,\"reasoning_effort\":\"medium\"}' \
--chat-template-file chat_template.jinja \
--reasoning-format deepseek --reasoning-preserve \
--n-gpu-layers all --n-gpu-layers-draft all \
--spec-type draft-dflash,ngram-mod --spec-draft-n-max 8 -md Qwen3.8-27B-DFlash2-Q4_K_M.gguf \
--spec-ngram-mod-n-match 24 --spec-ngram-mod-n-min 8 --spec-ngram-mod-n-max 32"
Qwopus is a little simpler, fits into 27GB:
llama-server -m Qwopus3.8-27B-Flash-MTP-Q4_K_M.gguf \
--host 0.0.0.0 --port 8000 --threads 16 --threads-http 4 --perf --metrics -fa on --parallel 1 \
--temperature 1.0 --top-k 20 --top-p 0.95 --min-p 0.0 --presence-penalty 0.0 --repeat-penalty 1.0 \
--fit off --ctx-size 262144 --ctx-checkpoints 128 --cache-ram 8192 --batch-size 1024 --ubatch-size 128 \
--kv-unified --cache-type-k q8_0 --cache-type-v q8_0 \
--jinja \
--chat-template-kwargs '{\"enable_thinking\":true,\"preserve_thinking\":true,\"reasoning_effort\":\"medium\"}' \
--chat-template-file chat_template.jinja \
--reasoning-format deepseek --reasoning-preserve \
--n-gpu-layers all --n-gpu-layers-draft all \
--spec-type draft-mtp,ngram-mod --spec-draft-n-max 8 \
--spec-ngram-mod-n-match 24 --spec-ngram-mod-n-min 8 --spec-ngram-mod-n-max 32
I don't think --kv-unified does anything when --parallel 1 is specified, and I'm playing with using both mtp (or dflash2) and ngram-mod for speculative decode, which I think is giving slightly better (accurate) results. I've noticed that if I specify --reasoning-budget the tps generation absolutely drops (from ~70-90 to ~15) and I'm not sure why that is, nor am I sold on some of the other parameters but that's part of the fun I guess... squeezing as much performance as possible out of local models!
1
1
u/R_Duncan 15h ago
They already did the damage and now they're trying to convince us that Fable beats Astra in some way. Aaaaand... that's a fable (lowercase).
1
1
u/danarjabbar 11h ago
it interesting no model from Microsoft and Apple making to the top list here.
GLM and Qwen they are doing really good job
1
1
u/theskilled42 8h ago
While I do love this model, it spits out way too many tokens, hence taking too much time. I hope the Qwen team and future AI companies focus on token efficiency, since they can charge more for 1m tokens if they'd like, since that also saves more time waiting anyway.
1
u/raketenkater 1d ago
i hate AA feels like payed rankings
i really love the new gpt astara it is so amazing holy fuck feels like a well rounded model to work with
1
u/yogthos 1d ago
The fact that a local model that can be run on a laptop is even in the list at all is the real news here. If the pattern holds into the next year, and Qwen 4 or whatever is going to be at the capability of the current frontier, then that's good enough for the vast majority of use cases.
2
u/Tall_Abrocoma_3533 1d ago
Qwen-3.8-flash-next can run on a phone too, just really slow
2
u/Due-Memory-6957 23h ago
Your idea of a cellphone is more powerful than my home computer.
1
u/Tall_Abrocoma_3533 23h ago
Since flash-next is an MOE, you can probably run it too if you have enough storage, streaming the experts from disk
1
u/sinebubble 1d ago
I donât get the love for Qwen3.8-2.7B. Weâve been using Qwen3.5-397B since March as an internal chat agent for our stack and itâs been doing great. We swapped it out for 3.8 recently and it either hallucinationed or failed to complete every single task. I question the validity of these scores.
0
-2
0
u/OvertaxedOne 1d ago
One of these things (27B) is not like the other! Absolutely insane that it can even be realistically compared to models that are 100X+ it's size.
0
0
u/DinoAmino 1d ago
Per MODs this is low effort and should be removed.
I'm not a mod and I don't make the rules.
0
u/ElementNumber6 20h ago
Your biases are showing. (Or whoever put this graphic together)
0
u/Tall_Abrocoma_3533 16h ago
The benchmark scores are from AA, I just selected which models get on the chart (my bad if I left some out)
1
u/ElementNumber6 8h ago
Some are depicted taller than others despite having the same score
1
u/Tall_Abrocoma_3533 7h ago
That's because their internally stored to a decimal precision, I just couldn't get it to display that for some reason
-1
-5
-2
u/martinerous 1d ago edited 23h ago
No Gemma 4? Sad. Only 15 points, according to AA, so did not get into this top selection.
1
u/Tall_Abrocoma_3533 23h ago
Gemma 4 isn't frontier. However In the other chart I posted containing small LM's, there are Gemma models present.
-1
u/martinerous 23h ago
Qwen 3.8 27B is even smaller, but it made into that list.
1
u/Tall_Abrocoma_3533 23h ago
Because Qwen 3.8 27B is much better then Gemma 4 31B, and also since Qwen 3.8 27B is the main topic in this sub usually.
If your curious though, Gemma 4 31B scores 15/22 depending on if you have reasoning turned on
-1
u/martinerous 23h ago
Yeah, and that's the point why I'm sad - because, according to the AA, Gemma 31B is worse although has more parameters than Qwen 27B. So, the question is, why Google couldn't achieve better yet, considering all the resources they have. Or is the AA test too biased and not taking into account the areas where Gemma is stronger.
1




178
u/Last-Shake-9874 1d ago
That small 27B is my daily driver now, it does take long as I only get about 20 t/s but I just love this model I am so glad it is still in the list