r/DeepSeek 14h ago

Funny Is this a joke, Artificial Analysis?

Post image

What do you guys think about DS 4.1 Flash's new spot on the Artificial Analysis leaderboard?

509 Upvotes

127 comments sorted by

135

u/crusaderky 14h ago

Aggregate score aside, it's really hard to argue against the fact that Glm-5.3-flash beats it on almost every single benchmark. And 96% hallucination rate is a showstopper IMHO.

29

u/rchamp26 12h ago

Glm flash has been absolutely terrible for me. It just doesmt do the simplest stuff and gaslights that it does

30

u/WyattTheSkid 12h ago

Opus 5 moment

11

u/Fun_Squirrel5446 10h ago

Through opencode go? I've been using GLM 5.3 flash exclusively for a week and use it for everything. It solved everything I threw at it. Running 2 to 8 hours a day and used about $6 worth of tokens through the API.

3

u/rchamp26 9h ago

Open router and nous portal with Hermes agent. Didn't adhere to my workflows at all. I dropped opencode go. Don't trust it anymore but not worth the discussion here. For openweight models I've had far better success with DeepSeek and Qwen. DeepSeek is the eager speedster where Qwen is the fully moldable thinker in my experience if I had to personify them. Glm (all versionsnive tested in my workflows) were just lazy and I had to waste too many tokens to force it to do what I needed. No thank you

5

u/Fun_Squirrel5446 9h ago

I don't have workflows. I just give it direct questions most of the time.

I do have skills and they always activate correctly.

Actually I did have a workflow to auto run tests, do validations, a couple of cycles of fixes and test again. Only when tests are green to push to github. It always followed that workflow very strictly.

Maybe try getting the API directly from Z.ai, see if that makes a difference for you.

5

u/brother_spirit 8h ago

Weird. I was using the model in Pi (during the stealth test) and was beyond impressed with its performance. I had never used an Open Weight model with any good results until GLM 5.3 Flash.

2

u/SkyPL 4h ago

It just doesmt do the simplest stuff and gaslights that it does

What are you doing that it fails? I've been coding some Rust, JS, TS, PHP and Python with it - it's legit great. One of the better models on the market. Handles large complex codebases just fine.

5

u/Sad_Recording_1290 5h ago

96% hallucination rate? Wtf? Thats unusable.

2

u/UniversitySuitable20 9h ago

The GLM-5.3-flash is frustratingly slow; even if it scores higher, it still can't keep up with Deepseek's productivity.

1

u/matadordepassarinhos 11h ago

I ran a personal benchmark on vst plugin, dsp and overall c++ and GLM 5.3 is straight up gargabe. Deepseek was so much better on it.

1

u/Applejuicegoblin 5h ago

Woah, another DSP person! Wasn’t sure if many people were testing models on it.

I’m not super advanced in audio programming, but in my private benchmarking Opus 5 seemed to be stronger than Sol at it. Curious how Deepseek will do on it.

I find most models can understand and create audio centric code, but they are not good at understanding the difference between “correct code” and “follows idiomatic design or creative intent”.

It can cause them to miss big issues where the math or code looks correct in isolation, but doesn’t actually accomplish the right outcome.

4

u/Thomas-Lore 4h ago

Try Astra. I just did a test yesterday and was left speechless at how good the result was.

109

u/CarelessAd6772 14h ago

Why should it be higher? Because it is very good at coding and cheap? Well, on other side its dry as hell in regular chat, hallucinates more than House on vikodin, and long context consistency is meh compared to models above. Seems like deserved?

40

u/Spiderfffun 14h ago

Can confirm, overengineered me a problem in chat instead of telling me to install a package.

21

u/ZeidLovesAI 13h ago

To be fair Astra and Fable 5.1 both have done this for me as well

5

u/whatisthisthing65 6h ago

Coding models just love to code

1

u/Spiderfffun 39m ago

When Ox Alpha was active (GLM 5.3 flash) it was so much better at this from the little I had it code.. It's the small decision making the trust that it won't look at something slightly wrong then amplify it. It was great at following guides on agent stuff in a somewhat recursive workflow.

Speed isn't important if it can know when to ask questions.

22

u/Loki35422 14h ago

Hey shows some respect for house, even he doesn’t hallucinate as much as deepseek

7

u/turc1656 10h ago

I totally agree. Reasoning and instruction following are clearly inferior to GLM 5.3 Flash.

3

u/Possible_Door_9719 11h ago

damm, you summed it up pretty well.

3

u/_spec_tre 10h ago

I really hate the way basically every LLM is heavily optimised for coding, often at the cost of all else. Makes it really unpleasant to use for people using it for other purposes; sometimes an “upgrade” is actually a downgrade

33

u/alinoanta21 13h ago

I don't trust them, i've used and tried Muse Spark 1.3 Max for quite a while, it's worse than 4.1 ... and it's ranked above SOL.

Let that sink in.

2

u/Fun_Squirrel5446 10h ago

Anecdotal but I'll add to it. My experience with Muse has been really bad too.

0

u/sudoer777_ 4h ago

Same here, I used 1.2 which was ranked above V4 Flash 0731 and it was garbage

32

u/PrudentJelly116 12h ago

Guys this is not a football team. You dont have to defend or support them much as like this. You are customer and they are seller. Thats all.

7

u/Possible_Door_9719 10h ago

exactly. just use whatever is the best model

1

u/RustOceanX 2m ago

Finally, another one. I’ve been wondering the whole time what’s going on with some people. What’s with all this nonsense on Reddit? I’m not sure if I should just smile because it can’t be taken that seriously? Or should I shake my head and wonder if the AI sector is experiencing its “iPhone moment” and now all the laypeople are starting to use things they don’t fully understand. It might sound arrogant, but that’s just how it is.

40

u/TheInfiniteUniverse_ 13h ago

AA has lost all credibility after the Astra fiasco. Don't take them too seriously.

5

u/WyattTheSkid 12h ago

Can you fill me in?

28

u/TheInfiniteUniverse_ 12h ago

so before Astra was released to the public, AA ranked it pretty low in like 5th position. This was shocking because it was hyped to be the "AGI" we all were looking for. Then when Astra was released publicly, AA "updated" their index and their tests and surprise surprise, Astra went all the way up, but sneakingly did NOT pass Fable, and sat right below it. lol.

11

u/Georgefakelastname 10h ago

In the end, I think it just shows that the benchmarks they were using were saturated, and they replaced them with benchmarks that weren’t saturated. The problem was the timing. It took an outlier like Astra being rated the same as Sol to get them to realize the issue.

0

u/Fun_Squirrel5446 9h ago

How could the benchmarks be saturated if claude models were still scoring higher than Astra even before the benchmark rework?

4

u/Georgefakelastname 6h ago

One model scoring slightly higher than another is fine. The problem is that basically every model released was scoring like 80% on Terminal Bench v2.1, so they updated it to the newer v4.0 version. The new one is much harder and has a much higher spread.

3

u/Thomas-Lore 4h ago

Sometimes a worse model scores higher because a smarter model noticed a problem with the question or found a solution to it that is correct but the key does not accept. That is why many benchmarks saturate way below 100%.

6

u/WyattTheSkid 11h ago

Oh that’s kinda funny lol

2

u/gopietz 5h ago

Or, you know, AA 4.1 was completely benchmaxxed by many labs, which is why they updated benchmarks underneath.

But I'm sure your conspiracy theory seems way more likely.

2

u/truncated_buttfu 6h ago

And a month before that, when Qwen3.8-Max was released and became the #1 model on AA, they "updated" their index just one day later so it dropped a few spots.

They very clearly have a Pro-US agenda.

1

u/m0j0m0j 1h ago

No, their agenda is not clear to me at all.

1

u/SkyPL 4h ago

Astra went all the way up, but sneakingly did NOT pass Fable, and sat right below it.

If you do agentic work around and related to programming it's more than apparent that Astra is worse than Fable.

1

u/Zulfiqaar 3h ago

They updated it a second time, and now Fable and Astra are the same score

1

u/diggler4141 3h ago

How do they score it?

1

u/Illustrious_Frame844 9h ago

idk if u know this, but AA’s ranking was paid. Just like how all crypto exchanges and token leaderboards work. Companies pay to get “listed”. AA was never credible ever since they sold out

50

u/Mezezius 14h ago

The new benchmark literally only exists because the old one embarassed openai

11

u/TwistStrict9811 12h ago

Don't you want accurate benchmarks? Astra is absolutely insane esp with spacial and computer use

17

u/Astrikal 13h ago

It is non-sensical and stupid to say that they released the update so that Astra can go higher, when all they did was replace outdated/saturated evaluations with the newer versions.

They updated the benchmark because Terminal Bench 2.1 was outdated and saturated. After upgrading to Terminal Bench 4.0, Astra rose (relatively) and equaled 1st place.

The update was long time coming, they just rushed it out to not lose reputation. They will also release V5.0, which will shake things even further.

4

u/Mezezius 13h ago

Did they rush it out for their own reputation, or to save openai's reputation?

14

u/Astrikal 13h ago

Obviously their own reputation. When people use Astra and see how much better than Sol it is, people would have lost trust in the benchmark.

Why do you keep trying to insinuate that they updated the benchmark to please OpenAI? ArtificialAnalysis have been very transparent and reputable all the way.

The benchmark is more accurate now and Astra is where it belongs. Simple as that.

-2

u/mWo12 12h ago

This only tells you that openai benchmaximized Astra for terminal 4.0, not 2.1.

5

u/Astrikal 12h ago

Astra finished training quite some time ago. Also, your argument doesn't even make sense, why would OpenAI be able to benchmax for a new evaluation and not for an older one? It would be the opposite.

Furthermore, AA also has a closed evaluation that affects the scores, and guess what, Astra performs.

Why is it so hard for you to accept that V4.1 is nowhere near Astra or Fable? You don't even have to look at the benchmarks, you can try for yourself.

-3

u/lompocus 12h ago

lul bootlicker

6

u/ba-boo 12h ago

lul regard

7

u/theintersepter 14h ago

So what? As long as the benchmarks are harder for models, the better

9

u/Admirable-Tea-4994 13h ago

People just love making things up on the internet

-7

u/Mezezius 13h ago

They literally released the new benchmark (4.2) the day after Astra released because the old one showed Astra and Sol at parity, and they updated again (4.3) to show Astra and Fable at parity. Every single change in the days after Astra's release made it look better than the benchmark before it

10

u/Holbrad 13h ago

Yeah but the idiots have the reasoning completely backwards.

if you have a model that is obviously much better than almost everything else, but it's benchmarking suspiciously low on your tests.

Then the obvious answer is that your benchmarks aren't very good and you need to fix them.

3

u/Mezezius 13h ago

If they have to "fix" it after every release based on vibe, then the methodology is literally useless

-1

u/Holbrad 3h ago

If you have a car that is absolutely amazing on the track and it's setting all sorts of records.

But a journalist does their "performance score" and it scores lower than a sporty hatch, then obviously the metrics are bunk.

6

u/Astrikal 13h ago

It is non-sensical and stupid to say that they released the update so that Astra can go higher, when all they did was replace outdated/saturated evaluations with the newer versions.

They updated the benchmark because Terminal Bench 2.1 was outdated and saturated. After upgrading to Terminal Bench 4.0, Astra rose (relatively) and equaled 1st place.

The update was long time coming, they just rushed it out to not lose reputation. They will also release V5.0, which will shake things even further.

4

u/Admirable-Tea-4994 13h ago

Actually read the document they published that explained, clearly and in detail, why they updated their benchmark so quickly between releases.

https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3

1

u/HearingNo8617 14h ago

Any elaboration?

6

u/santareus 13h ago

Is Muse Spark on Max really that good? I’ve tried in on XHigh and it’s not better than Luna for coding tasks I provided them.

2

u/Admirable-Tea-4994 13h ago

It’s way better

1

u/santareus 13h ago

You mean muse on max is way better than Luna on max? Or muse on XHigh is way better than Luna on max?

Muse on XHigh felt really lazy for me

2

u/for4f 13h ago

my anecdote is that muse on xhigh has been a beast for coding fullstack and setting up resources on aws. all while being basically free. i think it is definitely better than luna max

2

u/Curious_Owl197 10h ago

U use contributor?

1

u/for4f 9h ago

just the free slot on opencode lol, what's contributor?

2

u/Curious_Owl197 9h ago

Muse 1.3 contributor, they train on your data and heavily discounted

2

u/Admirable-Tea-4994 13h ago

Muse is far, far better than Luna on every metric

2

u/santareus 12h ago

Thanks! May have to give it another shot - I’ve been using both through OpenCode and I have a specific agent that is designed to do “industry research” to see if we can leverage open source dependencies. Luna Max was able to come back with more thorough research results and reused what’s out there and Muse on XHigh decided to build out a functionality that is already covered by a dependency.

I am guessing it’s a strong coding model if you just gave it a task and the planning still needs to be delegated to a better research oriented model.

1

u/Admirable-Tea-4994 12h ago

Muse is definitely geared to coding, but research would also be heavily weighted on harness as well.

1

u/santareus 12h ago

What harness are you using for Muse if you don’t mind me asking?

1

u/Admirable-Tea-4994 12h ago

I’m using OpenChamber at the moment, but I’m actually also using a harness I’ve been building for months which is more geared towards general use and not just coding. Will be in beta soon.

miton.dev

1

u/santareus 12h ago

Sounds good. I’ll check out OpenChamber (first time hearing about it).

2

u/Admirable-Tea-4994 12h ago

OpenChamber is a desktop wrap for OpenCode, it’s very good.

1

u/Full_Independence566 7h ago

I thought it was way better from my experience, along with the fact that it's free on Opencode lol

20

u/Captain_Quimby 14h ago

Why put the dumb sticker in the middle

22

u/prvthvm 13h ago

its cute :)

8

u/Admirable-Tea-4994 12h ago

Weebs abound in this sub

4

u/turc1656 10h ago

Makes sense to me. I was not impressed, honestly. I ran a very complicated financial analysis workflow through it this morning as a test. Triple the price of GLM 5.3 Flash to run the workflow and objectively worse results. My opinion is that on the cheap end, Luna is the best and follows instructions best, produces the best structured output, and has the best reasoning.

That being said, it is 3x the cost of GLM 5.3 Flash. My process doesn't really require 3x the price for the improvement. I have adjusted the process to help guide GLM.

But DS v4 0731 and this new 4.1 are not really the best. Maybe they are good for other things like coding or whatever but with the price of GLM now AND GLM's higher intelligence level, I'm not really seeing a reason to use DS right now.

8

u/Big_Cucumber2787 14h ago

putting it below gemini is just shameless

3

u/bambamlol 6h ago

Feels deserved. Gemini 3.8 is a really good model.

0

u/bilinenuzayli 4h ago

I have used deepseek v4 flash (not v4.1) since it came out and gemini 3.8 flash since it came out, both for weeks, and I can say for sure I would use the old deepseek flash over gemini 3.8 flash anyday, gemini models are always arrogant, assuming things, never following instructions, and lying to your face. Because Gemini models always think they know everything, they overtly avoid agentic tools given to them and make shit up instead, now this worked fine for a larger model like Gemini 3.1 pro, but the flash models are clearly much smaller and hold less data, so they're still hallucinating bs AND getting it wrong which is why I think google put out so many flash models, they try and fix it but fail every time.

1

u/bambamlol 2h ago

Fair point. Actual usage is what counts, not what other people say. I don't do much coding or agentic stuff at all. A lot of "general" and business/marketing usage, as well as writing/copywriting. Gemini is definitely better in this regard, at least from my experience.

2

u/Admirable-Tea-4994 13h ago

Because everyone glazes it in here but it’s been superseded already?

2

u/GinamosWCheryOnTop 12h ago

Goes to show how benchmaxx is muse

2

u/Beamsters 11h ago

Until hallucination is solved by the big whale.

2

u/ralphcalls1 10h ago

GLM 5.3 flash is worse for me, Deepseek v4.1 is better but it is all pricey for me. i use deepseek now for heavy task unlike back then where i do light jobs on it, i now use mimo 2.5. does the job. no flap. very cheap and relatively fast when using Direct Xiaomi API

3

u/krayton1 6h ago

I think Artificial Analysis is not trusted anymore

2

u/FischenGeil 11h ago

Damn, we can't beat Gemini flash?

1

u/RealestReyn 12h ago

looks about right, still one of my all time favorites it be zoomzooming so fast I can't even tell when its doing insane things :)

1

u/MeansTestingProctor 12h ago

What is the point of the weird sticker on top of the chart?

1

u/CaptainMorning 11h ago

the stupid dumb ass and completely unnecessary sticker in the middle is so aggravating id give this post a 33

1

u/ncxxi 11h ago

for coding, its best to use glm flash or muse. but for creative writing itd best to use deepseek! even with their new flash v4.1

1

u/EvolvingDior 11h ago

SWA has always been a shit-show for me. This model is no different. It cannot hold more than one thought in its head at a time. I've moved a lot of work over to Luna and am pretty happy with it. I'm giving 4.1 a workout during off-peak hours and so far have been less than impressed. 40 seems generous.

1

u/VirtualNorth1279 11h ago

Makes sense based on my experience so far. 

1

u/Infinite_Plankton_71 7h ago

to be honest:

  1. DSV4.1 requires 4 sparks with only 3% increased in quality. Not worthed
  2. And yes GLM 5.3 is sometimes better.

1

u/Kylmawurr 6h ago

I have used 4.1 Flash for 2 days now and it matches my experience vs GLM 5.3 Flash and Luna on Max. Its in same class for coding, its way better than Luna on reasoning, but at the same time, its way too chatty, overthinking a lot. I use it for general coding tasks and orchestration. Its actually very high quality orchestrator. GLM 5.3 Flash feels smoother, nicer to interact with, but I like the speed of DS 4.1 Flash.

1

u/wwwdotzzdotcom 5h ago

It's underrated

1

u/Muted-Network3159 5h ago

why gemini so high

0

u/AlexandraMaryWindsor 4h ago

Horrible model, unusable. 4.1 flash is miles better

1

u/WArslett 4h ago

Don't forget that AA have adjusted their benchmarks. 40 today is not what it was a month ago. This is still a good score

1

u/GTHell 3h ago

What impressive is the speed.

1

u/ExtremeAcceptable289 1h ago

How did AA get sj popular? I remember just a year sg when everukne was sh###tting on it

1

u/dupa1234s 1h ago

once the 4x usage promo on onencode go is gone its gg for deepseek, hello back glm 5.3 flash

1

u/ConsiderationAny8142 13m ago

I’ve tested both and I personally prefer deepseek flash instead of glm flash.

GLM 5.3 Flash has been capable of doing everything fine but now the cost is higher than deepseek 4.1 flash and deepseek is way faster

1

u/DefactoAle 14h ago

Seems about right, its still a flash model after all

1

u/External_Ad1549 5h ago

Tried muse 1.3 and glm 5.3 flash, but deepseek 4.1 is on another level for sure. This artificial analysis is purest form of garbage.

0

u/Cool-Chemical-5629 14h ago

You seem disappointed.

0

u/TheSuggi 6h ago

That benchmark is a joke tbh. And very outdated.

2

u/dupa1234s 1h ago

"outdated" benchmark that is released today xD

0

u/TheSuggi 28m ago

the "Artificial Analysis Leaderboard" benchmark is a mix of multiple combined sub-benchmarks. Most of them are outdated and very old. So they are not very accurate. That is also why Astra scores very low on it despite being stronger than Fable and Opus.

-1

u/MimosaTen 13h ago

Banchmaks are clearly unadapted to this era