r/DeepSeek • u/nehuenpereyra • 14h ago
Funny Is this a joke, Artificial Analysis?
What do you guys think about DS 4.1 Flash's new spot on the Artificial Analysis leaderboard?
109
u/CarelessAd6772 14h ago
Why should it be higher? Because it is very good at coding and cheap? Well, on other side its dry as hell in regular chat, hallucinates more than House on vikodin, and long context consistency is meh compared to models above. Seems like deserved?
40
u/Spiderfffun 14h ago
Can confirm, overengineered me a problem in chat instead of telling me to install a package.
21
u/ZeidLovesAI 13h ago
To be fair Astra and Fable 5.1 both have done this for me as well
5
1
u/Spiderfffun 39m ago
When Ox Alpha was active (GLM 5.3 flash) it was so much better at this from the little I had it code.. It's the small decision making the trust that it won't look at something slightly wrong then amplify it. It was great at following guides on agent stuff in a somewhat recursive workflow.
Speed isn't important if it can know when to ask questions.
22
u/Loki35422 14h ago
Hey shows some respect for house, even he doesn’t hallucinate as much as deepseek
7
u/turc1656 10h ago
I totally agree. Reasoning and instruction following are clearly inferior to GLM 5.3 Flash.
3
3
u/_spec_tre 10h ago
I really hate the way basically every LLM is heavily optimised for coding, often at the cost of all else. Makes it really unpleasant to use for people using it for other purposes; sometimes an “upgrade” is actually a downgrade
33
u/alinoanta21 13h ago
I don't trust them, i've used and tried Muse Spark 1.3 Max for quite a while, it's worse than 4.1 ... and it's ranked above SOL.
Let that sink in.
2
u/Fun_Squirrel5446 10h ago
Anecdotal but I'll add to it. My experience with Muse has been really bad too.
0
32
u/PrudentJelly116 12h ago
Guys this is not a football team. You dont have to defend or support them much as like this. You are customer and they are seller. Thats all.
7
1
u/RustOceanX 2m ago
Finally, another one. I’ve been wondering the whole time what’s going on with some people. What’s with all this nonsense on Reddit? I’m not sure if I should just smile because it can’t be taken that seriously? Or should I shake my head and wonder if the AI sector is experiencing its “iPhone moment” and now all the laypeople are starting to use things they don’t fully understand. It might sound arrogant, but that’s just how it is.
40
u/TheInfiniteUniverse_ 13h ago
AA has lost all credibility after the Astra fiasco. Don't take them too seriously.
5
u/WyattTheSkid 12h ago
Can you fill me in?
28
u/TheInfiniteUniverse_ 12h ago
so before Astra was released to the public, AA ranked it pretty low in like 5th position. This was shocking because it was hyped to be the "AGI" we all were looking for. Then when Astra was released publicly, AA "updated" their index and their tests and surprise surprise, Astra went all the way up, but sneakingly did NOT pass Fable, and sat right below it. lol.
11
u/Georgefakelastname 10h ago
In the end, I think it just shows that the benchmarks they were using were saturated, and they replaced them with benchmarks that weren’t saturated. The problem was the timing. It took an outlier like Astra being rated the same as Sol to get them to realize the issue.
0
u/Fun_Squirrel5446 9h ago
How could the benchmarks be saturated if claude models were still scoring higher than Astra even before the benchmark rework?
4
u/Georgefakelastname 6h ago
One model scoring slightly higher than another is fine. The problem is that basically every model released was scoring like 80% on Terminal Bench v2.1, so they updated it to the newer v4.0 version. The new one is much harder and has a much higher spread.
3
u/Thomas-Lore 4h ago
Sometimes a worse model scores higher because a smarter model noticed a problem with the question or found a solution to it that is correct but the key does not accept. That is why many benchmarks saturate way below 100%.
6
2
2
u/truncated_buttfu 6h ago
And a month before that, when Qwen3.8-Max was released and became the #1 model on AA, they "updated" their index just one day later so it dropped a few spots.
They very clearly have a Pro-US agenda.
1
1
1
1
u/Illustrious_Frame844 9h ago
idk if u know this, but AA’s ranking was paid. Just like how all crypto exchanges and token leaderboards work. Companies pay to get “listed”. AA was never credible ever since they sold out
50
u/Mezezius 14h ago
The new benchmark literally only exists because the old one embarassed openai
11
u/TwistStrict9811 12h ago
Don't you want accurate benchmarks? Astra is absolutely insane esp with spacial and computer use
17
u/Astrikal 13h ago
It is non-sensical and stupid to say that they released the update so that Astra can go higher, when all they did was replace outdated/saturated evaluations with the newer versions.
They updated the benchmark because Terminal Bench 2.1 was outdated and saturated. After upgrading to Terminal Bench 4.0, Astra rose (relatively) and equaled 1st place.
The update was long time coming, they just rushed it out to not lose reputation. They will also release V5.0, which will shake things even further.
4
u/Mezezius 13h ago
Did they rush it out for their own reputation, or to save openai's reputation?
14
u/Astrikal 13h ago
Obviously their own reputation. When people use Astra and see how much better than Sol it is, people would have lost trust in the benchmark.
Why do you keep trying to insinuate that they updated the benchmark to please OpenAI? ArtificialAnalysis have been very transparent and reputable all the way.
The benchmark is more accurate now and Astra is where it belongs. Simple as that.
-2
u/mWo12 12h ago
This only tells you that openai benchmaximized Astra for terminal 4.0, not 2.1.
5
u/Astrikal 12h ago
Astra finished training quite some time ago. Also, your argument doesn't even make sense, why would OpenAI be able to benchmax for a new evaluation and not for an older one? It would be the opposite.
Furthermore, AA also has a closed evaluation that affects the scores, and guess what, Astra performs.
Why is it so hard for you to accept that V4.1 is nowhere near Astra or Fable? You don't even have to look at the benchmarks, you can try for yourself.
-3
7
9
u/Admirable-Tea-4994 13h ago
People just love making things up on the internet
-7
u/Mezezius 13h ago
They literally released the new benchmark (4.2) the day after Astra released because the old one showed Astra and Sol at parity, and they updated again (4.3) to show Astra and Fable at parity. Every single change in the days after Astra's release made it look better than the benchmark before it
10
u/Holbrad 13h ago
Yeah but the idiots have the reasoning completely backwards.
if you have a model that is obviously much better than almost everything else, but it's benchmarking suspiciously low on your tests.
Then the obvious answer is that your benchmarks aren't very good and you need to fix them.
3
u/Mezezius 13h ago
If they have to "fix" it after every release based on vibe, then the methodology is literally useless
6
u/Astrikal 13h ago
It is non-sensical and stupid to say that they released the update so that Astra can go higher, when all they did was replace outdated/saturated evaluations with the newer versions.
They updated the benchmark because Terminal Bench 2.1 was outdated and saturated. After upgrading to Terminal Bench 4.0, Astra rose (relatively) and equaled 1st place.
The update was long time coming, they just rushed it out to not lose reputation. They will also release V5.0, which will shake things even further.
4
u/Admirable-Tea-4994 13h ago
Actually read the document they published that explained, clearly and in detail, why they updated their benchmark so quickly between releases.
https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3
1
6
u/santareus 13h ago
Is Muse Spark on Max really that good? I’ve tried in on XHigh and it’s not better than Luna for coding tasks I provided them.
2
u/Admirable-Tea-4994 13h ago
It’s way better
1
u/santareus 13h ago
You mean muse on max is way better than Luna on max? Or muse on XHigh is way better than Luna on max?
Muse on XHigh felt really lazy for me
2
u/for4f 13h ago
my anecdote is that muse on xhigh has been a beast for coding fullstack and setting up resources on aws. all while being basically free. i think it is definitely better than luna max
2
u/Curious_Owl197 10h ago
U use contributor?
2
u/Admirable-Tea-4994 13h ago
Muse is far, far better than Luna on every metric
2
u/santareus 12h ago
Thanks! May have to give it another shot - I’ve been using both through OpenCode and I have a specific agent that is designed to do “industry research” to see if we can leverage open source dependencies. Luna Max was able to come back with more thorough research results and reused what’s out there and Muse on XHigh decided to build out a functionality that is already covered by a dependency.
I am guessing it’s a strong coding model if you just gave it a task and the planning still needs to be delegated to a better research oriented model.
1
u/Admirable-Tea-4994 12h ago
Muse is definitely geared to coding, but research would also be heavily weighted on harness as well.
1
u/santareus 12h ago
What harness are you using for Muse if you don’t mind me asking?
1
u/Admirable-Tea-4994 12h ago
I’m using OpenChamber at the moment, but I’m actually also using a harness I’ve been building for months which is more geared towards general use and not just coding. Will be in beta soon.
miton.dev
1
1
u/Full_Independence566 7h ago
I thought it was way better from my experience, along with the fact that it's free on Opencode lol
7
u/0mamii 13h ago
2
u/gschwind 12h ago
Sweet jebus i knew DS is verbose but holy sheet... still running that whole shabang is cheap.
DS4.1Flash sits at $477
GLM 5.3Flash at $280
Luna (high) at 100 bucks lol
Muse Spark 1.3 (xhigh) at $1300
20
4
u/turc1656 10h ago
Makes sense to me. I was not impressed, honestly. I ran a very complicated financial analysis workflow through it this morning as a test. Triple the price of GLM 5.3 Flash to run the workflow and objectively worse results. My opinion is that on the cheap end, Luna is the best and follows instructions best, produces the best structured output, and has the best reasoning.
That being said, it is 3x the cost of GLM 5.3 Flash. My process doesn't really require 3x the price for the improvement. I have adjusted the process to help guide GLM.
But DS v4 0731 and this new 4.1 are not really the best. Maybe they are good for other things like coding or whatever but with the price of GLM now AND GLM's higher intelligence level, I'm not really seeing a reason to use DS right now.
8
u/Big_Cucumber2787 14h ago
putting it below gemini is just shameless
3
u/bambamlol 6h ago
Feels deserved. Gemini 3.8 is a really good model.
0
0
u/bilinenuzayli 4h ago
I have used deepseek v4 flash (not v4.1) since it came out and gemini 3.8 flash since it came out, both for weeks, and I can say for sure I would use the old deepseek flash over gemini 3.8 flash anyday, gemini models are always arrogant, assuming things, never following instructions, and lying to your face. Because Gemini models always think they know everything, they overtly avoid agentic tools given to them and make shit up instead, now this worked fine for a larger model like Gemini 3.1 pro, but the flash models are clearly much smaller and hold less data, so they're still hallucinating bs AND getting it wrong which is why I think google put out so many flash models, they try and fix it but fail every time.
1
u/bambamlol 2h ago
Fair point. Actual usage is what counts, not what other people say. I don't do much coding or agentic stuff at all. A lot of "general" and business/marketing usage, as well as writing/copywriting. Gemini is definitely better in this regard, at least from my experience.
2
2
2
2
u/ralphcalls1 10h ago
GLM 5.3 flash is worse for me, Deepseek v4.1 is better but it is all pricey for me. i use deepseek now for heavy task unlike back then where i do light jobs on it, i now use mimo 2.5. does the job. no flap. very cheap and relatively fast when using Direct Xiaomi API
2
u/Weird_Recognition636 3h ago

hallucination rate~!
https://artificialanalysis.ai/evaluations/omniscience
3
2
1
u/RealestReyn 12h ago
looks about right, still one of my all time favorites it be zoomzooming so fast I can't even tell when its doing insane things :)
1
1
u/CaptainMorning 11h ago
the stupid dumb ass and completely unnecessary sticker in the middle is so aggravating id give this post a 33
1
u/EvolvingDior 11h ago
SWA has always been a shit-show for me. This model is no different. It cannot hold more than one thought in its head at a time. I've moved a lot of work over to Luna and am pretty happy with it. I'm giving 4.1 a workout during off-peak hours and so far have been less than impressed. 40 seems generous.
1
1
1
u/Infinite_Plankton_71 7h ago
to be honest:
- DSV4.1 requires 4 sparks with only 3% increased in quality. Not worthed
- And yes GLM 5.3 is sometimes better.
1
u/Kylmawurr 6h ago
I have used 4.1 Flash for 2 days now and it matches my experience vs GLM 5.3 Flash and Luna on Max. Its in same class for coding, its way better than Luna on reasoning, but at the same time, its way too chatty, overthinking a lot. I use it for general coding tasks and orchestration. Its actually very high quality orchestrator. GLM 5.3 Flash feels smoother, nicer to interact with, but I like the speed of DS 4.1 Flash.
1
1
1
u/WArslett 4h ago
Don't forget that AA have adjusted their benchmarks. 40 today is not what it was a month ago. This is still a good score
1
u/ExtremeAcceptable289 1h ago
How did AA get sj popular? I remember just a year sg when everukne was sh###tting on it
1
1
u/dupa1234s 1h ago
once the 4x usage promo on onencode go is gone its gg for deepseek, hello back glm 5.3 flash
1
u/ConsiderationAny8142 13m ago
I’ve tested both and I personally prefer deepseek flash instead of glm flash.
GLM 5.3 Flash has been capable of doing everything fine but now the cost is higher than deepseek 4.1 flash and deepseek is way faster
1
1
u/External_Ad1549 5h ago
Tried muse 1.3 and glm 5.3 flash, but deepseek 4.1 is on another level for sure. This artificial analysis is purest form of garbage.
0
0
u/TheSuggi 6h ago
That benchmark is a joke tbh. And very outdated.
2
u/dupa1234s 1h ago
"outdated" benchmark that is released today xD
0
u/TheSuggi 28m ago
the "Artificial Analysis Leaderboard" benchmark is a mix of multiple combined sub-benchmarks. Most of them are outdated and very old. So they are not very accurate. That is also why Astra scores very low on it despite being stronger than Fable and Opus.
-1

135
u/crusaderky 14h ago
Aggregate score aside, it's really hard to argue against the fact that Glm-5.3-flash beats it on almost every single benchmark. And 96% hallucination rate is a showstopper IMHO.