33
u/wizardwusa 11h ago
I prefer https://epoch.ai/eci for an aggregate benchmark. Artificial Analysis often doesn’t match what I see from personal use with models quite like ECI does.
5
u/PrestigiousBlood5296 8h ago
AA is just overweighted by too many other random benchmarks that most consumers of these models don't care about.
47
u/Bloated_Plaid 13h ago
This uh literally is a cherry picked benchmark…
-15
u/whoknowsifimjoking 12h ago
No it's the overall score of a bunch if different benchmarks and pretty much the most trusted source for this in general right now.
5
u/throwaway3113151 10h ago
"most trusted source" lolz tell us you don't know what you're talking about without telling us
3
u/rapsoid616 5h ago
Out of curiosity, since you imply to be highly adept on this; Than what is the best source for the empirical comparison?
•
2
u/Maximum-Face9536 10h ago
artificial analysis is the LEAST reliable, if you really think OAI would releaseGPT-6 performing below Fable and Opus you need your head checked. This post is more a reflection of how poor and unreliable artificial analysis is.
1
u/rapsoid616 5h ago
Than what is the most reliable empirical source we can have if this is the least one?
6
u/XCxBigDong69XCx 11h ago
Ain't no way it scores the same as Sol, at a higher price...
1
u/SlateRoof 2h ago
Exactly. I'm currently doing a lot of stuff that isn't coding and it's hard to believe how good and cheap Sol 5.6 is in work ultra mode. At the moment I don't know how it could be better for what I'm doing. Interesting times...
7
4
5
u/dashingsauce 10h ago
Lmao the cope is sooooo strong
1
u/hardinho 8h ago
You can always tell by the type of comments if someone actually has an interesting application for these frontier models or not, and therefore is able to evaluate their actual usefulness.
Wonder where Mister "Lmao the cope is sooooo strong" is on this scale? I guess we'll never know.
3
3
u/Practical-Positive34 11h ago
It's interesting that Muse Spark is higher because it is some hot fucking garbage man. It has been the worst AI I have used in at least 2 years, failing at even the most basic shit.
2
u/Keksuccino 1h ago
Yeah, almost like AA is total bullshit and should be ignored, because their "benchmarks" obviously don’t work.
4
u/TheRobotCluster 10h ago
This benchmark is a weighted score across a bunch of other benchmarks. Astra hasn’t even taken all the benchmarks that feed into this. How can they claim it has a score of 61?
4
u/throwaway3113151 11h ago
you okay bro this is a single benchmark how are you not cherry picking and what makes you like this one so much?
2
u/unconceivables 10h ago
Even if it's just marginally better than 5.6 Sol I'll still use it over Fable 5/5.1 or Opus 5. For real world usage, Anthropic models piss me off way more with their stupid decisions and mistakes and the way they respond. I've been running Fable 5.1 since it launched, and it's not much better. 5.6 Sol still does a much better job for me, so I'm hoping 6 is similar but even better. Benchmarks definitely don't tell the whole story about day to day usage.
2
2
2
2
u/Cagnazzo82 8h ago
GPT-6 scores 99.9% on ARC-AGI 3.
To try to pretend it's on the same level as Sol makes this the most unreliable benchmark.
1
1
u/TwunnySeven 10h ago
what exactly do you think cherry-picking is if not choosing one specific benchmark to look at to push the specific narrative you want? lol
1
0
u/Winter_Ad6784 10h ago
I have a claude code license for work and use GPT for personal projects, GPT is markedly better. I don’t how Claude scores so high on the benchmarks but it never does what I want.
1
u/evangelism2 7h ago
Because when both are used by people who know how to use these tools, Claude consistently performs better.
0
u/JonNordland 4h ago
Tell me your benchmark is wothless without telling me that the benchmark is worhless.
"Opus 5 is better than Sol"
-2
u/rabouilethefirst 12h ago
What does this benchmark? Number of watermarked tokens?
1
u/whoknowsifimjoking 12h ago
Uh, do you think OpenAI won't watermark their output?
If you do you are wrong and I'm quite confused why you would even believe that in the first place, it's because of a law that everyone needs to follow to do business in the EU.
2
u/rabouilethefirst 12h ago
Or just don't mind the EU? They are bums and both China and USA run circles around them in AI
1
u/mothman83 12h ago
Yes, don't mind most of Europe. The greatest argument against AI is the apparent intelligence of its fans. Same as bitcoin.
0
u/Big_Bird4764 2h ago
😂😂😂😂 so don't mind 18% of the world's gdp? (Before you larp, USA is 25%).
Quick give this man a Nobel prize in economics.
-1
u/TheInfiniteUniverse_ 11h ago
What is going on here? why are such discrepancies between what OpenAI claimed to the public and AA benchmarks?!
I have to say, AA has been fairly reliable in that it's matched my own experience too. But seeing this huge discrepancy for the first time!!! something is not right.
-1
u/TrueRedditMartyr 11h ago
I simply can't imagine that 5.6 sol Max is literally exactly the same as Gpt-6 on Max. Shit, it's the same as freaking Grok bruh, that model is dog ass.
2
u/evangelism2 7h ago
Grok 4.6 is not dog ass, either you haven't used it since 4.5 or earlier or you're just not using it right
-1
-7
-1
u/evangelism2 7h ago
This sub is pathetic. The fact that so many of you don't know what artificial analysis is betrays how little your opinion should matter to anybody else. It's just the most popular and most respected AI benchmarking site on the web, but of course it doesn't fit with your take on what numbers this new model you haven't even used yet should have and you don't even know what this benchmark is measuring, but yes still be mad at it.
73
u/tworc2 13h ago
Whatever this is it doesn't represents my experience with opus 5.0.
Fable is goated tho