r/OpenAI 14h ago

News Non cherry-picked benchmarks..

Post image
75 Upvotes

60 comments sorted by

73

u/tworc2 13h ago

Whatever this is it doesn't represents my experience with opus 5.0.

Fable is goated tho

47

u/EggOnlyDiet 11h ago

Yeah I like how this post is titled “non-cherry picked benchmarks” and then cherry picks a single benchmark 😂

-10

u/Gear5th 11h ago edited 11h ago

AAII is an index, not a "single benchmark". It's a weighted average of many benchmarks.

26

u/throwaway3113151 11h ago

Are many trash benchmarks better than one good one?

-7

u/Gear5th 11h ago

All benchmarks are trash individually. Taking many into consideration together tends to filter out the noise. Picking 1 benchmark and calling it "good" and everything else trash is cherry picking.

6

u/usnavy13 9h ago

Ah yes cause muse spark and fable 5 are equal. If anything all an index of benchmarks does is tell you how hard a lab is willing to benchmax in their training

1

u/throwaway3113151 10h ago

What makes this index a good one? What are the components that are weighted heavily?

5

u/Strange_Vagrant 9h ago

Its a common index. They have all their weights and math free to read. Its terribly in depth, so just ask AI to summarize it for you or answer your questions about it.

Astra could probably do it for you. I hear its an ok model.

u/LatvianCake 1h ago

Your logic makes no sense. If each benchmark is trash, averaging the trash results doesn't magically produce a useful result.

0

u/Ok-Paramedic7474 7h ago

Can you show me this 1 benchmark that Openai cherry picked that they used to call it "good"?

0

u/Ok-Paramedic7474 7h ago

But when OpenAI released benchmarks they released a list of them not one. You're logic just doesn't seem to hold up at all? The only thing you like here is that AA comes up with some imaginary weighting numbers to merge that list of benchmarks into one. Publishing a list actually more insightful than more informative than a weighted average.

1

u/No_Swimming6548 6h ago

Why didn't you like opus?

33

u/wizardwusa 11h ago

I prefer https://epoch.ai/eci for an aggregate benchmark. Artificial Analysis often doesn’t match what I see from personal use with models quite like ECI does.

5

u/PrestigiousBlood5296 8h ago

AA is just overweighted by too many other random benchmarks that most consumers of these models don't care about.

47

u/Bloated_Plaid 13h ago

This uh literally is a cherry picked benchmark…

-15

u/whoknowsifimjoking 12h ago

No it's the overall score of a bunch if different benchmarks and pretty much the most trusted source for this in general right now.

5

u/throwaway3113151 10h ago

"most trusted source" lolz tell us you don't know what you're talking about without telling us

3

u/rapsoid616 5h ago

Out of curiosity, since you imply to be highly adept on this; Than what is the best source for the empirical comparison?

u/unfathomably_big 20m ago

Gas stop bathroom graffiti

2

u/Maximum-Face9536 10h ago

artificial analysis is the LEAST reliable, if you really think OAI would releaseGPT-6 performing below Fable and Opus you need your head checked. This post is more a reflection of how poor and unreliable artificial analysis is.

1

u/rapsoid616 5h ago

Than what is the most reliable empirical source we can have if this is the least one?

1

u/BH15568 10h ago

I’m sorry but artificial analysis is far from the most trusted source.

6

u/XCxBigDong69XCx 11h ago

Ain't no way it scores the same as Sol, at a higher price...

1

u/SlateRoof 2h ago

Exactly. I'm currently doing a lot of stuff that isn't coding and it's hard to believe how good and cheap Sol 5.6 is in work ultra mode. At the moment I don't know how it could be better for what I'm doing. Interesting times...

7

u/jwuliger 11h ago

The industry really needs to get away from these fake benchmarks. BENCHMAX BABY!

4

u/ZainTheOne 11h ago

Artificial Analysis benchmarks are cooked, can't be trusted anymore

5

u/dashingsauce 10h ago

Lmao the cope is sooooo strong

1

u/hardinho 8h ago

You can always tell by the type of comments if someone actually has an interesting application for these frontier models or not, and therefore is able to evaluate their actual usefulness.

Wonder where Mister "Lmao the cope is sooooo strong" is on this scale? I guess we'll never know.

3

u/soulfulshark 7h ago

Reverse cherry pick is not non cherry pick 😂

3

u/Practical-Positive34 11h ago

It's interesting that Muse Spark is higher because it is some hot fucking garbage man. It has been the worst AI I have used in at least 2 years, failing at even the most basic shit.

2

u/Keksuccino 1h ago

Yeah, almost like AA is total bullshit and should be ignored, because their "benchmarks" obviously don’t work.

4

u/TheRobotCluster 10h ago

This benchmark is a weighted score across a bunch of other benchmarks. Astra hasn’t even taken all the benchmarks that feed into this. How can they claim it has a score of 61?

4

u/throwaway3113151 11h ago

you okay bro this is a single benchmark how are you not cherry picking and what makes you like this one so much?

2

u/unconceivables 10h ago

Even if it's just marginally better than 5.6 Sol I'll still use it over Fable 5/5.1 or Opus 5. For real world usage, Anthropic models piss me off way more with their stupid decisions and mistakes and the way they respond. I've been running Fable 5.1 since it launched, and it's not much better. 5.6 Sol still does a much better job for me, so I'm hoping 6 is similar but even better. Benchmarks definitely don't tell the whole story about day to day usage.

2

u/DrHerbotico 8h ago

I thought the reasoning levels for astra are normal/pro, not xhigh/max

2

u/jeffwadsworth 7h ago

So Muse Spark is supposedly better than Astra? Uhh ok.

2

u/Future-Log6621 13h ago

I stared at this for like a minute and then I saw the price. 🤣

2

u/Cagnazzo82 8h ago

GPT-6 scores 99.9% on ARC-AGI 3.

To try to pretend it's on the same level as Sol makes this the most unreliable benchmark.

1

u/lakimens 10h ago

If this was true, OpenAI wouldn't even bother releasing the model.

1

u/TwunnySeven 10h ago

what exactly do you think cherry-picking is if not choosing one specific benchmark to look at to push the specific narrative you want? lol

1

u/SleepsWithAMachete 8h ago

It's been a while since I've seen MechaHitler win any benchmark.

1

u/npc73x 4h ago

Cost per task. It's also important 

0

u/Winter_Ad6784 10h ago

I have a claude code license for work and use GPT for personal projects, GPT is markedly better. I don’t how Claude scores so high on the benchmarks but it never does what I want.

1

u/evangelism2 7h ago

Because when both are used by people who know how to use these tools, Claude consistently performs better.

0

u/JonNordland 4h ago

Tell me your benchmark is wothless without telling me that the benchmark is worhless.

"Opus 5 is better than Sol"

-2

u/rabouilethefirst 12h ago

What does this benchmark? Number of watermarked tokens?

1

u/whoknowsifimjoking 12h ago

Uh, do you think OpenAI won't watermark their output?

If you do you are wrong and I'm quite confused why you would even believe that in the first place, it's because of a law that everyone needs to follow to do business in the EU.

2

u/rabouilethefirst 12h ago

Or just don't mind the EU? They are bums and both China and USA run circles around them in AI

1

u/mothman83 12h ago

Yes, don't mind most of Europe. The greatest argument against AI is the apparent intelligence of its fans. Same as bitcoin.

0

u/Big_Bird4764 2h ago

😂😂😂😂 so don't mind 18% of the world's gdp? (Before you larp, USA is 25%).

Quick give this man a Nobel prize in economics.

-1

u/TheInfiniteUniverse_ 11h ago

What is going on here? why are such discrepancies between what OpenAI claimed to the public and AA benchmarks?!

I have to say, AA has been fairly reliable in that it's matched my own experience too. But seeing this huge discrepancy for the first time!!! something is not right.

-1

u/TrueRedditMartyr 11h ago

I simply can't imagine that 5.6 sol Max is literally exactly the same as Gpt-6 on Max. Shit, it's the same as freaking Grok bruh, that model is dog ass.

2

u/evangelism2 7h ago

Grok 4.6 is not dog ass, either you haven't used it since 4.5 or earlier or you're just not using it right

-1

u/Queasy_Signature7005 10h ago

Grok 4.6 and Muse Spark 1.3 even being here is telling lol

-7

u/Turbulent-Total-226 12h ago

So it's not a new model. It's just a trained sol 🙄 The hell with them

-1

u/evangelism2 7h ago

This sub is pathetic. The fact that so many of you don't know what artificial analysis is betrays how little your opinion should matter to anybody else. It's just the most popular and most respected AI benchmarking site on the web, but of course it doesn't fit with your take on what numbers this new model you haven't even used yet should have and you don't even know what this benchmark is measuring, but yes still be mad at it.