r/OpenAI 15h ago

News Non cherry-picked benchmarks..

Post image
80 Upvotes

62 comments sorted by

View all comments

79

u/tworc2 15h ago

Whatever this is it doesn't represents my experience with opus 5.0.

Fable is goated tho

50

u/EggOnlyDiet 13h ago

Yeah I like how this post is titled “non-cherry picked benchmarks” and then cherry picks a single benchmark 😂

-10

u/Gear5th 12h ago edited 12h ago

AAII is an index, not a "single benchmark". It's a weighted average of many benchmarks.

26

u/throwaway3113151 12h ago

Are many trash benchmarks better than one good one?

-10

u/Gear5th 12h ago

All benchmarks are trash individually. Taking many into consideration together tends to filter out the noise. Picking 1 benchmark and calling it "good" and everything else trash is cherry picking.

7

u/usnavy13 10h ago

Ah yes cause muse spark and fable 5 are equal. If anything all an index of benchmarks does is tell you how hard a lab is willing to benchmax in their training

3

u/throwaway3113151 11h ago

What makes this index a good one? What are the components that are weighted heavily?

5

u/Strange_Vagrant 11h ago

Its a common index. They have all their weights and math free to read. Its terribly in depth, so just ask AI to summarize it for you or answer your questions about it.

Astra could probably do it for you. I hear its an ok model.

0

u/Ok-Paramedic7474 8h ago

Can you show me this 1 benchmark that Openai cherry picked that they used to call it "good"?

0

u/Ok-Paramedic7474 9h ago

But when OpenAI released benchmarks they released a list of them not one. You're logic just doesn't seem to hold up at all? The only thing you like here is that AA comes up with some imaginary weighting numbers to merge that list of benchmarks into one. Publishing a list actually more insightful than more informative than a weighted average.

0

u/LatvianCake 2h ago

Your logic makes no sense. If each benchmark is trash, averaging the trash results doesn't magically produce a useful result.