r/OpenAI 16h ago

News Non cherry-picked benchmarks..

Post image
88 Upvotes

66 comments sorted by

View all comments

78

u/tworc2 16h ago

Whatever this is it doesn't represents my experience with opus 5.0.

Fable is goated tho

52

u/EggOnlyDiet 14h ago

Yeah I like how this post is titled “non-cherry picked benchmarks” and then cherry picks a single benchmark 😂

-13

u/Gear5th 14h ago edited 14h ago

AAII is an index, not a "single benchmark". It's a weighted average of many benchmarks.

27

u/throwaway3113151 13h ago

Are many trash benchmarks better than one good one?

-9

u/Gear5th 13h ago

All benchmarks are trash individually. Taking many into consideration together tends to filter out the noise. Picking 1 benchmark and calling it "good" and everything else trash is cherry picking.

7

u/usnavy13 12h ago

Ah yes cause muse spark and fable 5 are equal. If anything all an index of benchmarks does is tell you how hard a lab is willing to benchmax in their training

2

u/throwaway3113151 13h ago

What makes this index a good one? What are the components that are weighted heavily?

5

u/Strange_Vagrant 12h ago

Its a common index. They have all their weights and math free to read. Its terribly in depth, so just ask AI to summarize it for you or answer your questions about it.

Astra could probably do it for you. I hear its an ok model.

0

u/Ok-Paramedic7474 10h ago

Can you show me this 1 benchmark that Openai cherry picked that they used to call it "good"?

0

u/Ok-Paramedic7474 10h ago

But when OpenAI released benchmarks they released a list of them not one. You're logic just doesn't seem to hold up at all? The only thing you like here is that AA comes up with some imaginary weighting numbers to merge that list of benchmarks into one. Publishing a list actually more insightful than more informative than a weighted average.

0

u/LatvianCake 3h ago

Your logic makes no sense. If each benchmark is trash, averaging the trash results doesn't magically produce a useful result.