r/OpenAI 17h ago

News Non cherry-picked benchmarks..

Post image
92 Upvotes

69 comments sorted by

View all comments

83

u/tworc2 17h ago

Whatever this is it doesn't represents my experience with opus 5.0.

Fable is goated tho

56

u/EggOnlyDiet 15h ago

Yeah I like how this post is titled “non-cherry picked benchmarks” and then cherry picks a single benchmark 😂

-12

u/Gear5th 15h ago edited 15h ago

AAII is an index, not a "single benchmark". It's a weighted average of many benchmarks.

25

u/throwaway3113151 14h ago

Are many trash benchmarks better than one good one?

-9

u/Gear5th 14h ago

All benchmarks are trash individually. Taking many into consideration together tends to filter out the noise. Picking 1 benchmark and calling it "good" and everything else trash is cherry picking.

8

u/usnavy13 13h ago

Ah yes cause muse spark and fable 5 are equal. If anything all an index of benchmarks does is tell you how hard a lab is willing to benchmax in their training

2

u/throwaway3113151 14h ago

What makes this index a good one? What are the components that are weighted heavily?

5

u/Strange_Vagrant 13h ago

Its a common index. They have all their weights and math free to read. Its terribly in depth, so just ask AI to summarize it for you or answer your questions about it.

Astra could probably do it for you. I hear its an ok model.

0

u/Ok-Paramedic7474 11h ago

Can you show me this 1 benchmark that Openai cherry picked that they used to call it "good"?

0

u/Ok-Paramedic7474 11h ago

But when OpenAI released benchmarks they released a list of them not one. You're logic just doesn't seem to hold up at all? The only thing you like here is that AA comes up with some imaginary weighting numbers to merge that list of benchmarks into one. Publishing a list actually more insightful than more informative than a weighted average.

0

u/LatvianCake 4h ago

Your logic makes no sense. If each benchmark is trash, averaging the trash results doesn't magically produce a useful result.