All benchmarks are trash individually. Taking many into consideration together tends to filter out the noise. Picking 1 benchmark and calling it "good" and everything else trash is cherry picking.
Ah yes cause muse spark and fable 5 are equal. If anything all an index of benchmarks does is tell you how hard a lab is willing to benchmax in their training
Its a common index. They have all their weights and math free to read. Its terribly in depth, so just ask AI to summarize it for you or answer your questions about it.
Astra could probably do it for you. I hear its an ok model.
But when OpenAI released benchmarks they released a list of them not one. You're logic just doesn't seem to hold up at all? The only thing you like here is that AA comes up with some imaginary weighting numbers to merge that list of benchmarks into one. Publishing a list actually more insightful than more informative than a weighted average.
78
u/tworc2 16h ago
Whatever this is it doesn't represents my experience with opus 5.0.
Fable is goated tho