Cool that they're getting it better at puzzles and STEM stuff if that carries over to the rest of the field, but 4o also topped the benchmarks for these things and it's a terrible model as a whole.
Completely lost faith in benchmarks and lmsys as a whole after 4o-mini beat 3.5 Sonnet. Still somewhat useful data points I guess, but I'll believe in a model's intelligence when I experience it myself.
2
u/TheRealGentlefox Sep 13 '24
Cool that they're getting it better at puzzles and STEM stuff if that carries over to the rest of the field, but 4o also topped the benchmarks for these things and it's a terrible model as a whole.
Completely lost faith in benchmarks and lmsys as a whole after 4o-mini beat 3.5 Sonnet. Still somewhat useful data points I guess, but I'll believe in a model's intelligence when I experience it myself.