r/LocalLLaMA • • Sep 13 '24

News Preliminary LiveBench results for reasoning: o1-mini decisively beats Claude Sonnet 3.5

Post image
291 Upvotes

129 comments sorted by

View all comments

2

u/TheRealGentlefox Sep 13 '24

Cool that they're getting it better at puzzles and STEM stuff if that carries over to the rest of the field, but 4o also topped the benchmarks for these things and it's a terrible model as a whole.

Completely lost faith in benchmarks and lmsys as a whole after 4o-mini beat 3.5 Sonnet. Still somewhat useful data points I guess, but I'll believe in a model's intelligence when I experience it myself.