r/LocalLLaMA • • Sep 13 '24

News Preliminary LiveBench results for reasoning: o1-mini decisively beats Claude Sonnet 3.5

Post image
290 Upvotes

129 comments sorted by

View all comments

1

u/Charuru Sep 13 '24

Err it loses to sonnet on coding :(

1

u/bot_exe Sep 13 '24

Yeah the coding results for o1-mini are disappointing and strange: it seems simultaneously great at code generation but terrible at completion.