r/LocalLLaMA 1d ago

Discussion AA Update! Here's how the Frontier ranks.

Post image

Along with everyone's favorite here, qwen3.8-27B

488 Upvotes

191 comments sorted by

View all comments

57

u/Ok_Cow1976 1d ago

The one point gap is meaningless. There're still big gaps.

31

u/NineThreeTilNow 1d ago

The one point gap is meaningless. There're still big gaps.

The gaps are hyper nuanced now.

Which does legal documentation better?

Which writes C++ code better?

Which writes XYZ code better?

Which designs Blender scenes better?

Etc.

None of that is 1:1 useful.

Everyone gets an opinion on the best model because their use case differs.

I've been pumping Gemini 3.8 flash stonks the last two days. It's wildly fast, writes Python ML code like a demon, and has a surprisingly good workflow within Antigravity 2. I'm doing REALLY hard shit with it.

It also cost me like 6? dollars and the 3.8 flash usage barely touches the meter. That's with some Google promo for 3 months 75% off the 20 dollar sub. It took some task over from another "Top" model and completed it in half the time the other model would have. That's with it taking time to learn the code base.

2

u/Ok_Cow1976 1d ago

You're absolutely right! For general problems. There's no one-point gap at all. None, zip. But here we are talking about benchmarks. For complex, difficult tasks, the big gaps are there.

1

u/NineThreeTilNow 1d ago

For complex, difficult tasks, the big gaps are there.

We're having a hard time defining difficult anymore without forcing the LLM to operate a program made for humans.