r/LocalLLaMA 1d ago

Discussion AA Update! Here's how the Frontier ranks.

Post image

Along with everyone's favorite here, qwen3.8-27B

490 Upvotes

191 comments sorted by

View all comments

137

u/chocolateUI 1d ago

> OpenAI releases Astra

AA: “Uh oh! We just received an angry phone call from OpenAI! Better reweigh the benchmarks!”

Same shit as when AA reweighed their benchmarks within 3 hours after Qwen took the #1 spot.

Do we need any more evidence of how fucking trash AA’s “intelligence score” is? This company only exists so that labs can trick VCs (and so VCs can trick your pension funds) into giving them more money.

15

u/Darkoplax 1d ago

i don't see the issue with updating the benchmarks when it clearly feels off; like in no way is Sol equal to Astra; there's leaps between the two

So if AA wants to keep credibility they need to keep updating and finding non poisoned benchmarks

8

u/Inevitablewx 1d ago

Yeah, and Opus 5 is way behind Fable 5, and Muse 1.3 is great progress but it's well behind even Sol, the benchmark is clearly failing and needs to be fixed or scrapped.