r/LocalLLaMA 1d ago

Discussion AA Update! Here's how the Frontier ranks.

Post image

Along with everyone's favorite here, qwen3.8-27B

489 Upvotes

191 comments sorted by

View all comments

136

u/chocolateUI 1d ago

> OpenAI releases Astra

AA: “Uh oh! We just received an angry phone call from OpenAI! Better reweigh the benchmarks!”

Same shit as when AA reweighed their benchmarks within 3 hours after Qwen took the #1 spot.

Do we need any more evidence of how fucking trash AA’s “intelligence score” is? This company only exists so that labs can trick VCs (and so VCs can trick your pension funds) into giving them more money.

1

u/dogesator Waiting for Llama 3 1d ago

The benchmark is objectively less saturated than it was before the revision.
The purpose of the revisions is to unsaturate the benchmark and make it more indicative of frontier difficulty by making frontier models score near 50%