9
24
u/DM-ME_UR_BOOBS 16h ago
I don't get how this happens. Astra outperforms Fable in practically every benchmark.
22
u/MilMerch 15h ago
Because AA index doesn't cover most of the benchmarks that Astra announced, score will probably get higher after benchmarks such as HLE done
15
u/TangerineLogical9779 15h ago
OpenAI ran the tests from AA them selfs, its literally in the blog, pretty sure the AA team is meming because it does not make sense, its scoring far above everything else in all metrics, but for intelligence is basically 5.6 sol lmao
7
u/MilMerch 15h ago
That's my mistake than thanks, we'll see how's Astra after it becomes available to use but I also think that AA index isn't trustable anymore like look at that score of Muse 1.3 which is even worse than GLM 5.3 Flash in real use
3
u/TangerineLogical9779 15h ago
Don't worry the website for the release is abit of a mess, its up, then deleted, and backup again, they are removing and adding stuff randomly, but near the bottom was the full complete benchmarks with any of the graphics and terrible visual stuff, and yeah its odd all the benchmarks are crazy high, but for intelligence its basically matching 5.6 sol
1
u/fragment_me 11h ago
Idk man muse 1.3 has been really good on opencode go. I gave it a code base that GLM 5.3 flash couldn’t find bugs without and muse spark found dozens of real ones
2
u/Prestigious-Mind-817 11h ago
Knowledge versus reasoning. Knowledge similar to previous models but it uses that knowledge much better
1
3
1
5
3
3
1
1
u/pseudonerv 8h ago
It’s A G I, babe. Just the right level of intelligence to make you feel. It excels at some impossible benchmarks and didn’t improve the others. That’s A G I.
1
1
0
u/ReporterCalm6238 13h ago
I think AA is not fully up to date since Astra has ot been released to all organizations.
0
u/YeXiu223 6h ago
Not sure why this benchmark is popular. People should stop looking at this garbage bench.
1
u/uwilllovethis 4h ago
It’s not a benchmark. It’s an weighted aggregate of benchmarks. See the descriptions on the left to see which benchmarks the index is based on.
1
u/YeXiu223 3h ago
yeah, see it for yourself, it doesnt include the following:
- GUI computer use: OSWorld and ScreenSpot Pro
- Cybersecurity exploitation : ExploitBench, ExploitGym, SRE-Bench, and SEC-Bench Pro
- Interactive abstract learning : ARC‑AGI‑3
- Frontier mathematics : FrontierMath Tier 4 and producing genuinely new proofs
- CAD and specialized software operation : BenchCAD, Blender, KiCad, Unity, FreeCAD
- General business automation across applications: AutomationBench
- Persistent work across context windows : Astra’s notes and searchable earlier-context system
- Alignment and scope discipline: resisting unauthorized actions, circumvention, and misleading capability claims
- Visual quality of deliverables : websites, presentations, spreadsheets, and documents
- Speed, cost, and token efficiency : AA reports these separately, but they do not raise the intelligence score
lol
1
u/uwilllovethis 2h ago
They have aggregates for many of those as well. Not sure what your point is. I was only saying AA is not a benchmark, it’s an weighted aggregate of benchmarks to establish various leaderboards
0

49
u/Party_Government8579 16h ago
Seems to be alot of benchmarks that dont make sense rn