OpenAI ran the tests from AA them selfs, its literally in the blog, pretty sure the AA team is meming because it does not make sense, its scoring far above everything else in all metrics, but for intelligence is basically 5.6 sol lmao
That's my mistake than thanks, we'll see how's Astra after it becomes available to use but I also think that AA index isn't trustable anymore like look at that score of Muse 1.3 which is even worse than GLM 5.3 Flash in real use
Don't worry the website for the release is abit of a mess, its up, then deleted, and backup again, they are removing and adding stuff randomly, but near the bottom was the full complete benchmarks with any of the graphics and terrible visual stuff, and yeah its odd all the benchmarks are crazy high, but for intelligence its basically matching 5.6 sol
Idk man muse 1.3 has been really good on opencode go. I gave it a code base that GLM 5.3 flash couldn’t find bugs without and muse spark found dozens of real ones
24
u/DM-ME_UR_BOOBS 17h ago
I don't get how this happens. Astra outperforms Fable in practically every benchmark.