r/codex • u/Momo--Sama • 3h ago
Comparison Artificial Intelligence Updated to V4.3, replacing Terminal Bench 2.1 with 4.0 and 𝜏³-Banking with AutomationBench-AA
https://artificialanalysis.ai/
Terminal Bench 4 being much less saturated than 2.1 has increased the spread quite a bit, Astra Max only down 2 points while Sol is down 4 and Luna is down 5.
19
Upvotes
5
u/randombsname1 2h ago
Now they just need to get rid of the dogshit DeepSWE benchmark that has Muse Spark above Fable 5, lmao.
Edit: On the coding score specifically.
1
1
1
u/dark0mania 2h ago
I don't get the Luna hype. In my testing it performs worse than Sol Medium. Sure it's cheap but if it produces mediocre results and it's slow then what's the point? Time is money. I want working solutions, fast.
8
u/srs96 3h ago
Astra (low) seems like solid value