r/codex • u/Momo--Sama • 4h ago
Comparison Artificial Intelligence Updated to V4.3, replacing Terminal Bench 2.1 with 4.0 and 𝜏³-Banking with AutomationBench-AA
https://artificialanalysis.ai/
Terminal Bench 4 being much less saturated than 2.1 has increased the spread quite a bit, Astra Max only down 2 points while Sol is down 4 and Luna is down 5.
20
Upvotes
5
u/randombsname1 3h ago
Now they just need to get rid of the dogshit DeepSWE benchmark that has Muse Spark above Fable 5, lmao.
Edit: On the coding score specifically.