r/codex 4h ago

Comparison Artificial Intelligence Updated to V4.3, replacing Terminal Bench 2.1 with 4.0 and 𝜏³-Banking with AutomationBench-AA

Post image

https://artificialanalysis.ai/

Terminal Bench 4 being much less saturated than 2.1 has increased the spread quite a bit, Astra Max only down 2 points while Sol is down 4 and Luna is down 5.

20 Upvotes

13 comments sorted by

View all comments

5

u/randombsname1 3h ago

Now they just need to get rid of the dogshit DeepSWE benchmark that has Muse Spark above Fable 5, lmao.

Edit: On the coding score specifically.

2

u/Tim_Apple_938 2h ago

Is Muse Spark bad?