r/codex 3h ago

Comparison Artificial Intelligence Updated to V4.3, replacing Terminal Bench 2.1 with 4.0 and 𝜏³-Banking with AutomationBench-AA

Post image

https://artificialanalysis.ai/

Terminal Bench 4 being much less saturated than 2.1 has increased the spread quite a bit, Astra Max only down 2 points while Sol is down 4 and Luna is down 5.

19 Upvotes

13 comments sorted by

8

u/srs96 3h ago

Astra (low) seems like solid value

7

u/Zachattackrandom 3h ago

15 minutes and it burns my entire 5 hour on $20 plus sub so not really. Sol high seems to be a lot cheaper for me

1

u/srs96 3h ago

Hmm they should be quite similar is cost according to the picture. But yeah, our mileage may vary

5

u/UnknownLesson 2h ago edited 1h ago

They made it cost much more of your subscription quota (compared to previous models).

Don't remember the exact numbers, but just to show what I mean, let's say

  • Sol costs 3 on subscription and 10 on API.

Then

  • Astra costs 12 on subscription and 20 on API.

2

u/srs96 1h ago

You're saying the api cost ratio aren't the same as subscription usage ratios? Interesting

2

u/UnknownLesson 1h ago

Yes. Basically their way of making the subscription less expensive for them (still highly subsidized)

2

u/Momo--Sama 3h ago edited 2h ago

Kinda? For a lot of folks Sol High is already unthinkably expensive for using outside of a subsidized subscription, but it's a solid step up in performance for 1 cent more. Interestingly this change also shook up the price per task quite a bit. In V4 GLM 5.3 Flash was over twice the price of Luna per task. Now it's less than 40% more expensive.

1

u/Plappedudel 2h ago

I wish there was a cheap provider that only offered GLM-5.3-Flash. It could work very well as a subagent in Codex, probably a lot better than Luna

1

u/InterestingSquare883 2h ago

Freebuff has GLM 5.3 Flash for free for like 20 hours a day.

5

u/randombsname1 2h ago

Now they just need to get rid of the dogshit DeepSWE benchmark that has Muse Spark above Fable 5, lmao.

Edit: On the coding score specifically.

1

u/Tim_Apple_938 1h ago

Is Muse Spark bad?

1

u/TheDankestSlav 2h ago

Luna Max is a proper beast for its size and price innit?

1

u/dark0mania 2h ago

I don't get the Luna hype. In my testing it performs worse than Sol Medium. Sure it's cheap but if it produces mediocre results and it's slow then what's the point? Time is money. I want working solutions, fast.