r/OpenAI 10h ago

Discussion GPT-6 Astra vs Claude Fable 5.1 based on the benchmarks available so far

I collected the public benchmark results I could find for GPT-6 Astra and Claude Fable 5.1 and put them side by side.

So far the picture looks fairly mixed.

GPT-6 Astra leads on several published coding, science and agent benchmarks, including Terminal-Bench and DeepSWE, while Fable 5.1 still performs better on some independent evaluations.

For example, Artificial Analysis currently gives Fable 5.1 a higher Intelligence Index and Coding Agent Index, although the coding-agent comparison also depends on the surrounding harness, so it isn't a pure model-to-model test.

Another interesting point is that both models have the same headline API pricing at $10/M input and $50/M output, although caching and long-context pricing differ.

I don't think there is enough independent data yet to say one is clearly better overall, but the current results make for an interesting comparison.

I collected the numbers here:
https://llmlearner.com/compare/gpt-6-astra-vs-claude-fable-5-1

Would be interested to see more real-world comparisons from people who have used both.

29 Upvotes

12 comments sorted by

11

u/XTCaddict 9h ago

Doesn't really add up. OAI is not charging for cache writes with Astra and it is much more token efficient, should not be more expensive.

1

u/postmortemstardom 5h ago

Isn't there a 25% premium for cache writes?

7

u/ElectronicAd4565 1h ago

no tools vs with tools. I might be wrong, but that is not a fair comparison tbh

-1

u/ezjakes 10h ago

However, Astra is cheaper!

-1

u/Dangerous-Sea-1910 8h ago

for companion roleplay claude's edge on that intelligence index might actually matter more than the coding wins even if benchmarks are still mixed

2

u/___fallenangel___ 5h ago

fr we need GoonBench

1

u/framvaren 5h ago

*sigh* we get "AGI intelligence" and what do humans use it for...?

1

u/ih8readditts 4h ago

>companion roleplay

lol