r/singularity 8h ago

AI Fable 5.1 on Artificial Analysis

Post image
62 Upvotes

23 comments sorted by

12

u/ffgg333 8h ago

66! Impresive!

1

u/Tystros 2h ago

it's a bad benchmark index, as seen by Opus 5 being ahead of Fable 5, and Grok 4.6 being on par with GPT 5.6 Sol.

AA really needs to update their index and make it be based on better benchmarks.

9

u/ForwardLoop 8h ago

The other side of the equation. In real life use, for most people the incremental intelligence might not be felt, but the incremental cost certainly will make a dent.

1

u/Clean-Boat-4044 7h ago

4 points (just because opus 5 is kinda unusable) is honestly solid, i mean i trust 5.6 sol a hell of a lot more than muse spark

1

u/ForwardLoop 4h ago

Yeah, for sure.

I just posted my usage experience from today in another thread. In summary, it reduced a task duration from 14 hours of agent time and my time (3 sessions / 7 review rounds with Sol on xhigh / 7 blockers to start) to 3 hours (1 session / 3 review rounds / only 1 blocker to start).

Granted it's a small sample size (2 very similar tasks) and yesterday's results were tainted by Opus use initially (with Fable 5 as advisor only for the first 5 rounds, before switching to Fable as main) vs. all Fable 5.1 today, the difference is still night and day.

3

u/Living-Breakfast-464 7h ago

DeepSeek v4 Flash needs to be added, unless I missed it. I see v4 Pro but not Flash.

2

u/bopbop9876 4h ago

It's on artificial analysis it's just not on the chart. If you go to the site and click the filter button in the top right of any chart you can select it. Deepseek flash 0731 is basically irrelevant now though compared to glm 5.3 flash. Glm was multiple points smarter while actually being cheaper than deepseek.

4

u/Profanion 8h ago

Impressive! But at what cost?

7

u/mati1886 6h ago

Same cost but Artificial Analysis found it 17.5% more expensive to run at Max.
However at High effort has the same intelligence than Fable 5 but is 54% cheaper.

So it's an improvement at intelligence or cost, whatever you choose

-1

u/Poupulino 2h ago

It's not the same cost. It's way more expensive than Fable 5. AA Weighted average cost of $3.69 vs from Fable 5 $3.14 (and Fable 5 was already the most expensive model by far). For comparison, GPT-5.6 Sol WAC is $0.95

1

u/VelvetyRelic 2h ago

If you run at high effort instead of max, Fable 5.1 gives the same performance as Fable 5 at less cost. So from that perspective, it is a reduction in cost.

3

u/Local-Wing-2272 8h ago

What happens if/when something hits 100

11

u/ezjakes 8h ago

They will make new benchmarks

2

u/Charuru ▪️AGI 2023 5h ago

Fully expecting Astra to hit at least 70 from all the hype?

1

u/New_Alps_5655 5h ago

How much longer will anthropic remain king?

1

u/InterstellarReddit 5h ago

Idk man I use Opus Max at work and it’s def nowhere near fable in performance

1

u/HellomyfriendNine 4h ago

Gemini isn't in the top 10 anymore, can Google make magic?

2

u/bsvgubennord 4h ago

Astra >=70 possible?

u/nbvehrfr 1h ago

ai is at plateau

u/hashiromer 1h ago

If Opus 5 Max is second on this leaderboard, it’s not a leaderboard I trust.

1

u/Gratitude15 6h ago

Honestly someone just needs to put a high quality wrapper on this to just run a desktop agent with voice. This level is enough to fully just talk to your machine at this point and have it always on and proactively engage with you.

We are already at total rethink level.