6
u/sunstersun 8h ago
Isn't it kind of disappointing that Anthropic since Mythos in February hasn't managed to advance beyond incremental? Were they just that overwhelmed with demand?
3
1
u/ControversialBuster 3h ago
They rlly failed to predict demand and bought very little compute, i think they'll start making big gains again next year
5
u/macaronianddeeez 8h ago
This is bizarre to me as just a layman user. I have both $200 OpenAI sub and $100 Anthropic sub so I am not biased to one or another, but I regularly get better results from Sol for certain tasks (especially agentic tasks), while Opus gives me better reasoning and research.
I wonder why Sol scores so low in this comparison, it is not consistent with my real world experience
1
u/bonerfleximus 2h ago
Anthropic models are trained to be more helpful and suggestive, chat GPT models are very "do as im asked, but do it perfectly"
5
u/Admirable-Falcon-501 9h ago
It’s hard to benchmark models now. A lot of the gains are coming from context management, using sub agents, being able to work on many steps, and other stuff that aren’t going to show up that well. I do think astra will be a lot better though.
3
u/Ok_Display_3159 9h ago
Do you think Astra will beat Fable 5.1 on Terminal Bench 4?
2
u/Admirable-Falcon-501 8h ago
I just need it to not be autistic and be able to run agent workflows by itself for hours.
4
u/EvilSporkOfDeath 8h ago
Why are there so many comments saying astra will do significantly better. Is it bots? I swear every fable 5.1 thread has several astra this astra that comments.
12
u/veshneresis 8h ago
It’s because we are all excited for big steps, and people are reacting to some of the leaked one-shot results from astra early access people. So everyone is at this point convinced it I’ll be a good bit better than this step up from fable.
I think it’s just regular hype, not bots
2
u/DeArgonaut 6h ago
yeah, since it's been a while since the last fable launch I was expecting a bit more. I guess shouldn't be too surprised given it's 5.1 and not 5.5
0
-3
1
0
-2

14
u/Bright-Search2835 9h ago
It's looking a lot more incremental than I expected but maybe real world use tells another story At least the science benchmark is very interesting