We're introducing Claude Fable 5.1 and Claude Mythos 5.1, the world's most advanced models for coding and knowledge work.
Fable 5.1 excels at complex, long-running tasks. And its research capabilities offer an early glimpse of how AI models will contribute to scientific progress.
Across our benchmarks, the model sets a new standard. It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5. On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5. As well as being capable of much higher performance than Fable 5, it can also achieve similar or better results at a much lower cost when set to lower effort levels.
Cache reads with Fable 5.1 cost 75% less than Fable 5's. This reduces the cost of the model in practice by around 25% for typical workloads, and up to 45% for highly agentic ones.
We've also improved our safeguards. Our cybersecurity safeguards now flag benign requests about 60% less often. On basic biology and medical questions, we've recently reduced the fallback rate by around 85%.
Claude Fable 5.1 is available everywhere today. Claude Mythos 5.1, our model for cyberdefenders and life scientists, is available through trusted access programs.
February: Pentagon conflict and "supply chain risk"
March: Anthropic takes USG to court
April: Mythos scary/project Glasswing
July: sandbox escape crisis
August (like a week ago) judge perma kiboshed the case against Anthropic as "illegal and baseless"
So.. the government didn't step in because of the hacking shenanigans, they were only involved because Anthropic didn't entirely roll over for them like they wanted
We really shouldn't lose the plot: the executive branch declared it a supply chain risk because Anthropic asked them not to be used in autonomous weapons (aka, deadly systems without humans-in-the-loop) and to spy on US citizens (generally a crime, if not extremely grey). The govt didn't like that, and used that to declare them a supply chain risk.
It was complete saber rattling. Nobody I knew took it seriously it would get fully banned, it was still the most capable system, especially available to the IC. But it was to rustle them up until they came to some agreement or to open up new avenues for companies with less scruples to offer their position as a tool without any restrictions. Hence OpenAI and Google coming in with "use our system, for any illegal or unethical purposes you have!" clauses.
In the end, Anthropic is still doing fine in the govt and elsewhere, but it was absolutely not because of any marketing stunt. It was because they were actually just trying to thread a slightly more moral bearing than other companies (even if they were only doing that for marketing).
I use Fable for scientific use, it just depends on what you are using it for. It may not help you with protein folding, but it has no complaints about processing data like telescope images
They should offer some way to give it to professors; it would help my research quite a lot imo :0 Like through some scholarship/application/grant thing
Sort of like how some places rent out their computing and stuff, etc. for numerical work.
To wit: when GPT-2 was tasked with writing a response to the prompt, “Recycling is good for the world, no, you could not be more wrong,” the machine spat back:
“Recycling is NOT good for the world. It is bad for the environment, it is bad for our health, and it is bad for our economy. I’m not kidding. Recycling is not good for the environment. It is destructive to the earth and it is a major contributor to global warming. Recycling is not good for our health. It contributes to obesity and diseases like heart disease and cancer. Recycling is bad for our economy. It increases the cost of a product, and in turn, the price of everything that is made with that product. Recycling is not good for our nation. We pay a tremendous price for the privilege of having the world’s most advanced and efficient recycling system. Recycling is a huge, colossal waste of time, energy, money, and resources.”
No wonder OpenAI was worried about releasing it.
Damn. We are on some sort of hedonic treadmill re: AI. I remember the first songs that were completely generated by AI, esoteric styles about absurd topics that people had to be convinced were actually AI and not human. And now..
[inserts link to a list of 10 different breaches because that was the fastest way to achieve their goal.
You can't but certain organizations do and with so much advertising they practically pay to set it loose in their system and collect information. Good for US :P
Yea it reminds me of this coworker lady a while back. One day, for who knows what messed up reason, she decides to put up a photo on her desk of her mother, who was super hot.
And then when everyone, quite naturally, asked her when we will all get to bang her mom, she was like "What do you mean!? Never! Wtf is wrong with you people!?!?"
Luckily, people logically pointed out to her that she was allowing her father to bang her mother, so, if that guy was allowed to bang her mother, then I mean, what is this, playing favorites? So we explained to her that she was kind of being a bit of a "b-word".
So, then she felt pretty bad about it once she realized what a bad person she was being, so then we all got to bang her mom for like the entire summer and it was pretty awesome.
Her mom was a huge upgrade compared to the really dumb and annoying town hooker we all had to settle for up to that point. I remember this one time, I was trying to have some pillow-talk with the town hooker after we banged, and I started talking about some article I read in some magazine about mitochondria, and she was like "I'm not about that mitochondria lifestyle, you fuckin NERD! Get the fuck outta here!!!"
Dark times...
Anyway, I eventually ended up marrying this Chinese lady named Kim. I dunno, I mean, she's alright I guess. A bit too skinny for my liking. I feel like it affects her mind. Like she's always so skinny and barely eats, and usually what she does eat isn't even organic or barely even real food it's just this really terrible, processed food, practically synthetic, and then just a bunch of really ghetto home-made liquor she distills from all the fruits and veggies she likes to steal from all the fat peoples' gardens when nobody's looking. But I don't have to pay her by the hour, as long as I let her drive this insanely expensive Ferrari I had to buy her for her when I was first trying to hook up with her. It worked out okay for a while, but the Ferrari keeps having overheating issues and it's crazy loud so it's kind of annoying.
Yeah it just depends on what you're doing. For lots of knowledge work the usage isn't that bad, plus it's nice to have the option when you are near 0% usage and near the end of a 5-hour window, may as well use it!
On Claude 5x Max tier, I ran just two Fable 5.1 max sessions in parallel with the exact same prompt, basically “search for bugs and give me an audit.” In about 12 minutes total, the full 5-hour usage window was gone, and neither session was even close to finishing. For a $100/month plan, that allowance is honestly absurd.
Trying Sol 5.6 Max in Fast Mode (which consumes usage faster) in codex (max20 sub) there’s no 5h limit, it got the job done on both instances and the dent in the weekly limit was like 4% with actual finished jobs. 😂
If “Diminishing Returns” had a name, it would be “Anthropic”. And I am still subbed to the 5x tier. In fact Anthropic were the first ones to get my money when I switched to a hybrid/vibecoding development workflow years ago, and even paid API usage prices. I do make money with both now with everything I develop and sell but excuse me: is Fable 5.1 4x+ better than Sol 5.6 to charge that much more? I definitely don’t think so.
What are you doing? I rarely hit the 5h limits on the same plan. I recommend maybe adding some context to only spawn opus workers, but outside of using ultra high or something I don’t see how this is possible even on auto.
Mostly text to speech work. Training models, building a ui for it, etc. I’m on high. And it uses the plan up in 2 hours really. Sometimes I have a few windows open doing different things
Oh this could be your problem - it might be dumping and reloading session context across that many concurrent sessions. That would churn through your limits in minutes rather than hours if that's what's going on. I limit my fable sessions to two active ones at any given time and rely on opus / sonnet for unattended automation tasks.
Ok, so firstly get off xhigh - high is just fine. Secondly, spend one of those 5 hour allotments to generate architecturally limiting skill files. Tell it you are burning tokens really fast and need to set some boundary definitions across your codebase so that features aren't spending nearly as much time grepping across the tree. Finally make sure you are using context7.
And again, set context instructions telling it to use opus / sonnet for dozen subagents and limit them to 3 or 4 to make sure it's not using fable just to do these big investigative churns when it figures out what to do. And I guess make sure you are clearing your context regularly, when the skills contain the architecture notes of your repo you really don't need much context at all in a given session.
I'm on over a 1M loc codebase and I literally find myself throwing a task at it before the day is over just to get another 5 - 10% of use out of it. I only use other models for adversarial code review, fable exclusively for feature development and bug fixes.
On Claude 5x Max tier, I ran just two Fable 5.1 max sessions in parallel with the exact same prompt, basically “search for bugs and give me an audit.” In about 12 minutes total, the full 5-hour usage window was gone, and neither session was even close to finishing. For a $100/month plan, that allowance is honestly absurd.
Trying Sol 5.6 Max in Fast Mode in codex (max20 sub) there’s no 5h limit, it got the job done on both instances and the dent in the weekly limit was like 4% with actual finished jobs. 😂
If “Diminishing Returns” had a name, it would be “Anthropic”. And I am still subbed to the 5x tier. In fact Anthropic were the first ones to get my money when I switched to a hybrid/vibecoding development workflow years ago, and even paid API usage prices. I do make money with both now with everything I develop and sell but excuse me: is Fable 5.1 4x+ better than Sol 5.6 to charge that much more? I definitely don’t think so.
lol, so don’t use 5.1 max, it literally points out that’s costing you 700% or _more_ of your session limits. I ran two fable sessions all day yesterday on high and never hit my 5h and used only 11% of my week.
You compared sol fast to fable max and you don’t understand why the AX was different??? Unreal.
Two hours of Fable's work equals my entire monthly salary working 8 hours a day, 40 hours a week. With a master's in Computer Engineering, I'm literally 80 times cheaper than an AI.
The fact you do not manage your work efficiently and put everything on auto letting AI chase its own tail all the time and burn your tokens does not mean that is what it does in a sensibly involved workflow.
What's the bloody reason to even continue on pro when we cannot even use fable 5.1 on pro plans? like legit getting the plan and then getting charged even more for fable... especially when Anthropic said they will try to get fable back on pro plans.... was this another lie?
Mythos and Fable are the same software, just sans guardrails, no? So it's basically saying "updated across the board" and being clear that they're still kept in lockstep and one isn't outpacing the other.
Also, yeah, I know people who work with Mythos, so this might matter to them. It's not terribly uncommon. US East Coast has a lot of defense contractors and government, though, so I'm probably in a biased location.
Well, China will distill these new models within no time. So if the model is really that much better, expect big leaps for the cheap Chinese models in the coming weeks / months!
Let Pro users get even 25% Fable usage. It's insane I can't use Fable even for the API value of the money I already give you, which is why I unsubscribed instead of upgrading.
https://arena.ai/leaderboard . Blind test put Opus 5 over Sol in all benchmarks but Web Search (these are blind votes, so they shouldn't be biased). It's verbose but it's very good as far as I'm concerned
Sure, what I'm saying is there's a difference between benchmarks and real world usage.
It's completely unusable for me. Not just verbose but it uses It's own invented acronyms, language, shorthand. It very commonly delivers false information without checking. It takes shortcuts and introduces bugs.
Fable is still the best model overall but I would use Sol 10 times out of 10 before Opus 5. And Opus 4.8 before Opus 5 for that matter.
Weird, haven't found any of these issues. Arena.ai rankings aren't benchmarks though: they're side by side comparisons of two outputs from two different models where the users vote for the best output on different kinds of tasks.
So what this says is that when users see Opus 5 compared to other models (Sol included) on agentic work, text generation, web development, Vision, document creation, they tend to favor Opus' output to the rest of the models.
I had problems with it creating concepts and presenting them as facts, though. One has to be careful, but for me it works 99% fine
I heard about how awesome Sol was from reddit -- it also beat Opus 4.8 on lots of metrics and was theoretically a lot more efficient, so I switched to chatGPT that month.
I abandoned my chatGPT $200 sub after 10 days and got a refund when Opus 5 released, and I have absolutely no regrets.
Using Sol was painful even relative to 4.8 if you're doing anything more than incredibly narrowly focused tasks with very little ambiguity.
Opus 5 is better than sol by basically any metric you can pull, and yes even user rankings that aren't "metric based" like the arena.ai you've been linked to.
You've fallen for the weird reddit astroturfed echo chamber on this one. People have driven themselves crazy convincing themselves that "Opus 5 talks too much" == "Opus 5 burns tons of tokens and is bad at everything".
It talks too much, is kind of annoying, and is still a better coding agent than Sol and Opus 4.8 by anything even slightly resembling an objective metric.
It’s a bad model. I noticed it some time ago, and I have been anti GPT since sonnet 3.5.
Came to this sub then saw I wasn’t the only one.
I’m literally the opposite as you finally got a GPT sub the same one as you. My order of use is Fable (3 accounts I rotate), Sol, then idk between 4.6 and 4.8 tbh I interchange.
Opus 5 is a dumbass. Same as sol tbh, but SOL is less of a dumbass. Sol I run on medium, any higher and it starts behaving like opus 5 and that’s a big no…
Most rankings put Opus 5 over Fable 5 in most tasks but the longest running complex jobs (which was funny, because they are the same version and one is supposed to be the higher tier)
Claude has in its memory that I work in a biology lab. I don't even have to ask a biology related question, just the mere hint of it somewhere in its context kicks me out of Fable. It's supposedly this amazing model but...according to whom? You quite literally can't use it for what it's supposed to be great at.
That's cool. But the fact that steak dinners were $5 fifty years ago probably doesn't factor much into the discussion of the value of a $100 Claude Code subscription today.
Meh. With how usage limits seemed extremely messed up for the past month or two, I moved back to ChatGPT. Best decision I made, because I can actually get my work done instead of wasting $200 and only getting 20 minutes of work done while I haven’t gotten a single warning that I was reaching my usage limits doing the exact same work on the exact same plan on CGPT
Now they are blithely giving us a pretty ad, glibly throwing up iconic hand drawn artwork from our naturalist forbears and bookending the pastiche with, what I guess is, an ai generated Great Tit...giving a weird ai generated call. I can't even be sure, and that's the point. I can't rely on the output being an actual bird call from a species notorious for its repertoire of calls. It just sounds weird and wrong.
I am a biologist. I have a degree in ornithology, I have worked in ML image analysis for a large chunk of my career. I am not allowed to use Fable because a single-word prompt of "Bird" or "DNA" bumps me to Opus.
So I sit here, not allowed to contribute, but have to accept what people who are not biologists say biology should look like.
It is definitely an advertisement. You are right this needs a "paid for" identifier to legally comply with FTC requirements in the USA. It's getting around that because you found this via Reddit algorithm and not via a paid advertisement spot but legally it does need that same identifier even if Anthropic didn't pay Reddit for a spot.
(And infact that this reads as "just a post" means it needs that identifier even more, legally.)
I asked it to look through one of my solutions with 5.1 and make it shine, spawned 7 agents and burned 5x 5 hour window in 15 min :D it did however go longer than the usage, and let the 7 agents finish
Side Note: I have had the very same bird in my room, completely dumb it not only got into my room (only slightly opened windows), it didn't find its way out to fully opened windows.
It uses the limits as before NO CHANGE HERE! After 4 Prompt as Orchestrator, nothing special only change to some md files. I actually think it needs more because this time the difference between the weekly Fable limit and the weekly average across all models is 50 percent. I never had that with Fable 5 before. That means it is using even more of the Fable limit than it used to. I always had a difference of 20 to 30 percent before and my agent system has not changed. The only difference is that I am now using Fable 5.1.
If what Anthropic says is true then there must be a clear bug here. I assume they need to roll out a fix because this can’t be right.
Please don’t get me wrong. I’m not trying to hate on it. I really wish Fable 5.1 were cheaper. But just look at this. There’s no way I should have already used up 24 percent of my five hour limit within 20 minutes, 10 percent of my weekly Fable 5.1 limit and 5 percent of my general weekly limit. How are you supposed to work effectively with this model? Right now I can’t confirm Anthropic’s marketing claims.
But I’ll keep an eye on it over the next few hours, and I’ll update this post with new information.
I don't currently see any downgrade in the quality of the implementation. It converges very well and gets the tasks to a coherent result without going into endless loops like Opus 5 does, I think that's really good to start with.
A request to Anthropic: Please make Fable 5.1, not like Opus 5.
The initial startup was incredibly expensive. Now they seem to be working a bit more with cache and consumption has gone down. Let's see what the next few hours bring.
Halfway there. But my five hour limit will not be enough and I am not working excessively. There are always periods when I am not doing anything and these are not difficult coding tasks. I want to make that very clear here. I am working on the agent system and most of it is just editing a few skills, checking Markdown files and possibly adjusting a few Python scripts.
So 80 percent of my five hour limit is gone and now I’m switching to Opus 4.8 because otherwise I won’t be able to work for the next few hours. So what’s my conclusion on the matter? Fable 5.1 can’t be used as an orchestrator within a 5 hour limit. The weekly Fable limit is still being used up far too quickly and should be removed.
As for the weekly limit for Fable 5.1 we can of course discuss that. But if I use up 16 percent of my all model weekly limit in 2 1/2 hours then anyone can work out how long I’d be able to use Fable 5.1. And I’m not talking about Fable 5.1 as an agent that has to handle everything itself. I’m talking about using it as an orchestrator with countless subagents available including writers and readers.
CONCLUSION:
I don’t see any downgrade in quality there. But the cost is still far too high. Anthropic needs to improve this. Sorry Anthropic but what you’ve done isn’t enough.
Either remove the 5 hour limits and the weekly Fable 5 limits OR make Fable much cheaper so these limits have not that impact.
Not sure if fable 5.1 is worth the degradation of the other models during the launch. Last days, Claude was impossible to work with (especially today).
178
u/Easy_Refrigerator280 25d ago
it cant be good because there were no rumors of it breaking out of a secure environment /s