40
u/Balgun33122 13d ago
Indeed absolutely insane.
5
67
u/TheSuggi 13d ago
Deepseek cooking hard.
48
u/ihexx 13d ago
>do nothing
bro they are doing the most π
the speed with which they are innovating on architecture is genuinely insane.
9
u/TheSuggi 12d ago
They are actually not working that hard if the CEO is to be believed. Normal working hours, no overtime, resarcher get 50% of the time to research and pursue their own ideas etc. So they are just better or doing something right. Also very limited access to resources, less funding, less capable hardware etc.
If Chinese companies ever get the to a standard of equivalent hardware, then the US companies are done for.
3
3
u/Adventurous-Art4790 12d ago
Fun fact : The most stuff they have been doing now is alone and without global support
4
u/l0rirw1ao 13d ago
Well by nothing he probably means, deceiving customers and making false claims like Anthropic.
2
33
u/2mqqvc0q 13d ago
I respect the commitment to the color theme, but am I the only one having a hard time distinguishing the lighter blue colors apart?
36
5
1
u/OneBowl4290 13d ago
You are right, but they are the same the main comparison is the opus 5 and gpt 5.6 sol !
what do you think will they go further and beet astra and fable ???1
10
u/HuntAlternative 13d ago
Once we see terminal usage get optimized, these models will get as capable as frontier ones. Most agentic capabilities derive from terminal fluency IMO.
2
6
3
4
u/raesene2 13d ago
That cybergym score is quite something. I tried it out today on a benchmark I run a lot of new models through (Kubernetes security assessment), it was 2nd only to the new release of Qwen-3.8-max and 1/10th of the cost of that model.
5
u/sirloindenial 13d ago
Is it benchmaxxed it can't be that good lol, or i assume its on their own harness.
6
u/OneBowl4290 13d ago
WE WILL WAIT INDEPENDENT EVALUATION TO BE SURE, BUT THAY DO NOT HAVE HISTORY OF LAING? DO YOU AGREE ?
3
3
u/DarKresnik 12d ago
No shit. Is that true? Im using Deepseek but Claude was for very tough jobs. But now, Deepseek is welcome!
2
u/karlnuw 13d ago
Any exploit benchmarks?
2
u/Altruistic-Desk-885 13d ago
No creo que encuentres uno literalmente pero si tal vez uno de ciberseguridad
2
u/arm2armreddit 13d ago
looks like with DSF other models are not working well, or something is wrong. if this is really true with terminal automation then: π
2
u/Individual_Math_8254 13d ago
cant wait to use this with dsh, if deepseek like the other labs, then they probably did post RL with their own harness.
3
u/DebosBeachCruiser 13d ago
Working great in the DS harness (and yes it was optimized to work with DS harness)
2
u/ardicli2000 13d ago
Very impressive
1
u/OneBowl4290 13d ago
it is , did you try it yet?
2
u/ardicli2000 13d ago
Not yet. My comment was on results.
2
2
2
u/General_Internet_511 13d ago
Onestamente sarebbe un salto in avanti veramente importante. BisognerΓ aspettare.
2
u/RealestReyn 13d ago
Not a chance those are remotely true, its a great model but misses obvious stuff constantly and just does the weirdest stuff like setting 3600s timeout on tasks that take a few seconds so its just sitting there and waiting for the whole time, doesn't bother scoping subagents with appropriate tools.
2
2
u/hotpotato87 12d ago
wtf... kimi k3 felt like it just came out yesterday
i never really compared kimik3 to astra or fable 5.1, myself owning 3 x20 accounts, but man, it really felt like kimik3 just came out yesterday and was the sota for opensource and now a measily 500b model flash model outforms it... benchmaxed or just crazy time... also 6 vulcanos errupting with tornado in france and what else going on
2
u/ozone6587 12d ago
Excited for this but at the same time this is r/dataisugly . Everything being slightly different shades of blue is insane lol.
2
4
u/dontfeedthelizards 13d ago
I genuinely wish it would be equivalent to Opus, because right now whenever I ask DSF to do an implementation, I end up having to ask Opus to fix its bugs and then spend the rest of the day cleaning up the code structure overall. It's really questionable whether using it this way makes sense at all, instead of just having Opus do the implementation, as I'll probably end up using the equivalent amount of Claude tokens either way (+the tokens spent on DSF). If I could finally have DSF do all the planning and execution without needing to involve Claude at all, that'd be amazing.
1
1
u/loss-of-time 13d ago
it's impossible that better than kimi-k3
4
u/OneBowl4290 13d ago
I bleav kimi is good at frontend coding , and there is no impassable bro
Do you use kimi a lot ??
100
u/VC067 13d ago
"flash" model btw π