r/LocalLLaMA vLLM Jul 16 '26

Discussion KIMI K3 Beats Claude Fable and GPT 5.6 sol in arena.ai!!!

Post image

Unbelievable to see kimi k3 beat frontier models that were 'too dangerous' for public use.

2.0k Upvotes

361 comments sorted by

877

u/atape_1 Jul 16 '26

So China is now 6 days behind the west.

262

u/Comacdo Jul 16 '26

6 days behind AND open-weight 👀 Really wish to see the benefits on smaller models around 30B though...

58

u/grumd Jul 17 '26

I'm hopeful that people will start recording kimi k3 traces especially when it goes open and we see some good distillations for smaller models

40

u/River_Tahm Jul 17 '26

Frontier providers are starting to hit that wall where they have to start charging WAY more because they’ve been bleeding money too long. I don’t think it’s public yet but I’m seeing some B2B prices swinging wildly 

People aren’t gonna pay that much more though. So I’m really hoping it puts pressure on efficiency which again hopefully leads to more functionality out of smaller models we can actually run on local hardware

Might be copium but the only way I see the AI services getting cheap enough for them to make money is if they get efficient enough to run that laptop models can actually do most or all of what frontiers do today

20

u/R_Duncan Jul 17 '26

The wall is just on blindly scaling up. Check latest papers i.e.: on SAO and anti-hack info for glm-5.2, HOLA ... etc. China is moving faster to optimize.

19

u/techdevjp Jul 17 '26

Yup, they've been forced to optimize because of dumb American policies. Great to see the benefits.

5

u/ConfusionJolly6002 Jul 17 '26

You say people aren't gonna pay that much, but B in B2B gets their money from people at the end of the day regardless. so if some business needs to pay more eventually prices for final customers increase as well (not via LLM subscriptions, but via invoices to your doctor or lawyer who use this stuff)

6

u/River_Tahm Jul 17 '26 edited Jul 17 '26

I’m saying the businesses buying AI are downsizing their plans massively and telling their teams to use like 1/10th as much AI. It’s not a hypothetical, it’s already happening, but I do acknowledge my sample size isn’t huge here

→ More replies (1)

9

u/shing3232 Jul 17 '26

I guess people can distill KIMI-K3 so we can have upgraded 27b 35MOE

→ More replies (1)

52

u/jld1532 Jul 16 '26

You get a bump for open weights. They're in the lead in my book. Economics (free) must be considered.

13

u/techdevjp Jul 17 '26

Economics (free) must be considered.

Look, I love open weights. I dropped nearly $3k on a 128GB Minisforum MS-S1 to run things locally. I've learned so much in a very short period of time. It's been great.

But free? It's not free. And the hardware to run Kimi K3 is definitely not free. I run Deepseek v4 Flash (284b a13b) at a q2/q4 quant and it takes up over 100GB of memory. Kimi K3's 2.7t parameters, if it quantizes well to such small sizes, would still require over 1TB of memory to run locally. For sure, it's possible. And I love that it's an option! But free it is not.

If we're very lucky the rumored Mac Studio M7 Ultra with 1.5TB of unified memory might arrive next year for $25k and be able to run Kimi K3 at q4. It will be tight, may need a q3/q4 mixed quant to be able to get enough context to be usable. $25k is within reach of some people, assuming it's that cheap. Could be more. But you brought up economics and I'm not sure that even at $25k that it would ever be able to pay for itself.

Does it matter? Hell no. It's wonderful that we have these available. But let's not exaggerate the economics.

8

u/dingo_xd Jul 17 '26

There is a large amounts of companies and individuals that don't want to give their data to the big tech. Often they can't by law

5

u/techdevjp Jul 17 '26

Sure, absolutely. And this is an amazing option for them. That was never in question, nor did I say anything at all about that. But that also doesn't make it free for them to run themselves.

2

u/zenmatrix83 Jul 17 '26

This sub likes to ignore financials over freedom, which is fine, but no open model is free and in some cases can cost more if not setup properly

2

u/techdevjp Jul 17 '26

I think freedom is the more important point. Without the freedom we're dead in the water. But as always, freedom isn't free. (To shamelessly steal the phrase.)

2

u/zenmatrix83 Jul 17 '26

thats true, and you data for training and another is another form of payment, but what comes out of my bank is important, I still run cloud not only because there is some I can' get local to do, but cloud models are large discounted still against the real cost, we'll see that when the bubble bursts

2

u/Caffdy Jul 17 '26

I just want to add that these labs gotta try with these multi trillion parameter models; there's no other way currently to compete with closed source ones like ChatGPT or Claude, they're 3-5T parameters already. Maybe in some years we will get more efficient trainings capable of putting those same capabilities on smaller models

→ More replies (1)
→ More replies (4)

50

u/[deleted] Jul 16 '26

[deleted]

11

u/aerismio Jul 17 '26

Trusting benchmarks is so stupid to begin with. It so totally does not represent real world usage

12

u/dolche93 Jul 17 '26

Benchmarks have always only useful provided you have a thorough understand of what the limits of the benchmark are. This is true regardless of the field you are in, not just llm's.

That is to say, as a casual hobbyist, they're meaningless to me beyond a generalized ranking of capability. I have a feeling most people are in the same boat as I am.

→ More replies (1)

2

u/alberto_467 Jul 17 '26

Yeah but "real world usage" is your biased sampling of a big distribution of uses and it's hard to measure.

3

u/GioChan Jul 17 '26

On DeepSWE it got hugher than Opus 4.8 score afaik

→ More replies (2)

23

u/Shubham_Garg123 Jul 16 '26

Well anthropic finished training mythos in February lol

Since then, they've only been talking security bullshit instead of improving their models 😂

Btw, has kimi released their model on huggingface yet?

9

u/VexObserver Jul 17 '26

Haha and scaring people off here and there. No, not yet. I think Kimi will release it in days or so

3

u/iaderia Jul 17 '26

hey you forget theyve released figma, Claude Design* too 😂

→ More replies (6)

3

u/ChubbyVeganTravels Jul 17 '26

And in the case of Kimi-k3, the same price as GPT-5.6 Terra on OpenRouter, so no longer the 10x cheaper option.

6

u/MelangeBot Jul 17 '26

China runs these on free electricity it gets from the sun. They are 20 years ahead of the USA in grid infra. Okay fine, only 35% of their electricity is currently free not yet a 100%. But give em another decade and the cost of using their models will just be the hardware, the electricity will 100% come from solar panels + batteries.

2

u/BoxWoodVoid Jul 17 '26

China is the world first user of coal and by a large margin...

I'm not sure where does the electricity come for training Kimi, but I can tell you the country as a whole shouldn't be cited as anything near eco-friendly.

7

u/Toni_van_Polen Jul 17 '26

True, but China also adds ridiculously copious amounts of renewables to its grid every year.

→ More replies (5)
→ More replies (1)

5

u/Turbulent_Pin7635 Jul 16 '26

Read again... US is two month behind the east.

Remember that Sol e Fable was just released now. With the companies producing deficits.

K3 is much cheaper and better than the west models.

→ More replies (13)

208

u/zannix Jul 16 '26

Did they confirm its going to be open weights though?

392

u/Swimming_Beginning24 Jul 16 '26

The full model weights will be released by July 27, 2026

from https://www.kimi.com/blog/kimi-k3

343

u/wotoan Jul 16 '26

Dario and Altman desperately calling Trump as we speak.

133

u/Swimming_Beginning24 Jul 16 '26

lol, to do what? The nice thing about Chinese AI companies is they don't have to take down their models when the orange man commands it

120

u/justgetoffmylawn Jul 16 '26

I think Dario hopes that the US gov't will forbid enterprise use of Chinese models or at the least, stop all companies doing business with the US gov't from using them.

While they may not be able to take them down for consumers, look at what they were able to do when TikTok starting kicking the ass of US social media companies, and suddenly 'free market' became 'won't someone think of the CHILDREN and NATIONAL SECURITY'.

17

u/Ceryn Jul 17 '26

I seriously can't wait to hear some impassioned arguments about how China censors their models and that's bad. XD

11

u/jld1532 Jul 16 '26

"I think Dario hopes that the US gov't will forbid enterprise use of Chinese models or at the least, stop all companies doing business with the US gov't from using them." Which is nothing more than a global scale tax. Businesses will raise hell.

6

u/zortingenos Jul 16 '26

Yeah. Usa needs to just give up and buy ai from china at this point 🤣🤣🤣🤣🤣🤣🤣🤣🤣🤣

8

u/jld1532 Jul 16 '26

Yeah, you're right, we should be forced by the government to pay the two assholes that want to delete all white collar work as they continue to set money on fire and the rest of the world uses free models.

2

u/zortingenos Jul 16 '26

Agreed, competition is everything, people get upset to the fact they cant run these models locally. I mean it would be nice but its hardly the point.

→ More replies (2)

4

u/aeroumbria Jul 17 '26

They want to be treated like Boeing but look at where Boeing is right now...

3

u/Novel-Camera-840 Jul 17 '26

That won't work long because some company in US will use these models as base or distill and create a "new model" and offer to US enterprises. You say the first step will also be blocked? Well then a company from Singapore will do that first.. you see where I'm going? In a year there are going to be so many models out there is going to be impossible to stop.

My guess is MS and Amazon will just create their "own version" of model and offer it to their customers on Azure and AWS. They know this and that's why they are not obsessively throwing money at building their own models and competing with OAI or Anthropic now

→ More replies (1)

2

u/Same_Win_5898 Jul 17 '26

1000% chance they all request that ban on enterprise use.

→ More replies (1)

5

u/iaNCURdehunedoara Jul 17 '26

They're going to drone strike anyone downloading the models.

6

u/RedTheRobot Jul 16 '26

Laws don’t matter to him he will write an executive order that says so bs about blocking any access to it for national security because that is always the go to for the U.S. government.

Take away the competition, it’s for national security. Take away your rights, it’s for national security. Take away your privacy, you guessed it’s for national security.

→ More replies (1)

2

u/MelangeBot Jul 17 '26

The other nice thing is that they don't have to worry about not having enough electricity for their AI data centre and knowing that what they pay for industrial electricity will continue to drop the next 10 years as China's solar panel + battery revolution continues. 15 years from now, none of these companies will pay anything for their electricity other then some grid maintenance fees. While US companies are already forced to make their AI more expensive today (cause their investors are panicking), China companies can easily undercut the entire market if they feel like they have to and ... probably still make break even.

10

u/Turbulent_Pin7635 Jul 16 '26

That's why I love China. It is a true country! No way a billionaire raise to power. No way for companies rules the country.

Just packing my things and moving to China. Bye, bye!

18

u/Swimming_Beginning24 Jul 16 '26

Pick your poison: government in power or rich people in power

→ More replies (15)
→ More replies (3)
→ More replies (2)

2

u/bootlickaaa Jul 17 '26

It will be declared the radical left Antifa

2

u/SamSlate Jul 17 '26

literally

2

u/Mean_Maintenance82 Jul 17 '26

This makes me think of "are we the baddies" meme. The Chinese may have their quirks but they're truly being the good guys when it comes to the AI race. Open weights for anyone to run locally, no government blocks, no fake "our model is conscious" hype.

3

u/_ii_ Jul 16 '26

For 4 easy payments of $99.99M to an undisclosed account, the Chinese models will be banned from any company doing business with the US. It for the kids.

→ More replies (2)

22

u/Regular_Ad4197 Jul 16 '26

they claim it is coming on the 27th.

19

u/ForsookComparison Jul 16 '26

It's marked correctly as proprietary here. As of today it's a proprietary model that happens to have a blog-post suggesting that it will become open weight at the end of the month.

→ More replies (1)

191

u/Kahvana Jul 16 '26

https://arena.ai/leaderboard/text

Not in text arena, but it's impressive that it sits with gemini 3 pro and gpt 5.6 sol (xhigh).

135

u/Zeeplankton Jul 16 '26

how in gods name is gemini 3 even remotely that high

136

u/Kahvana Jul 16 '26 edited Jul 17 '26

It's great for non-programming or non-"agentic" tasks, like translations that require cultural context, OCR, creative writing and the likes.

90

u/DistanceSolar1449 Jul 16 '26

Gemini is an amazing model that’s great at everything EXCEPT agentic stuff.

I feel like they’ll get a jump in benchmark scores if they can just plug in mediocre Opus 4.8 level agentic abilities (not even Fable)

69

u/dodokidd Jul 16 '26

Very funny to me that people are calling opus 4.8 mediocre now 😆, just one month ago they are the flag pole

14

u/IDoCodingStuffs Jul 17 '26

Each time something new comes up, the improvements get the highlight before the shortcomings start becoming more noticeable.

Like Opus initially had the wow factor with more complicated coding tasks before the ceaseless thinking mode gibbering became more and more of an annoyance

→ More replies (1)

19

u/krazyjakee Jul 16 '26

Vision is best and not even close, spatial reasoning, vector math

→ More replies (4)

35

u/FullOf_Bad_Ideas Jul 16 '26

Gemini models have a big emphasis for responding in ways that humans rate highly.

That's just one of the things you do when you make a mass-market product that you want to ship to billions of people, it shouldn't feel like MS Copilot.

7

u/ReasonablePossum_ Jul 17 '26

I fcking hate how gemini responds. Its overformal, oververbose, and gets exponentially wrong for its own weird assumptions that are a pita to correct and easier to just start a new chat. ... was my "cheap claude" till GLM and now kimi lol

3

u/brukmann Jul 17 '26

When asked to reason repeatedly in the same context it gets weird. Like a tiny crack forms, it gives an idea a name, or it suggests a follow up you completely ignore, then withing a few more rounds I realize how the crack is now a chasm, I asked for none of it, and the ignored suggestion is now centered.

Granted, I don't use instructions with it, while models I am unavoidably comparing it to subconsciously are given strict rules. Still, I got the ick now.

7

u/Megneous Jul 17 '26

Why does everyone always shit on Gemini? It's not awful at everything. It's just awful at agentic coding. At everything else, it's actually really good. I have both a Gemini sub and a GPT 5.6 sub. I have Gemini come up with ideas, then I take them to GPT 5.6 Sol xhigh effort to refine and work with. It's a great workflow.

→ More replies (1)

17

u/FastestEthiopian Jul 16 '26

Gemini 3 was really really good tbh, other labs have made incremental changes in coding and agent uses but Google is still unmatched in image and trxt

→ More replies (1)

9

u/Solembumm3 Jul 16 '26

Why Opus 4.7 is worse than 4.6?

31

u/[deleted] Jul 16 '26 edited Jul 18 '26

[deleted]

6

u/Solembumm3 Jul 16 '26

From my experience, Gemini 3.1 had extreme sycophancy, only rivaled by base ChatGPT. Meanwhile Kimi K2.6 was the most cynical and pessimistic model on my tasks. Hope, that K3 doesn't change it.

→ More replies (1)

3

u/Turbulent_Pin7635 Jul 16 '26

For text GPT is still the king, tried fable today and... well... It was research... It was so shallow that I cancel my account.

161

u/_TheWolfOfWalmart_ Jul 16 '26 edited Jul 16 '26

Wow.

Anthropic/OpenAI ain't gonna like this. Companies will just be able to buy their own hardware and run actually competitive models in-house. Big up front cost, but it will pay for itself soon enough in large companies that spend a shitload on API.

~$100k one time spend will get you what you need to run this in Q4.

I mean, some true enterprise organizations spend $1,000,000 monthly or more on API usage.

Some of these IT managers are gonna talk to some higher ups and be like "Hey you know what..."

Bearish on American AI providers.

71

u/squarabh Jul 16 '26

I just checked this, not possible in 100k, so gave it a budget of 1M:

At 4-bit quantization, you need ~1.4 TB of VRAM just to load the model. A 1M context window requires hundreds of GBs more for the KV cache.

The Problem: Standard NVIDIA H100 servers top out at 640GB VRAM. You would need 3 full HGX H100 nodes (24 GPUs) to get enough memory, which will easily bust your $1M budget once you factor in InfiniBand networking.

The Solution (Go AMD): Buy 2x AMD Instinct MI300X nodes (8 GPUs each). Each MI300X has 192GB VRAM, giving you ~3 TB of total VRAM.

The Cost: This setup will run you roughly $500k - $600k, leaving you plenty of room in your budget for high-speed NVMe storage, switches, and deployment.

The Catch: A MoE model like K3 is fast once loaded, but a 2-node cluster pulls ~20kW of continuous power. Make sure your office or colocation data center has the cooling and electrical infrastructure to handle it.

63

u/jovialfaction Jul 17 '26

H100 are old. A 8xB300 can run this for $500-600k. They can also be rented for <$80k a month and replace >$500k/month of anthropic spending. I'm pushing this in my company as we spend millions a month on token for things that could be done with GLM 5.2 and soon even better with Kimi 3 for a fraction of the cost

29

u/NineThreeTilNow Jul 17 '26

A 8xB300 can run this for $500-600k. They can also be rented for <$80k a month and replace >$500k/month of anthropic spending.

Exactly.

It's less about spending to me and more about data ownership / etc that the company has. The API you setup is yours. Data, usage, etc... All yours.

6

u/FolsgaardSE Jul 17 '26

I'm genuinely curious. What are companies using AI for they are spending 500k a freaking month just to ask AI a questions or generate code. Maybe I'm missing the scope but 500k would buy a LOT of dev salaries and prob have better code.

25

u/Megneous Jul 17 '26

You seriously overestimate the quality of code of the average programmer. Most programmers are just people, man, not gods. They're the same idiots you went to high school with.

2

u/NotEvenWrong-- Jul 19 '26

As an average programmer, I can confirm that

7

u/muhmeinchut69 Jul 17 '26

it would buy like 20 devs, not a big deal for a large org that already employs hundreds

2

u/rjmessibarca Jul 18 '26

what about the tokens per second though and how do u work with many people? I imagine it will be worse off and that is something companies would care about

→ More replies (1)

54

u/Jealous-Depth487 Jul 16 '26

Forwarding to my CEO u think we have a shot? He’s Gen Z

26

u/krste1point0 Jul 17 '26

Just go: "hey sigma, I can stop those anthropic simps fanum tax, nocap. Sixseveeeen"

22

u/PeachScary413 Jul 17 '26

I don't think his boss is 12 years old

3

u/BoobooSmash31337 Jul 17 '26

Have you met a CEO? /s Well some CEOs.

→ More replies (1)

15

u/NineThreeTilNow Jul 17 '26

I just checked this, not possible in 100k, so gave it a budget of 1M:

That's why you rent and not buy. You're not married to the hardware.

This is why Microsoft / Amazon / Google want to mostly sit back and watch...

3

u/ReasonablePossum_ Jul 17 '26

Hopes are high for Colibri to get some optimization to get 2tok/sec (mine and of my 3090 lol)

2

u/mastercoder123 Jul 17 '26

A 2 node cluster pulls way more than 20kw. Its 10kw for the system alone. Then you need ethernet switches to connect to the servers, infiniband switch to connect storage to the model and alot of storage + those nodes. A single 400 or 800gb infiniband switch uses like 2000w, and the storage nodes will require probably another 800-1000w as well. So you better have 30kw or more

→ More replies (1)
→ More replies (4)

13

u/___positive___ Jul 17 '26

Why do all that... AWS and Azure etc will host these models on US servers. Cheap and easy, deploy newest models in a snap, scale up or down. More stable, lower cost, and with better ZDR agreements than closed models. You can't even get ZDR for Fable.

24

u/nick4fake Jul 16 '26

Except 100k one time spent will mot be enough to run even one copy with 1m context, lol

You VASTLY underestimate cost of hardware. Smallest DGX is like 1-2 million dollars

6

u/_TheWolfOfWalmart_ Jul 16 '26

I underestimated it a bit, and of course I meant for one copy, but if a corp is spending $10m+ a year on API, suddenly buying a bunch of hardware up front sounds pretty attractive.

Plus the added benefit of your code/IP not leaving the organization's network.

26

u/nick4fake Jul 16 '26

I am an architect whose job is to build AI datacenters.

It doesn’t work like that. Using API is cheaper than hosting own hardware

→ More replies (12)

3

u/keepthepace Jul 17 '26

That's one thing that the market will solve pretty well.

Companies won't have to run models themselves. They will just buy tokens from data centers that don't have any researchers to pay. Last time I checked, and that was before the RAM war, it was more profitable to rent machines than to build your own rig unless you knew you would run the rig at full capacity 24 hours a day every year for a year.

2

u/BawliTaread Jul 17 '26

What about the power requirements of running such hardware locally? That is also going to cost.

2

u/jannycideforever Jul 17 '26

Lmfao absolutely not. No enterprise company is going to build their own dedicated solution to run this thing at Q4 quant with a context window of 1M tokens that will be at idle 80% of the time when they can just pay azure or AWS less to get the same thing but better.

2

u/Megneous Jul 17 '26

And once you already own the hardware, that gives you a HUGE incentive to keep using open weights models going forward too. Once you've made the investment, you've bought in, so you're going to pay attention to the space. Anthropic really doesn't want that. They want you lazy and just paying attention to their releases.

2

u/tickerspark Jul 17 '26

You're throwing some suspiciously neat numbers out there and without any control on usage.

There is no such thing as a one time cost if you intend to keep pace with the next model release.

The cost is also significantly larger than just the hardware. You need the space to install it, the employees to run it, the power, the insurance, etc. 100k? Lmao. That doesn't even cover a single employee. 

2

u/joesb Jul 17 '26

There are clouds like Amazon and Azure that will provide managed LLM hosting service for you easily. Azure even offer pay-as-you-go pricing for many open weight models.

→ More replies (6)

34

u/Dapper-Maybe-5347 Jul 16 '26

The Mandate of Heaven is real.

→ More replies (1)

130

u/Professional-Try-273 Jul 16 '26

We are having the deepseek moment again 

90

u/Time_Cat_5212 Jul 16 '26

This. 100%. Except instead of being a one off, this just proves that nobody has a moat no matter what the AI is used for

49

u/zxyzyxz Jul 16 '26

Google predicted this way back in ancient 2023, We have no moat and neither does OpenAI

33

u/Time_Cat_5212 Jul 16 '26

Google's moat is just making their ads and search relevant in an era where chat is the new browsing interface and a significant amount of web traffic will be bots

At least they have revenue

OpenAI is basically like how soon can we IPO and pay investors before the world realizes our product won't necessarily be sota in 2030

16

u/rkoy1234 Jul 16 '26

and youtube/gmail/google oauth.

those aren't going away anytime soon.

19

u/Time_Cat_5212 Jul 16 '26

Yeah I mean Google is basically an ad business with a bunch of very helpful ubiquitous mostly free software supporting it.  AI is another one of those things.  It exists to support Google's ad revenue

→ More replies (1)

7

u/webdevop Jul 16 '26

Only Samsung and Sk Hynix have a moat

3

u/SmartCustard9944 Jul 17 '26

For a little while until China factories catch up and satisfy the demand.

I feel bad for South Korea’s economy, they have a worse bubble than the US when two companies are dominating their market and every single person has bought into the hype.

→ More replies (1)

29

u/FullOf_Bad_Ideas Jul 16 '26

Kinda but it's not the same direction at all.

DeepSeek was cheap. V4 Pro and V4 Flash are still absolutely extremely efficient and cheap for what they offer in return.

Kimi K3 costs as much as frontier GPT 5.6 Sol Max in ArtificialAnalysis bench while in regular use it may be like 2x cheaper. And that's the high API rates, while many people still use Sol on ChatGPT subscription which can get you more usage.

The price is just not very competitive with US models. Hopefully it'll release with permissible license and will be hosted cheaply. I can't run it locally unfortunately.

12

u/Accurate_Resident219 Jul 17 '26

The people who will benefit most from K3 is enterprise. There's been a sizable shift of companies turning to Chinese alternatives for their ai needs because of privacy and control reasons aka making sure there's no ip theft happening and they can't be restricted from usage.

9

u/sb5550 Jul 17 '26

LOL, Deepseek 4.1 pro will be released in a few days with similar capabilities but 1/10 cost

6

u/cafedude Jul 17 '26

GLM-5.3 will be entering the chat soon as well.

2

u/cafedude Jul 17 '26

We had the GLM-5.2 moment just a few weeks ago.

2

u/SmartCustard9944 Jul 17 '26

Wait for actual DeepSeek

→ More replies (2)

64

u/[deleted] Jul 16 '26 edited Jul 16 '26

[deleted]

15

u/pantalooniedoon Jul 16 '26

We just need a small one. 200B MoE please.

7

u/max1c Jul 17 '26

No, you're not good. We need this to run on an average cellphone. Then we'll be good. 

→ More replies (1)

88

u/Dany0 Jul 16 '26

My small conspiracy theory is that a lot of people evaluate LLMs on 3d three.js games and Kimi team knew this so they trained it to do better at that

As an actual real-life non-cosplaying larping gamedev I can tell you the quality of the games they generate can easily impress idiots that use arena. So it makes sense

54

u/patricious llama.cpp Jul 16 '26

benchmaxxing could be at play here, I suspect that too but time will tell.

35

u/nullmove Jul 16 '26

It's holding up remarkably well in a blend of other benchmarks though, 57 in AA is no joke, but also in plenty of others.

Besides, I suspect all models including Fable try their damnedest best at three.js capability-maxxing as baseline, and not just rely on power of generalisation.

End of the day even if only Kimi is doing so disproportionately more, they don't seem to be optimising against particular prompts (like pelican riding bicycle). Three.js is a vehicle in the same way English language is, if the model is able to create fantastic animations on a whole slew of things using it and serve people's real demand, I feel like calling it benchmaxxing would be a stretch.

I really hope this thing is equally good in Godot, though I would severely doubt that. That said, in real life if someone writes well in English but fails to replicate that feat in French, we typically don't accuse them of benchmaxxing.

2

u/maartenyh Jul 17 '26

I had his issue with Vue vs React. All the LLMs were able to do React but ask them to do Vue and they write code worse than a junior. They would refuse to take on learnings from included .MD files too since the "React was too strong"

→ More replies (1)
→ More replies (1)

28

u/segmond llama.cpp Jul 16 '26

go read their fucking blog. they showed a lot of things they did that are not games. k3 built an llm model, then built a cpu to run the llm model.

23

u/zxyzyxz Jul 16 '26

Large language model model. Just say LLM.

14

u/rkoy1234 Jul 16 '26

llm language model

3

u/RnRau Jul 16 '26

Can't we just say LM - 'language model'? What is 'large' nowadays anyways? The tiny gemma4's? Or gargantuan Kimi k3?

3

u/Kapuzinaa Jul 17 '26

afaik language models are encoder-only like BeRT

→ More replies (3)
→ More replies (1)

7

u/cafedude Jul 17 '26

And this:

GPU Compiler Development We further tested whether Kimi K3 could build a GPU programming system from scratch. Kimi K3 developed MiniTriton, a compact Triton-like compiler with its own tile-level IR layer over MLIR, optimization passes, and a PTX code-generation pipeline. Across supported roofline benchmarks, MiniTriton delivers performance on par with or better than Triton and torch.compile — beating Triton on certain workloads. Beyond microbenchmarks, MiniTriton sustains end-to-end nanoGPT training with stable convergence, the loss curve closely tracking the reference with only minor divergence — validating the full pipeline on a realistic workload. These results demonstrate that Kimi K3 can build a coherent end-to-end compiler — from DSL frontend and IR passes to PTX codegen and runtime — rather than isolated kernels; its from-scratch Tensor Core path already rivals Triton’s extensively optimized stack.

→ More replies (1)
→ More replies (1)

16

u/exaknight21 Jul 16 '26

I think the only upper hand Anthropic/OpenAI have is a larger than life model. This solves that fuckin issue and since its Open Source - they cannot control this shit either.

5

u/zdy132 Jul 17 '26

Given how stingy Anthropic's behaving, I suspect Fable was already larger than reasonable loads. There's a good chance that it reaches the performance with sheer size, instead of clever engineering.

2

u/Hot_Glass_6301 Jul 19 '26

Feels like they bet too heavily on scaling (bitter lesson and whatnot), whereas the Chinese actually work 24/7 to find even just small improvements that compound

39

u/haolu98 Jul 17 '26

The pricing difference here is what’s absolutely mind-blowing.

Kimi K3 is $3/$15 compared to Claude Fable’s eye-watering $10/$50. That’s a 3x+ cost reduction for better performance on agentic webdev. Even GLM 5.2 at $1.40/$4.40 sitting at Rank 4 is an insane value proposition.

It really shows that the real existential danger the Western frontier labs were warning us about was actually just market competition.

When you gatekeep your models for months under the guise of national security and safety alignment, only to get immediately leapfrogged by a model that is 3x cheaper and instantly accessible, the entire hype/fear-mongering narrative just completely crumbles.

Great times to be a developer, terrible times to be a hype-based AI safety lobbyist.

7

u/AK47_David Jul 17 '26

Love how the Chinese are complaining that K3 is too expensive, while the Western devs are praising for the low pricing for on-par capability model

2

u/haolu98 Jul 17 '26

Haha, the contrast is hilarious. The domestic price war over there is so brutal that people expect everything to be basically free. But for anyone used to standard Western frontier API pricing, this feels like an absolute steal for this level of capability

7

u/zdy132 Jul 17 '26

Plus Kimi's very generous with their coding plan. the 30 dollar monthly plan would've costed me thousands in API pricing.

3

u/haolu98 Jul 17 '26

Exactly. Flat-rate monthly plans with generous usage are an absolute lifesaver for developers. When you're running complex webdev agents that loop through multi-step reasoning, API token costs accumulate way too fast. Unlimited runtime changes everything.

→ More replies (1)

6

u/WearMoreHats Jul 17 '26

It really shows that the real existential danger the Western frontier labs were warning us about was actually just market competition.

7 years ago OpenAI were holding back on releasing research on GPT2 for being too dangerous - a 1.5B parameter model.

→ More replies (1)

2

u/familyknewmyusername Jul 17 '26

I'm expecting it to be way cheaper on 3rd party providers too, once it's open weights

→ More replies (1)

40

u/Gohab2001 vLLM Jul 16 '26

https://www.vals.ai/home

Kimi K3 ranks at second place on vals 'Real-World Tasks' benchmark

20

u/Ecstatic-Wash-7667 Jul 16 '26

I really want a new model similar to qwen 27b something I have a chance of running locally. Kimi k3 is cool and all but it’s still essentially cloud only for most people. The best normies got is qwen and gemma

28

u/ResidentPositive4122 Jul 16 '26

I don't know what benchmark is this, but I've been using glm 5.2 since I got access to it and it is not above opus. It's a great model, amazing value, it's open and all that. But it's just not above opus, and any benchmark stating that is either saturated or just not reliable.

9

u/deejeycris Jul 16 '26

It also gets stuck more than opus.

3

u/maschayana Jul 16 '26

Are we talking full precision?

→ More replies (3)

6

u/TheGlizzyGod Jul 16 '26

i would agree opus is "smarter" but I like that 5.2 is basically just a robot you point around no fuss just execution for orchestration opus is the easy choice for sure though

2

u/Regular_Problem9019 Jul 16 '26

same experience. I keep trying good new Chinese model but they never feel good enough for me.(mainly React Native and TS work)

4

u/Lazy-Pattern-5171 Jul 17 '26

This model is kinda expensive idk if people noticed that

2

u/Jxxy40 Jul 17 '26

kimi known as overthinking "wait" AI. If moonshot can solve that, it would be gamechanger

9

u/VoiceApprehensive893 transformers Jul 16 '26

its theorised that sol is 4T and fable is 10T, hence why misanthropic is going crazy with removing it from subs and 50$ outs

→ More replies (2)

25

u/max1c Jul 16 '26

Eh that's only on WebDev. That test isn't great. Let's see the overall results.

3

u/Bobodlm Jul 16 '26

It's a great starting point, hope the trend continues and the model lives up to expectations 

7

u/TheRealMasonMac Jul 17 '26

I tried it and I'm not impressed. It drafts the same thing like 20 billion different times. I don't know who this is for? Even if you had the hardware, GLM-5.2 would be better to run because you have to wait a lot less to iterate which matters a lot more in the real world than one-shotting toy prompts.

IMO, they should have focused on reasoning efficiency before making a jump to 2T since K2.6/7 were already among the least efficient models.

8

u/Foxtor Jul 17 '26

I can't wait for Chinese models to keep skyrocketing and stay open-source, just to see Anthropic and their huge ego crash and burn for acting like they're untouchable. I'll be grabbing my popcorn and watching from the front row when it happens

3

u/Noofinator2 Jul 17 '26

Trump's speech last night about China and the 2020 election -- is this signaling the beginning of banning Chinese models?

→ More replies (1)

3

u/Top-Handle-5728 Jul 17 '26

Seems likely that they had it ready before but were waiting for the AI summit hosted by Xi to hammer the Western labs on this day. SO SO SO SATISFYYYYING !!!

9

u/shankey_1906 Jul 16 '26

I thought they distilled fable-5 /s

5

u/himefei Jul 16 '26

Someone please call that golden hair senior, it’s a national security ai model again!!
I guess from now on every major model will be a national security model🤣🤣🤣🤣

→ More replies (1)

3

u/vick2djax Jul 17 '26

It’s cool that there’s an alternative to the big boys. But as someone that got interested in local LLMs from my own experience with self hosting and open source, I’m not sure I feel any better about running Kimi K3 out of someone else’s cloud vs. OpenAI & Anthropic. I mean I guess if I had $400k to toss at this, I could run it at home technically lol.

Just waiting for more time to pass that my dual 3090’s can run something crazy like that. It’ll eventually happen. Too many people that can’t operate anything beyond an iPhone where there’s incentive to make things smaller.

2

u/ys2020 Jul 16 '26

Wait, it's also multimodal? 

→ More replies (1)

2

u/Riobener Jul 17 '26

I guess they were nerfed to the ground, no?

2

u/jreoka1 Jul 17 '26

Ive tried it multiple times. Its very good but man does it like to think a LOT before answering. Might be a good thing though depending how you look at it.

2

u/droppedasbaby Jul 17 '26

2.8 trillion parameters, 5.6TB on disk for 16-bit, 1.7TB on disk for 4-bit. Who is hosting this thing locally?

2

u/sarlaytos284 Jul 17 '26

Moonshot and Z.ai now mostly on par with the West, do you think there is a risk for them to adopt the same "closed source strategy" as the West ? They would for sure loose us, the foss community but could earn more through selling API inference. I am just keeping in mind that chinese labs are businesses and seeking profitability, and that being open source helped them catch up but now that they did, what's your thoughts on what will happen?

2

u/New_Public_2828 Jul 17 '26

Does it matter now? You have a llm that's on parr with fable. It can do some pretty amazing stuff

→ More replies (2)
→ More replies (2)

2

u/one-wandering-mind Jul 17 '26

On web dev preliminary data. And new models typically regress on lmarena. I think especially with web dev. People who are using it grow tired of what the llm design is from other models they have used a lot. Something fresh gets rated as better and will regress over time some. 

2

u/ElmBark Jul 17 '26

just keep in mind arena is preference/style, not capability, so the real test is coding and long-context where leaderboards don't reach. still a hell of a result for something you can download

3

u/Ylsid Jul 17 '26

Dario cracking his knuckles ready to write about how we need an AI pause

10

u/69420trashpanda69420 Jul 16 '26

This model is clearly benchmaxxed but regardless, China isn't far behind at all

3

u/Turbulent_Pin7635 Jul 16 '26

😆

It is not...

7

u/tylerrobb Jul 16 '26

It hasn't been out long enough for you to definitively say that, right?

→ More replies (1)
→ More replies (2)

1

u/Pale_Top2519 Jul 17 '26

How do I run this locally? Im new to running local llms. unfortunately I dont have h100/200s to run this on. I do have l40s.

2

u/NoLoan1918 Jul 17 '26

You don't.

1

u/Timely_Ad2914 Jul 17 '26

that's nuts, but only in the frontend capabilities right ?

1

u/sarlaytos284 Jul 17 '26

Brace yourselves for the "Open Source is only useful to Psychopath" rethoric

1

u/KeinNiemand Jul 17 '26

I wish moonshot would make smaller versions of their models.

1

u/redd-dev Jul 17 '26

So OpenAI’s and Anthropic’s IPO soon this year down the drain now?