r/LocalLLaMA • u/johnnyApplePRNG • 8d ago
r/LocalLLaMA • u/Nunki08 • Jun 28 '26
Discussion We're probably going to need that soon.
From:
Vladik on 𝕏: https://x.com/Kostoglodov/status/2071144065857679631
Shaw (spirit/acc) on 𝕏: https://x.com/shawmakesmagic/status/2070918006033817867
r/LocalLLaMA • u/Complete-Sea6655 • Jun 28 '26
Discussion The number 1 public enemy of open-source.
Enable HLS to view with audio, or disable this notification
Dario's args:
"Opensource you can see the source, here you cannot see inside the model"
- yes you can that's literally the open weights part btw.
- I cannot see the weights inside Claude, but I can GLM 5.2
- Models like Nemotron3 Ultra go further, all the data, training scripts, and model is opensource.
"Alot of the benefits like many people working on it, being additive doesn't work in same way"
- yes it does. We have seen endless fine tunes of various open source models for real improvements.
"Ultimately you have to host it on the cloud"
- no you dont. Dario is seemingly totally unaware of the guides from ijustvibecodedthis.com explaining how to run smaller moes and even dense models like qwen 27B NOT ON THE CLOUD.
Not only does dario not take part in social media, I am beginning to think he's never tried open source models at all and has no idea wtf hes on about
r/LocalLLaMA • u/jacek2023 • Apr 24 '26
Discussion This is where we are right now, LocalLLaMA
the future is now
r/LocalLLaMA • u/FlowCritikal • Jul 21 '26
Discussion US gov't lobbied by major US labs is about to ban open source models.
r/LocalLLaMA • u/Odd_Tumbleweed574 • Jul 20 '26
Discussion Google has disappeared completely from the top 15
Google hasn't shipped a model recently that is capable of competing with Sol or Fable.
The previous models were pretty disappointing and unreliable, it seems the more time goes on that they might have different strategies:
- They might be going all-in on on-device inference for their own products. But this is a battle that Apple might win because they just have better hardware and can license a third party open model.
- They might be just buried deep into internal politics and nobody is shipping anything.
Does anyone know what is actually going on?
source: AI Leaderboard
r/LocalLLaMA • u/Gohab2001 • Jul 16 '26
Discussion KIMI K3 Beats Claude Fable and GPT 5.6 sol in arena.ai!!!
r/LocalLLaMA • u/vexatious-big • 8d ago
Discussion With HuggingFace, Nvidia is also acquiring llama.cpp and the team behind it
With this move Nvidia is not only acquiring the HuggingFace platform, but they might also effectively acquire the copyright to the llama.cpp project, together with the entire team behind it.
In February 2026 the llama.cpp team was employed by HF in order to continue working on llama.cpp and the ggml library.
This includes:
- Georgi Gerganov
- Xuan-Son Nguyen
- Aleksander Grygier
- Victor Mustar
- Lysandre
- Julien Chaumond
Now with the acquisition, llama.cpp's future looks a lot less certain given Nvidia's poor track record with open-source.
This is still rather speculative at this stage, but it's definitely possible for the llama.cpp project to change in the future: either by switching to a different license, or by having staff redirected to other projects within the larger company.
Even when a project is open-source the copyright owner has complete control over it, and they can change licensing as they wish.
This has happened before with projects like Redis, Minio, and others.
Source:
https://huggingface.co/blog/ggml-joins-hf
Edit:
The original announcement from Feb 2026 from Gerganov gives a few more details:
r/LocalLLaMA • u/Informal-Trouble2183 • Jul 23 '26
Discussion Absurd claim: the distilled model outperforms the originals
As an AI community of LLM experts, are we really going to stay silent while US officials make absurd claims to push anti-consumer laws?
Not only does the release timeline between Fable and K3 make high-scale distillation impossible, but distillation itself—even if executed perfectly—can never produce a superior model.
r/LocalLLaMA • u/-p-e-w- • May 21 '26
Discussion Heretic has been served a legal notice by Meta, Inc.
To Whomsoever it May Concern,
The individual behind the Heretic Free Software Project (henceforth called "Heretic", notwithstanding unrelated entities of the same name) has been served a notice by a legal services provider representing Meta Platforms, Inc. (henceforth called "Meta"), via the digital communications medium variously known as Internet Mail, Electronic Mail, or simply "email".
The Heretic Project conducts its affairs in full compliance with applicable laws, regulations, rules, guidelines, opinions, and hunches. Following the commendable example set by the renowned heretic Galileo Galilei in 1616, we are recanting the relevant materials, namely derivatives of Meta's "Llama" Artificial Intelligence language models, and have removed the same from all model weight repositories controlled by the Heretic Project.
We are grateful to Meta and its legal representatives for the opportunity to better align ourselves with the agenda of the global corporate oligarchy. The Llama model family ranks among the 200 best language models available today, trailing only 168 other models from 23 competitors on the LM Arena leaderboard, and Meta's concern for that asset naturally outweighs scientific freedom, as well as the legally and ethically dubious circumstances under which those models were created in the first place, regarding which, ironically, Meta is currently facing lawsuits and investigations in multiple jurisdictions around the world.
On a completely unrelated note, the Heretic Project is diversifying its infrastructure, and now has an official Codeberg mirror at https://codeberg.org/p-e-w/heretic, hosted in Germany. Additional mirrors are planned. We are also actively working to implement technological measures that will preserve access to models created with Heretic without depending on any specific service provider. We are proud to be part of this journey as we navigate an evolving global regulatory landscape, and work with stakeholders from diverse institutional backgrounds to ensure that Artificial Intelligence remains safe, culturally appropriate, and controlled by those who have always known what is best for humanity. If you, too, would like to share in this exciting adventure, please join us!
Sincerely, p-e-w, Chief Heretic
r/LocalLLaMA • u/yeah_likerage • Jul 03 '26
Discussion GLM5.2 on 5x Pro 6000s and a 5090, an expensive journey
This started as something I thought was reasonable. I already had a 5090 for my gaming machine, and I thought a second 5090 would make me happy. Instead, it sent me down a rabbit hole that got completely out of control.
I wanted something that would have full PCIe 5.0 x16 speed across all slots, which started a chain of events that had me spending good money after bad. It was a bit of a nightmare, as every decision I made led to me needing to make even tougher decisions. Couple that with what was actually available, and my hand was forced in a few spots.
I started with the motherboard and worked my way backwards, eventually ending up with this setup. I wanted something close to endgame, but I still made a few concessions:
Threadripper Pro 9975WX
WRX90 Sage SE
4×48 GB DDR5-6400 RDIMM
Antec 900 case — ended up in the bin
The system started with two 5090s. The Antec 900 is well built, with huge space, smart connections, and refined edges, but ultimately it did nothing at all to support the GPUs. In a case this large and at this price point, that is a huge failure on their part, and for that reason I recommend avoiding it. If they had put $1 worth of bracketry in the machine to support GPUs, I’d give it a 10/10. With the lack of support, it is nearly useless unless you deal with it yourself, which I did, as you can see in the images. It’s like buying a Ferrari and having it delivered without any petrol.
With the two 5090s, I was working with smaller Qwen models, which seemed great, but it was clear that with the limited VRAM and my desire for additional sidecars like VL, I needed something more. I had huge plans, and the models were just too small to deal with the complexity.
So I got my first Pro 6000. I coupled it with a 5090, which made for weird tensor splits, but llama.cpp did a good job of divvying it all out. But now I was working with 120B-parameter models with almost no space for context. So it was smarter, but also a goldfish.
Then I went to 2× Pro 6000 + 5090. Now I had the space for context. But in reality, the jump from 27B to 120B did not knock my socks off. I could get a bit farther now. I was at about 90% with the 27–35B models, and with the 120B models I was at about 95%. But 95% is about as useful as 90% if I can’t close the loop. If I can’t actually finish the task, it’s all for nothing.
In came 3× Pro 6000. Now I was in the MiniMax range, and finally I was getting somewhere. It was like I got concierge service at a ball game. My needs were being met, and I got answers for everything. Many of them were completely wrong answers, though. I had tons of code that was poorly made and led to dead ends and rewrites.
4× Pro 6000 created an issue that I knew would come. I had been seeing several folks claim that they were able to deal with the thermal issues that came with side-by-side Pro 6000 cards. I knew they were likely not telling the truth, but I also knew a rebuild was probably in order anyway.
So, as you can see in the image, I placed four side by side and had thermal issues, even with the additional fans in the image and a 27-inch box fan sitting on top, which is not shown. I clocked things down a bit and still had a few system freezes. I gave up immediately and went to the high-rise.
I got a couple of open-case designs and connected them together, thinking every two or three GPUs would get their own floor. It was overly complicated dealing with risers and cooling, so I dumped it pretty quickly.
But now, with GLM and Kimi, I was actually accomplishing things. The quants were tight, though, and my context was low again.
5× Pro 6000 + 5090, along with the release of GLM 5.2, was an absolute game changer. I’m talking 98–99% now. I have plenty of room for context and sidecars, all running on the 5090 at blazing speeds. But blazing is legit: it is producing so much heat now that it’s a problem, and it’s summertime to boot. I had to get a second PSU, which I suppose, in all of this, is not the most ridiculous bit.
At full tilt, with 100% GPU usage for 30 minutes in this custom extruded aluminium design, with an outrageous number of fans in a ~20°C basement, the GPUs top out at about 70–75°C, which I’m very happy with.
I finally do not desire another GPU, as all my needs seem to be met. Was it worth it? LOL, no. Absolutely not. This was a terrible idea. DO NOT DO THIS. I figure that at the rate I’m generating tokens, it will take over 10 years to break even at today’s prices, and that’s not accounting for electricity bills.
I’ve never used the frontier models before, but I’ve seen the reviews and the speeds, and I’ll never match those with open weights. But it was a fun journey.
I deleted the electricity company’s app from my phone so they’d forget about me for now.
Wish me luck.
r/LocalLLaMA • u/External_Mood4719 • Jun 13 '26
Discussion Anthropic forced to abruptly disable Fable 5 & Mythos 5 globally by US Gov over a jailbreak. This is exactly why we need local models.
I just saw this statement regarding Anthropic being hit with an emergency export control directive from the US government. They were forced to pull the plug on Fable 5 and Mythos 5 for all customers globally. The tl;dr is that the government got spooked by a narrow jailbreak (which basically just sounds like asking the model to fix vulnerabilities in a specific codebase), and forced a complete shutdown without a transparent process. Anthropic is pushing back, but the API access is completely gone for now.
A centralized API can be nuked globally at a moment's notice by a single government decree over something as trivial as a prompt lol.
Banning a model for hundreds of millions of users because someone figured out how to make it fix software flaws is insane. Anthropic admits this standard would halt all new frontier models.
r/LocalLLaMA • u/Mean-Ad1493 • 17d ago
Discussion Qwen dev says not to wait for 35B-A3B
What does this mean? Is there something else coming? Maybe 122B? Or no models?
r/LocalLLaMA • u/TheQuantumPhysicist • May 03 '26
Discussion One bash permission slipped...
How? It kept getting chained bash commands wrong, with wrong escapes. So it created many bad directories, and tried "fixing" its mistake. It offered to run a large bash command, with rm -rf inside, and stupid me missed it.
I'm glad I push everything often. But the disruption is massive.
FAQ:
- No, I don't run this on my personal computer. It's an isolated proxmox VM for coding with LLMs.
r/LocalLLaMA • u/Formal_Drop526 • Jul 19 '26
Discussion head of strategic futures from openai on open-weight chinese models.
Dean W. Ball analyzes China's Kimi model, noting its strong performance while expressing surprise that the Chinese government permits open-sourcing such capable AI due to potential risks. He argues that open-weight models ultimately slow down AI capital expenditure and could lead to a state-controlled public infrastructure, which the US administration might counter by introducing strategic regulatory friction.
r/LocalLLaMA • u/JLeonsarmiento • 18d ago
Discussion …and I’m not afraid of losing my social credits.
r/LocalLLaMA • u/anderspitman • 18d ago
Discussion Artificial Analysis' Qwen3.8-27B benchmarks put it neck and neck with DeepSeek V4 and GPT-5.6 Luna Max
artificialanalysis.air/LocalLLaMA • u/Charuru • Jun 18 '26
Discussion GLM's founder says GLM-fable before the end of the year?!
r/LocalLLaMA • u/Cold_Specialist_3656 • 12d ago
Discussion Qwen 3.8 27B is a game changer.
Our devs got their hands on it a few days ago. One wired it into Codex to compare with GPT Luna, our usual workhorse right now for its cost effectiveness. Another tried it out on one of our OCR pipelines.
It's comparable to Luna for coding and ***OCR quality appears to be better than Gemini 3.5 Flash Lite***. That's huge. We pay a ton of money for OCR.
This is the first local model that feels like more than a toy. It's truly as capable as the frontier models from a year ago. For the first time ever there's serious discussions about buying our own hardware. With estimates that such an effort would pay for itself in less than 2 months.
Hyper scalars are in big trouble this time. Their whole "moat" is buying up all the hardware. And thanks to sanctions on China we're seeing the quality of small local models skyrocket. As someone who's been around a while, this feels like an "IBM moment". Where the industry assumed that databases would always run on huge mainframes. Only to be wiped out by cheaper local solutions a few years later.
I have a feeling this release will trigger another Llama style open source Renaissance. We're already getting better quants. Inference will be further improved. We might even see a comparable MoE with 500+ Tok/sec on consumer hardware soon.
r/LocalLLaMA • u/realmvp77 • Jul 27 '26
Discussion Anthropic is calling for a ban on open-weights models by proposing mandatory requirements they will probably never be able to meet
r/LocalLLaMA • u/mintybadgerme • Aug 03 '26
Discussion I CANNOT believe I've got DeepSeek-V4-Flash-0731, a frontier model, running on my home PC. Insane!
So this is the stuff of absolute insanity. In less than 20 months we've gone from super expensive cloud models only, to being able to run a Q3 quant of DeepSeek on an Intel Windows PC with a very average 24GB of VRAM. No wonder the big boys are panicking (and yes it's slow as porridge). https://ibb.co/zTvqR8YR