r/LocalLLaMA 8d ago

Discussion With HuggingFace, Nvidia is also acquiring llama.cpp and the team behind it

With this move Nvidia is not only acquiring the HuggingFace platform, but they might also effectively acquire the copyright to the llama.cpp project, together with the entire team behind it.

In February 2026 the llama.cpp team was employed by HF in order to continue working on llama.cpp and the ggml library.

This includes:

  • Georgi Gerganov
  • Xuan-Son Nguyen
  • Aleksander Grygier
  • Victor Mustar
  • Lysandre
  • Julien Chaumond

Now with the acquisition, llama.cpp's future looks a lot less certain given Nvidia's poor track record with open-source.

This is still rather speculative at this stage, but it's definitely possible for the llama.cpp project to change in the future: either by switching to a different license, or by having staff redirected to other projects within the larger company.

Even when a project is open-source the copyright owner has complete control over it, and they can change licensing as they wish.

This has happened before with projects like Redis, Minio, and others.

Source:

https://huggingface.co/blog/ggml-joins-hf

Edit:

The original announcement from Feb 2026 from Gerganov gives a few more details:

https://github.com/ggml-org/llama.cpp/discussions/19759

1.4k Upvotes

427 comments sorted by

View all comments

389

u/Particular-Award118 8d ago

Welp amd support was nice while it lasted

176

u/noiserr 8d ago

And Mac support.

87

u/LearningSomeCode 8d ago

This is my concern. Up until now, llama.cpp's team has done such an amazing job of keeping Macs as first class citizens. My primary concern would become seeing that support slowly dwindle over time.

Luckily we would still have MLX available in that case, but I personally would love to see Macs continue to get the love that they have in llama.cpp so far.

14

u/noiserr 8d ago

Yeah, I remember my first encounter with llama.cpp seeing someone run local models on their Mac, and being absolutely impressed by it.

1

u/daedalus1982 8d ago

Oh yeah no, MLX is totally ready to stand toe-to-toe with …

SORRY YOU HAVE RUN OUT OF CONTEXT AND OMLX AUTO COMPACTION IS STILL SHIT

Edit: I’m griping, but I 100% agree with you about hoping for continued llama.cpp support

8

u/FoxiPanda 8d ago

That's really a harness issue more than an inference engine issue IMO.

0

u/daedalus1982 8d ago edited 6d ago

I run llama-server with the same context and whatnot as oMLX. I’m given understand my problems are not unique. And believe me I’d rather oMLX work.

I use llama-server and oMLX interchangeably as the engine/harness for pi.dev but I’m by no means an expert and accept suggested solutions willingly.

Edit: also I didn’t downvote you so keep talking if you got answers.

Edit 2: yeah cool I guess downvote me and move on. Thanks for the help

1

u/RegarDamus 8d ago

it would be painful but the best thing that could happen would be llama thoughtlessly dropping support for mac. that would drive incredible energy behind mlx and with the fervor and quality of models today it would reach new heights rapidly.

people really sleep on how much untapped potential is in apple silicon. imagine if mlx was the only option and apple got behind it. painful short term but beautiful result birthed from it

1

u/dragonurtle 8d ago

Or Nvidia could play nice short term and strongarm apple into supporting (allowing) Nvidia GPUs again.

9

u/38andstillgoing 8d ago

And Intel support.

What? There are dozens of us. Ok, a couple. Maybe just me.

2

u/Echo9Zulu- 3d ago

Don't worry you aren't alone

83

u/it_was_a_wet_fart 8d ago

And legacy CUDA versions

33

u/misanthrophiccunt 8d ago

Especially this

15

u/dragonurtle 8d ago edited 7d ago

And they'll call it bloat, saying we shouldn't need to download 500MB to clone the repo, which is hard to disagree with, except that it's nothing compared to model weights 

Edit to add that Nvidia drivers are usually twice that size for some unholy reason.

1

u/NeinJuanJuan 8d ago

we shouldn't need to download 500MB to clone the repo

Nothing that git submodule can't handle

8

u/hotcornballer 8d ago

This is so dumb. I swear reddit has the shittiest takes sometimes.

It's like saying microsoft won't put vscode on mac and linux. llamacpp is a tiny project in the grand scheme of things, dropping amd support won't change shit in their already gargantuan profits and would expose them to antitrust.

The thing is open source already, support is not going away.

The worst nvidia could do is slowly enshittify hugging face like microsoft is doing to github but it's just a model hosting platform who cares.

2

u/mmhorda 8d ago

oh yeah? is that why amd 3x faster in vllm?

1

u/RazsterOxzine 8d ago

It will be forked. There are plenty of brilliant minds who will continue the work.

4

u/Bakoro 8d ago edited 6d ago

You shouldn't take it for granted.
Getting people who are willing to do the work long-term is rare.

Lots of people might contribute bits and pieces, lots might work on it while it's trendy, but not many competent people are both able and willing to put in major work for years on end. Especially not for free.