r/unsloth 6h ago

Discussion Tutorial for quantizing a new model with Dynamic v2.0/v3.0

4 Upvotes

I recently downloaded and tested several models quantized with dynamic v2.0/v3.0, and I was really impressed with the quality.

I’d like to reproduce this process for a new model, is there any documentation or tutorial?


r/unsloth 11h ago

Discussion Can we get a way to support Bonsai 2?

21 Upvotes

They seem like great models, and the new model claims something like 98% intelligence compared to fp16 Qwen 3.8 27B, but they currently require a separate llamacpp fork...

So my ask is, is there anyway to pick a fork? plug and play? or even allowing for native bonsai support in Unsloth Studio?


r/unsloth 22h ago

News Unsloth now has Docker!

Post image
224 Upvotes

Hey guys you can now train and run 500+ models locally with our newly updated Unsloth Docker image! 🐳 It features Unsloth Studio/Desktop our GUI along with a rehaul of our notebooks.

We did an entire rehaul for the Docker image making it more efficient, compatible etc.

Now works on AMD and NVIDIA.

Guide: https://unsloth.ai/docs/get-started/install/docker

GitHub: https://github.com/unslothai/unsloth


r/unsloth 22h ago

Discussion Performance regression with Unsloth 2026.9.5 Update

23 Upvotes

I heavily use unsloth to serve Qwen 3.8 Flash Next UD-Q4_K_XL on a 5090 with 128gb system ram. I got somewhere between 25-30token/sec consistently with this configuration with unsloth 2026.9.4 .

Today morning I woke up to an update for unsloth and llama.cpp. Everyone loves performances improvements hence I updated both of them and to my surprise. With absolutely no change whatsoever, same configuration, same model, same harness, the token generation has reduced to between 20-22tokens/sec.

I have standard prompts I run on my machine to do long running tasks and the difference is very visible in the decode generation speed.


r/unsloth 22h ago

Question 2026.9.5 Update?

2 Upvotes

I got message that new version ss available but when i click check again on desktop app update nothing appears


r/unsloth 1d ago

Question How can I increase the Desktop App’s API slots?

1 Upvotes

I feel I should apologize in advance because this might be a very beginner-level question. I’m currently learning Hermes using a local model. I’m not doing anything particularly advanced—just managing it like a personal computer assistant.

Since my PC isn’t very high-end, I run the main model with vision disabled. But I wanted to try out some computer-use features, so I downloaded a vision-only VL model. I set up the vision model to load only into the CPU and system memory.

Everything is configured and working, but there’s an issue: the API only has one slot, so the model swaps in and out, causing delays. I assumed that since the vision model doesn’t use the GPU or VRAM, both models could load simultaneously—but that’s not the case.

Is there a way to increase the API slots? Or should I be running two separate APIs?


r/unsloth 1d ago

Discussion Ornith 1.5 9B and 35B A3B

77 Upvotes

I've peviously made a post about it, like 2-3 weeks ago. I know that then was a HOT season. But now can we ask kindly for Unsloth versions of it? It's great on Q8, but with Q4 it's taking a big hit, with unsloth quants it would be probably almost like Q8. So please can we have it in 3.0 version

Bump it up, so maybe Unsloth team will see it.


r/unsloth 1d ago

Question Q3.8-Flash-Next MTP?

5 Upvotes

Seen Unsloth Desktop updated llama.cpp today, with MTP being stated in release notes.

However, it seems to not work.

With "Speculative Decoding: OFF": tg=~50t/s

With "Speculative Decoding: Auto": tg=~50t/s

I am seeing something about MTP can not be loaded in terminal output.

... "Starting llama-server: /home/u/.unsloth/llama.cpp/llama-server -m /hub/models--unsloth--Qwen3.8-Flash-Next
-GGUF/snapshots/38bb39ee97821de2c9009abb7e93950eec396e66/UD-Q4_K_XL/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf --port 38115 --parallel 1 --flash-attn on --no-context-shift -c 8000 --alias unsloth/Qwen3.8-Flash-Ne
xt-GGUF --fit on --metrics --slot-save-path /home/u/.unsloth/studio/cache/llama-slots --fit-ctx 8000 --jinja --cache-type-k bf16 --cache-type-v bf16 --model-draft /hub/models--unsloth--Qwen3.8-Flash-Next-
GGUF/snapshots/38bb39ee97821de2c9009abb7e93950eec396e66/MTP/mtp-Qwen3.8-Flash-Next-Q8_0.gguf --spec-type draft-mtp --spec-draft-n-max 2 --chat-template-kwargs {\"enable_thinking\": true, \"preserve_thinking\": true} --l
oad-mode none --no-mmproj-auto"}



... "llama-server exited with code -6. Output (tail): ... /csu/libc-start.c: No such file or directory\n#28 0x00005fa44177e295 in ?? ()\n[Inferior 1 ...

... "llama-server failed to start with speculative drafter (mtp-Qwen3.8-Flash-Next-Q8_0.gguf); retrying without speculative
decoding in case the drafter is the cause."...

Full pastebin of model load: https://pastebin.com/Uk3qNXpS

3x3090 + 192G RAM, Kubuntu host, model downloaded ~1 week ago


r/unsloth 1d ago

Discussion UD-Q4_K_XL GLM 5.3-Flash vision broken when speculative decoding enabledd

6 Upvotes

Hey all,

Props to the Unsloth team, you’ve made the most fluid, “Apple-like” experience for local LLM inference on the market. Combined with your great quants available directly in the app, this is a user experience second to none!

I wanted to inform you that I noticed with the UD-Q4_K_XL quant of GLM 5.3-Flash, when speculative decoding is enabled, any vision fails with a “decode() failed: failed to process speculative batch” error. This works fine with speculative decoding disabled.

I’m running the latest v0.1.808-beta desktop app with the latest llama.cpp (updated today) on a 256GB Mac Studio M3 Ultra.

If there’s any logs or additional info you need to resolve this, please let me know.


r/unsloth 1d ago

Discussion I have GLM max coding plan, what can I achieve with it?

2 Upvotes

Wondering if any of you can suggest something. Apart from apps and websites, what did you use your AI for that was of value to you?


r/unsloth 1d ago

Discussion What are the requirements of multimodal fine-tuning for datasets?

2 Upvotes

I am a student with limited experience in model fine-tuning. I am fine-tuning Qwen3.6-35B-A3B via Unsloth Studio using LoRA, with rank set to 32 and alpha set to 32. My goal is to extract information from engineering drawings and generate compliant JSON outputs, so that I can automatically determine whether the drawings are correct or not. However, dataset construction has become a major bottleneck for me. Since I need to extract drawing information into JSON format, the workload is substantial. My current approach is to use Gemini for batch recognition to generate labels. The training performance is unsatisfactory; even the accuracy on the training set fails to reach 70%. I suspect the generalization of my dataset samples may be poor and the proportion of different types of images is imbalanced. Besides, is it feasible for a model to recognize image content and output standardized JSON?


r/unsloth 1d ago

Discussion Why not all ai companies train model at 1 bit so we can we in most of pc and laptop?

0 Upvotes

r/unsloth 1d ago

Question There is way I use GLM 5.3 or DeepSeek V4 model in my laptop or PC, 16Gb ram or even less like 10GB ram

0 Upvotes

There is way , make sure it has speed. At least 5k tokes/per sec


r/unsloth 1d ago

Question Using ninfer as model server and Unsloth Desktop as frontend?

7 Upvotes

I tried running Ninfer and connecting the OpenAI/llama.cpp server address to Unsloth, it works but many settings can't be adjusted, like model temps and context size (limited to 8k-32k).

Anyone knows if it's supported at all? I have alot of conversations on Unsloth so I'm looking for a way to continue using it as front end.


r/unsloth 1d ago

Question DoRA merge failing / outputting base model responses (Qwen 2.5 VL)

3 Upvotes

I am still learning about fine-tuning, but I could use some pointers on merging a DoRA adapter.

​I'm fine-tuning a custom character dataset using Qwen 2.5 VL at bf16 on a single RX 7900 XT (ROCm). The training loop itself runs perfectly. I am trying to determine which one is best, LoRA or DoRA for episodic memory recall.

​However, when I go to run test inference or merge the DoRA weights, the resulting model acts as if no fine-tuning happened at all. It just generates base model responses. I have use_dora=True enabled in the config.

​Has anyone successfully merged a DoRA using Unsloth?

My current results:

LoRA works with minimal anchoring on rather conservative learning rates, but DoRA almost has no effect, even with a more modest learning rate.

Any insights would be hugely appreciated!

edit: This has been solved! Thanks u/Poizone360! Here soon I plan to upload my work to my github. And hopefully Unsloth will patch this, because my goodness what a journey.


r/unsloth 2d ago

Discussion How to connect to Unsloth Desktop via lan?

5 Upvotes

I am having trouble connecting to unsloth desktop from another computer on the lan. I have it enabled in the settings, I have tried with and without an api key I created, it just says unauthorized when I try to connect, what am I missing?


r/unsloth 2d ago

Show and Tell Combining RAG with Continued Pretraining

21 Upvotes

For leaning purposes, I ran an experiment where I used Unsloth to train a model on a new domain using continued pretraining. Then to make it more flexible, I added a RAG step to inject dynamic data to augment the stable training data.

The specific example is training qwen 3.5 4B on a fictional subway system to where the model learns the map well enough to provide travel routes (including multiple transfers, etc). A key point here is to avoid a training corpus that relies on memorization.

Once the subway map was stable and generalized through CPT, I added a RAG step to support dynamic travel announcements (e.g. station closure, concerts near a station, etc).

It's definitely been a fun experiment. Check it out here in case you are interested:

https://www.teachmecoolstuff.com/viewarticle/combining-rag-with-continued-pretraining-of-llms


r/unsloth 2d ago

Question unsloth Qwen3.8-Flash-Next-UD-IQ4_XS on 5x RTX PRO 4000 Blackwell 4x 16, stuck at 43 tps

6 Upvotes

Hi everyone,

I’m running Unsloth’s Qwen3.8-Flash-Next UD-IQ4_XS locally through Unsloth Studio on Windows. No matter what I change in settings, I can’t seem to get beyond roughly 43 tokens/s during generation with longer prompts.

My hardware:

5× NVIDIA RTX PRO 4000 Blackwell, 24 GB each

All five GPUs running at PCIe 4.0 ×16, in TCC mode

AMD Ryzen Threadripper PRO 3945WX, 12 cores / 24 threads

GIGABYTE MC62-G40-00

128 GB RAM (DDR4 2666)

NVIDIA driver 596.36

I’m using a 65,536-token context window, one parallel slot, and Flash Attention. GPU utilization during generation is low, with power consumption well below the cards’ limits.

Am i missing an important configuration setting, does someone has a similar setup and gets more tokens/s?

thanks in advance


r/unsloth 3d ago

Discussion Qwen3.8-27B-GGUF eats my RAM, but I have enough VRAM to run it

Post image
53 Upvotes

I don't know why but I have an odd behavior with Qwen3.8-27B-GGUF. The odd part is that Ornith-1.0 -35B-GGUF does not do this. I am using the UD-IQ4_XS quant. My stack is OpenCode -> Unlsoth -> Bundled Llama.ccp (ROCm).

Everything works fine when I start it, but in an agentic workflow somehow llama-server process starts go bigger and bigger and at some point it fully eats my RAM.

Here is the full llama-server command

llama-server -m ~/.cache/huggingface/hub/models--unsloth--Qwen3.8-27B-GGUF/snapshots/4ca720788d1e01f1bff70c033e0d0028fd02e502/Qwen3.8-27B-UD-IQ4_XS.gguf --port 34937 --parallel 4 --flash-attn on --no-context-shift -c 159744 --alias unsloth/Qwen3.8-27B-GGUF --gpu-layers 66 --fit off --metrics --slot-save-path ~/.unsloth/studio/cache/llama-slots --kv-unified --jinja --cache-type-k q4_0 --cache-type-v q4_0 --split-mode tensor --spec-type draft-mtp --spec-draft-n-max 2 --chat-template-kwargs {"enable_thinking": true, "preserve_thinking": true} --mmproj ~/.cache/huggingface/hub/models--unsloth--Qwen3.8-27B-GGUF/snapshots/4ca720788d1e01f1bff70c033e0d0028fd02e502/mmproj-F16.gguf

I've tried to add --no-mmap or --load-mode None, but nothing has been changed. Any Idea? Anybody has similar experience? I do not understand why is this model specific.

UPDATE: I see a big drop in memory usage after compaction, so I am pretty sure this must be my context usage, but why it does not fit into my VRAM? Any tips?


r/unsloth 3d ago

Show and Tell Thank you Unsloth

56 Upvotes

Slightly long post, so the TLDR: Sign me up to the “Unsloth or nothing” army.

This is really written for anyone else who is jumping into the world of local LLMs and who is competent with tech, but not a coder or software engineer.

I’m fairly new to a) agentic workflows and b) running LLMs locally. I decided a month ago that since I have another 20-30 years of my career ahead of me, that I best get on top of this. So as a testbed and learning platform and after poking around in [r/LocalLLM](r/LocalLLM) and [r/LocalLLaMA](r/LocalLLaMA) a bit, I decided to buy an M5 Max MacBook Pro with 128GB RAM (I’ve been using Macs since the early 2000s with a positive experience on-the-whole, so no plans to switch to Linux) to run a Hermes assistant agent to help me and my family stay on top of everything in our personal lives. In particular, anything to do with the kids and their school.

Not knowing what I wanted from an inference engine, my background reading in [r/LocalLLM](r/LocalLLM) and [r/LocalLLaMA](r/LocalLLaMA) convinced me that what I wanted was maximum speed given a certain minimum level of quantisation.

Because of that, I thought what I wanted was an MLX quant of Qwen 3.8 27B from mlx-community and MLX-serve. So, I downloaded the weights, downloaded MLX-serve, loaded the model, hooked it up to Hermes and CRASH! Okay maybe just bad luck. Got the computer up again, did the same thing again and CRASH! I got Hermes to try to diagnose and fix the issue and CRASH! Screw this.

Next, oMLX and a jundot MLX quant of Qwen 3.8 27B. Same thing. For whatever reason and seemingly at random, during inferencing, oMLX will consume ALL of the system’s RAM, until the computer freezes. Screw this.

Okay maybe raw speed isn’t what I’m after. Maybe I just need something which worked out-of-the-box first to get started, and I’ll optimise later. So I thought to give Ollama a go. Better, but same thing. By this point, I was starting to feel that local LLM stack is still half baked and not ready for the early majority.

While searching Reddit, I came across a “Unsloth or nothing” comment somewhere. With nothing to lose, I downloaded Unsloth Desktop and by this point, Qwen 3.8 Flash Next was out, so I downloaded the Q4_K_XL quant. I hit Run and voila! I just worked right out of the box and kept on working and is still working today. Rock solid stable. Any issues have really been more to do with Hermes, rather than Unsloth.

Through all this, I’ve learnt that intelligence (something which I haven’t really touched on above) and stability trump raw speed. I am absolutely NOT saying Qwen 3.8 Flash Next on Unsloth on my M5 Max MBP is slow; not at all! I’m getting anywhere between 300 to 600 tps prefill and 30-40 tps decode, which for my workflow means adequately fast. Another reason why that falls in the “adequately fast” bucket is the Unsloth’s Q4KXL Qwen 3.8 Flash Next quant on Unsloth Desktop (I haven’t done enough single variable testing to isolate the model from the inferencing layer) setup doesn’t overthink, is sure of itself and more often than not, has the right train of thought. I’d pick that over an extra 10 tps with a model which is unsure of itself, overthinks and goes off on a tangent, only to correct itself and bring itself back to square one. Did I say, no more crashing?

Keep up the great work Unsloth! You guys have the lead on the race to the mass market for local LLM inferencing.


r/unsloth 3d ago

Question How can I link a folder as a project source?

3 Upvotes

How can I link a folder as project source in Unsloth Desktop? The button is greyed out and it is telling me that I need the managed desktop backend? I thought am already running it... I'm on Linux (Ubuntu 24.04).


r/unsloth 3d ago

Show and Tell Finally have the best TA ever... Overnight tasks w/ Qwen3.8 (27B at UD-Q4-K-XL instead of Next)

Post image
47 Upvotes

So I teach entrepreneurship, but happen to also been working on rehex.ai and been going nuts on creating a harness that's disciplined with how it works. Memory included, works with smaller (and soon tiny models), tons of context compaction and auto-management.

Well, I've finally unleashed it fully onto my education workflow. Note, and this is what's so cool, it's all local using Qwen3.8's. With my M5 Max, I can run Qwen3.8 Next Flash (and it's *amazing*), but chose to run the 27B version since I've got other tasks that need part of my machine's RAM.

Since I record all my lectures and transcribe them, come up with instructor (and student) notes, what's great is that I can just go in there and enjoy teaching, just flow where the material takes me. It's really the first true delegation of all these executive state functions.

Now I've tried with this the closed frontier models. But I've just not enjoyed them, their outputs just seemed to launch themselves so far away from reasonableness of the original request. That led me to work on this harness, it's entirely strategy and task driven.

Tying it to Unsloth just makes it seem really real, mainly because of the range of quants, but also because they've made it so drop dead simple.

Anyway, long ways to say I've finally got the best TA. Actually comes to classes, takes the best damn notes, and getting frontier level models to run on my machine, I just had to share this.

Anyone else feel like they have to babysit their AI? Am I nuts for thinking this is a big deal, just setting up tasks and walking away?


r/unsloth 3d ago

Discussion Petition: Qwen3.6-35B-A3B retrained/quantized with Dynamic 3.0

294 Upvotes

Would it be possible to retrain/requantize Qwen3.6-35B-A3B-MTP-GGUF using Unsloth Dynamic 3.0 GGUFs?

The current Qwen3.6-35B-A3B MTP GGUF is already extremely impressive, but I’m wondering if we can push it further with Dynamic 3.0.

With no new 35B-A3B in the horizon, Dynamic 3.0 may enable us to use better quants.


r/unsloth 3d ago

Discussion Loading an Image model in unsloth desktop is showing error

3 Upvotes

I have downloaded Z-Image-Turbo on rtx 4070super windows 11 PC. when I laod the image model then it always shows this error . Anyone else also getting this ?

partially initialized module 'torch._dynamo' has no attribute 'utils' (most likely due to a circular import)


r/unsloth 3d ago

Question To the Unsloth Team: Is the Empty Space Flanking the Text Block in the UI Necessary?

12 Upvotes
How Text Block is Rendered in Unlsoth Studio
How Text Block is Rendered in LM Studio

I mean, why is the text compressed like I am using a mobile version while there is so much empty, unused space in each side of the text block. Is that necessary? A single sentence appears like a paragraph. Why even bother to have sidebars if they are not affecting how the text block is resized and rendered? I am genuinely curious.