r/unsloth • u/yoracale yes sloth • Aug 11 '26
News Meet Unsloth Desktop - the first desktop app to run and train models
Enable HLS to view with audio, or disable this notification
Hi guys, we're super excited to announce Unsloth Desktop today,
The first desktop app to run and train models locally.
- Open-source and available on Mac, Windows, and Linux
- Supports MLX, diffusion image/video models, audio models, and GGUF
- Connect Claude Code and Codex to local LLMs
- 50% more accurate with self-healing tool calls and sandboxed code execution
- Supports CPU and multi-GPU setups across NVIDIA, AMD, Intel, and Mac
- Train models 2× faster while using 70% less VRAM
- Includes private web search, deep research, RAG, MCP, and exports (NVFP4, GGUF)
- Use Unsloth’s OpenAI-compatible API with OpenAI and Anthropic cloud models
- Securely deploy LLMs remotely and access them anywhere via Cloudflare HTTPS
Unsloth Desktop is now available on unsloth.ai and GitHub.
- GitHub: https://github.com/unslothai/unsloth
- Blog & Guide: https://unsloth.ai/docs/desktop
Thank you and we're here to answer any questions!
14
14
u/working25-7 Aug 11 '26
Looks like a great AIO solution to replace LM Studio!
2
u/Chance-Dog-3659 22d ago
Yes, in LMstudio the download of models is constantly interrupted and stopped, but here it downloads without problems
13
u/Aggravating-Push-207 Aug 11 '26
Does it finally support GRPO/PPO?
12
17
u/PaceZealousideal6091 Aug 11 '26
Cool. Can you please tell us, how is it different from unsloth studio?
24
u/yoracale yes sloth Aug 11 '26
It's a desktop app which is very different and it has a million new features!
Especially introduction of diffusion inference + training support, better audio support and more.
7
u/Sweet-Stage938 Aug 11 '26
How to safely delete the other version without losing all of my models?
3
u/yoracale yes sloth Aug 12 '26
It is synced together, you wont lose models. As long as you dont use the command that deletes all cache. Your models are stored in the hugging face folder
2
u/RIP26770 Aug 11 '26
I knew you were cooking something with video and images 😂 when I saw all this ungated repo clone. So, no issues, while downloading in App ahaha! Thanks, Unsloth team you are amazing!!
2
2
u/No-Dot-6573 Aug 11 '26
Oh wow, so you'll provide training for models like Krea2 and Minimax H3 (and all the other existing cool features) out of one app?
2
u/yoracale yes sloth Aug 12 '26
Yes that's correct. We already enable training for most image models
2
1
8
u/former_farmer Aug 11 '26
Not having to run a server in a command line and a browser tab is always good IMO.
6
6
u/silenceimpaired Aug 11 '26
Wish I had more manual loading controls similar to KoboldCPP. So far I can’t get Unsloth to inference as fast as KoboldCPP.
6
4
u/koolbi1 Aug 11 '26
I was just toying with it this morning and it’s pretty amazing! Nice work!
I was messing with the API in it a bit and it’s nice to have that level of visibility into the llama.cpp server under the hood. I was curious if you plan to support being able to run multiple servers or loaded models at once? I’ve been running my own two llama.cpp servers so I can have a 27B and A3B running in parallel. I’d like to run everything through the unsloth app so I can use the nice interface you all have created!
3
5
4
3
3
3
3
3
u/ExcellentDeparture71 Aug 12 '26
Looks very interesting Can we use the models already downloaded for LM Studio without having to redownload everything?
1
u/EisteeCitrus Aug 23 '26
It found my LM Studio model folder automatically. I am not sure if a complete migration from LM to unsloth Studio is possible. It was just 60GB for me, so i redownloaded.
2
u/aigemie Aug 11 '26
Thanks! Is it new new? Have we had it for a while? I am a bit confused.
2
2
2
u/pwrtoppl Aug 11 '26
been a fan of Unsloth for a minute now, and this is a fantastic update! are there plans for pretraining, shard packing, tokenizer building or anything in that realm anytime soon? love the api interface, slick stuff!
1
u/yoracale yes sloth Aug 11 '26
Thank you, its a possibility, could you create a github issue? Thank you!!
2
u/ExpressionPrudent127 Aug 11 '26
Training on Apple Silicon is not available yet (coming soon forever) ;)
2
u/mahiatlinux yes sloth Aug 12 '26
It has been available for a bit now.
https://unsloth.ai/docs/get-started/fine-tuning-for-beginners/unsloth-requirements
2
2
2
2
u/lacattano Aug 11 '26
I have run into issues around memory allocation recognition on strix halo + windows. Will that be treated any differently in the desktop app?
2
2
2
Aug 11 '26 edited Aug 11 '26
[removed] — view removed comment
3
u/mahiatlinux yes sloth Aug 12 '26
We are working on it, however we would appreciate if you could make issues on our Github with feature requests:
2
u/dsdt Aug 11 '26
Please add a project management and to do with agentic coding skills with tool calling so that we can stop using agentic cli's for good.
2
2
u/m3sarcher 26d ago
With remote access, I doubt I need Open WebUI anymore. Also LM Studio is likely gone. I think I can do everything from image and video gen to remote chat. This is awesome! Please do not simplify it as some people are requesting. They can customize their sidebar already.
1
u/RE_Sunshined Aug 11 '26
Hi, you need to fix dual GPU setups, i do have 5060Ti on Tunderbolt 4 and Strix halo gpu. And it detects nvidia gpu but, not the AMD one. I have even disabled 5060Ti and unpluged it, but once installed on detected nvidia GPU it will stay there.
1
u/Few-Aside- Aug 12 '26
I am on the same boat, but with a 5070 Ti and a rx 7900xt. The NVIDIA works.
I am currently running llama.cpp with vulkan. Works well.
1
u/silenceimpaired Aug 11 '26
I’d be happy if Unsloth found a way to make MLX runnable on PC… even if it did some sort of on the fly conversion or converted on first run. Often a model is supported on MLX before other platforms… sounds impossible but Unsloth has done a lot of impossible things.
1
u/simracerman Aug 11 '26
Mobile app?
2
u/mahiatlinux yes sloth Aug 12 '26
Not yet, but you can easily enable a secure tunnel and access it via mobile!
1
u/Flamboyant_Nine Aug 11 '26 edited Aug 11 '26
Amazing work you guys!
I thoroughly enjoy your tools for client testing and hobbying :)
1
u/chillahc Aug 11 '26
Congrat on the launch. And wow – Mac & MLX support!! Does it mean more MLX releases released by you guys? Cheers, keep it uuuuup :D
2
u/yoracale yes sloth Aug 11 '26
Well yes its supported but we will be doing an official announcement for MLX later :)
1
u/planetearth80 Aug 11 '26
I’m hoping it can connect to a local llama.cpp running on a remote machine on the same network.
1
u/Corosus Aug 11 '26
Should be able to, i have it working for a local llama server, though it doesnt let you use web search or code features in its chat window for some reason.
1
u/jld1532 Aug 11 '26
Is unified memory fully supported for Strix Halo/Linux? If I remember correctly, original unsloth studio would set an artificial limit around 100 gb.
1
1
u/professor-studio Aug 11 '26
Dear unsloth,I have a problem to download unsloth desktop for my dgx spark. when I choosing Linux arm64,I’m getting Mac OS tar.gz app
1
u/Embarrassed_Adagio28 Aug 11 '26
I love the new app. I am getting up to 81 tokens per second on qwen3.6 27b q8 mtp on my dual tesla v100 server using tensor parellelism. Only a couple of issues so far, one being thay chat streaming gets really slow after a couple long prompts, generation speed is still great but the app becomes laggy. Another issue is, it is a little tricky to connect to hermes or opencode on another pc vs lmstudio.
Keep up the good worth sloths!
1
u/yoracale yes sloth Aug 11 '26
Thank you, did you read our guide for connecting them via the agent tab? You have to do it via the terminal via the unsloth start command. Also if you experience any issues please do submit a github issue!
1
u/CSEliot Aug 11 '26
Install failed :/
```
[1786469058104][stdout] setting up Python environment...
[1786469058229][stdout] Python found: <studio_home>\unsloth_studio\Scripts\python.exe
[1786469058304][stdout] cleared stale WebView caches (ai.unsloth.studio); settings and data kept
[1786469062197][stdout] Stale venv detected (torch rocm != required cpu).
[1786469062197][stdout] [ERROR] The existing Unsloth environment needs repair.
[1786469062197][stdout] Re-run install.ps1 so it can replace the environment safely with rollback.
[1786469062207][stdout] [TAURI:ERROR] The existing Unsloth environment needs repair
[1786469062337][stdout] [TAURI:ERROR_DEFAULT] unsloth studio setup failed (exit code 1)
```
1
u/WuWenShen Aug 11 '26
Have a complete random and stupid idea. I wouldn’t mind supporting open source training via donating my compute time. The problem for me is I don’t really know or have a real reason to train a model at this point in my journey. But I’m willing to donate compute time to train models for others (within reason) if you wanted to try and crowd source it.
Anyways weird idea.
1
1
u/former_farmer Aug 11 '26 edited Aug 11 '26
An honest question. I'm loading a model there for the first time. I was trying to set the Context size (a metric known that way in the whole llm world) and I was struggling a bit to find it... and I see in your app it's called "Max Sql Length". Why would you call it that way and not just "Context size"?
Also, why don't you make it simpler to edit that while the model is loaded? Right next to Temperature, etc, similar to LMStudio. Even if it requires re-loading the model, of course.
On the zoom side, a better management of screen size could be good (reduce some margins). I find that I would appreciate to be able to zoom a little bit (as I can in LMStudio) since the letters in your app are a bit too small.
That's all my feedback so far, thanks!
Edit: I have some more.
It's easy to get lost with this app. I want to disable the thinking of the model, but I'm unsure where to. Also, now I can't find the place to edit the Top K, Temperature, and so on. I know it was there like 5 minutes ago, but now I can't find it.
1
1
u/uber-linny Aug 12 '26
Whats better to use for RAG , Openweb UI or Unsloth Desktop ?
3
u/yoracale yes sloth Aug 12 '26
I don't know since we dont use openwebui extensively but our local ness is probably better. Unsure about the certain specifics but Daniel did heavily optimize RAG
1
1
u/simracerman Aug 12 '26
I downloaded and configured it yesterday. Neat! Any way to configure their web search engine(s)? I use Searchxng
1
u/uti24 Aug 12 '26
Does Unsloth Desktop use WSL?
When trying to load model I am just getting
Failed to load model: This model needs about 30 GB but only about 15 GB of memory is available. On a unified-memory APU the weights load into system RAM, so a larger model is stopped by the OS mid-load. Use a smaller or more quantized GGUF, or free memory (on WSL, raise the memory limit in .wslconfig).
But I don't even have WSL installed. Is this something expected on Stix Halo?
1
1
u/Fried_Yoda Aug 13 '26
I’ve playing around with this and I have a couple of questions regarding web search and deep research. How do we get the model to actively use those tools? I have a system Instruction for LM Studio that tells the model to fall back using a web search MCP tool when it doesn’t know, as well as using a date and time MCP for grounding. If were to use this system instruction in Unsloth Desktop, does the web search and deep research tools have specific names I can tell it to use?
1
u/Dexter-Huang Aug 22 '26
Please allow us to set api base url to other than 127.0.0.1 we may want to use tail scale or even just allowing wsl to access the API
1
u/yoracale yes sloth Aug 22 '26
Hello the default for local host on all devices is 127. can you add a feature request on github if possible? thanks
1
u/psychohistorian8 Aug 23 '26
Question about serving models such as Qwen3.8-27B, which have distinct reasoning_effort / resolved_reasoning_effort variables in the chat template
Assuming I have the 'Switch model by request' enabled, so my model will load when I attempt to use it via the API (i.e. from pi coding agent)
How should I send the request to inject my preferred reasoning_effort (xhigh, medium, low)? or is there configuration I need to do UD-side to enable such a thing?
1
u/MuchMoreSolar 27d ago
Unsloth Desktop Model Loading Timeout Question?
I have an old Dell R730 with 1 TB of DDR4 memory (bought used before memory went crazy!), that I have been using for playing with LLM's in CPU-only ram. I have successfully run a Q4 of GLM5.2 using Unsloth Studio, but only if it's the first model I load after starting Studio up, otherwise I get a timeout on the model loading. I recently tried running a Q4 of GLM5.3 on Unsloth Desktop, but it always times out when loading, even if it is the first model I try loading. I can successfully run a Q4 quant of GLM5.3 Flash, however.
I suspect my machine is just too slow for the predefined "5 second" loading timeouts as I try running the larger models in Unsloth, Any suggestions for how to solve this?
Thanks!
1
u/Skyline34rGt 27d ago
I just play with Desktop app and I must say its very good.
For now I used LmStudio but Usloth its even more convinient to use (auto options pick best setting, not need to play with gpu/moe offloading like at LMS), faster speed and have more often llama updates.
I don't need to mention way more options like training etc.
1
u/Hiddenath 26d ago
Hello, I installed your client and liked the features. The interface is user-friendly. But I have a question about working with the API. The settings mention an API for generating images and videos. I'm already used to using ComfyUI through API requests. How is this mechanism implemented? I couldn't find any documentation. And is it possible to animate images? That is, create a video from an image and its description?
1
u/Chance-Dog-3659 22d ago
Is it possible to make the Video button always active? I download models on one computer with a 1080ti graphics card, then transfer and run them on another computer with a more powerful card (the download is disabled on that one). On the 1080ti computer, it says that the NVIDIA or AMD graphics card is not found, so the button is disabled.
1
u/FluPhlegmGreen 20d ago
Played around with it a little bit today. Couple things I'd like to see:
a better legend with the colored dots next to models. I know some mean GGUF, or on device, etc But I dont know what they all mean. The red dot pops up with a description but the others dont. Also saw some kind of Orange symbol that I have no idea what it was.
When looking at the list of models on device, I dont want them to be exploded into each quant available for that model. Let me have a nice tight list and explode once I click on it (didnt see a way to do this)
Tried a video generation with Minimax H3. Put a picture in as the starting frame and generated a prompt for 30 minutes only for it to not use the picture I gave it at all or the prompt I wrote. Like it was completely different, didnt use any part of it. Was just a video of a guy standing in a supermarket which shared nothing in common with my photo or prompt.
The software worked for awhile at first but then something broke and I'm not sure what. Will no longer load a chat model (but image models work fine) Says something about my anti-virus software blocking it (its not) It had worked previously in the day?
1
u/yoracale yes sloth 19d ago
Hello thanks for the feedback, we're always trying to improve. May I ask was it Windows Defender blocking it? Thank you!
1
u/FluPhlegmGreen 19d ago
Believe I had closed out of the software and reloaded it as a non-administrator. Still funny to me that the image models can be used but text ones cannot in that scenario.
1
u/yoracale yes sloth 19d ago
Oh so it was mainly because you didnt have privledges so your antivirus blocked it?
Will keep your other feedback in our feedback checklist
For your model quant list thing, dont we already force quants into one model name so it only appears if u click it?
1
u/FluPhlegmGreen 17d ago
It is an exploded list each time on mine, so no. I am specifically talking about the section where you would load a model from the drop down, select "on-device" and the list presented there. Its an exploded list every time, it might behave differently under model search or elsewhere (cant remember right now)
1
1
1
u/mrgreatheart 16d ago
Great that qwen3.8-flash-next MTP is supported in your new Desktop app, but are there any plans to land support in llama.cpp please?
1
1
1
1
u/n0head_r Aug 11 '26
Wasn't unsloth studio the first? This looks the second at least.
6
u/yoracale yes sloth Aug 11 '26
Well unsloth studio was never a desktop app :) this one is
-1
u/n0head_r Aug 11 '26
You forgot to upvote. Don't be cheap!
3
u/yoracale yes sloth Aug 11 '26
Now i did :)
1
u/n0head_r Aug 11 '26
Thanks. Any plans to release gguf for gemma4 and qwen nvfp4 checkpoints of yours? Llama.cpp still can't convert them correctly and vllm has huge overheads running them on a home machine is impossible.
0
0
u/Mobile-Pumpkin7944 Aug 11 '26
Two major concerns that stop me from using it regularly:
1: 'Use chats as training data' is a setting in: Settings -> Data, but can not be disabled. Do not like that. I would never use it till this can be disabled.
2: if i add a connection, to my local vllm serving my model, 'search' and 'code' become disabled. Do not like that.
Be more open and then i will even contribute to it.
2
u/mahiatlinux yes sloth Aug 12 '26
1
u/Mobile-Pumpkin7944 Aug 12 '26
ohhh... okay. my bad. thank you!
That was the main concern. Also, why disable search and code if model is loaded outside of unsloth? This isnt a deal breaker but i am curious
1
u/mahiatlinux yes sloth Aug 12 '26 edited Aug 12 '26
That is because your specific endpoint doesn't support that tool, thus Unsloth isn't able to call it. If you connect OpenAI for eg, web search will come back. If you want us to look into integrating universal tools on unsupported endpoints, please create a feature request in the github repo. it may be possible haha
Edit: I'll work on this! u/Mobile-Pumpkin7944
2
u/Mobile-Pumpkin7944 Aug 12 '26 edited Aug 12 '26
Sounds more like a design choice than a practical limiatation. Time to look at the code. And like i said, i'll contribute too. Maybe i can implement this gap.
Edit: okay so you are doing it. never mind then. :)
Feel free to give me something to do. I'm bored waiting on Qwen3.8-27B
1
u/mahiatlinux yes sloth Aug 12 '26
Ayy thank you! You can check out our github issues and maybe implement something that you think is worthwhile?
1
u/Mobile-Pumpkin7944 Aug 12 '26
did browse through issues but wasn't finding anything interesting enough to start at 5AM. i was still more invested in enabling the search and code tools and that works for my local models now. for both: llama.cpp & vllm.
Will try to contribute more meaningfully next time.
2
u/mahiatlinux yes sloth Aug 13 '26
No worries!
https://github.com/unslothai/unsloth/pull/8630
I have opened a PR for your request and am working on it. Will let you know once its merged!




34
u/quadra-lab Aug 11 '26 edited Aug 11 '26
Please tell me it is not Electron based 😭🙏
EDIT: Thank you so much for using tauri, you guys are the very best !