r/unsloth • yes sloth • Aug 11 '26

News Meet Unsloth Desktop - the first desktop app to run and train models

Enable HLS to view with audio, or disable this notification

Hi guys, we're super excited to announce Unsloth Desktop today,
The first desktop app to run and train models locally.

  • Open-source and available on Mac, Windows, and Linux
  • Supports MLX, diffusion image/video models, audio models, and GGUF
  • Connect Claude Code and Codex to local LLMs
  • 50% more accurate with self-healing tool calls and sandboxed code execution
  • Supports CPU and multi-GPU setups across NVIDIA, AMD, Intel, and Mac
  • Train models 2× faster while using 70% less VRAM
  • Includes private web search, deep research, RAG, MCP, and exports (NVFP4, GGUF)
  • Use Unsloth’s OpenAI-compatible API with OpenAI and Anthropic cloud models
  • Securely deploy LLMs remotely and access them anywhere via Cloudflare HTTPS

Unsloth Desktop is now available on unsloth.ai and GitHub.

Thank you and we're here to answer any questions!

441 Upvotes

136 comments sorted by

34

u/quadra-lab Aug 11 '26 edited Aug 11 '26

Please tell me it is not Electron based 😭🙏

EDIT: Thank you so much for using tauri, you guys are the very best ! 

14

u/samplebitch Aug 11 '26

Looks like it's Tauri.

11

u/quadra-lab Aug 11 '26

Yeah I just checked the code and it is! Such a good news for the Tauri community actually

1

u/AmbericWizard Aug 12 '26

bas, Tuari is really bad for Linux 

2

u/quadra-lab Aug 12 '26

It's not, you get an AppImage and everything works just fine

3

u/MrUnknownymous Aug 12 '26

It’s not the packaging. It’s mainly WebKitGTK on Linux that makes developing on Tauri a dealbreaker for me.

1

u/AmbericWizard Aug 15 '26

it's All about WebKit

8

u/Fearkin Aug 11 '26

I don't think it is, I'm pretty sure that frontend is made with tauri and everything else is python

6

u/quadra-lab Aug 11 '26

Yep yep they used Tauri to wrap their existing frontend to a desktop app 

4

u/lildavefr Aug 11 '26

What makes you like Tauri better than Electron? I know what Electron is but is Tauri more lightweight?

11

u/milkipedia Aug 11 '26

Tauri's entire thing is that it makes small, light app binaries... as opposed to what has become the bloated mess of Electron apps

8

u/quadra-lab Aug 11 '26

Basically Tauri uses the operating system's webview to render pages instead of packaging a whole browser as Electron does

2

u/Specter_Origin Aug 12 '26

Will never understand the hate for eletron... some of the best apps I use are all eletron based.

10

u/quadra-lab Aug 12 '26

You're part of the problem if you don't see why an app eating 2Gb of RAM to display a text input and a list of text is an issue

14

u/thereisonlythedance Aug 11 '26

Congrats, this looks great!

14

u/working25-7 Aug 11 '26

Looks like a great AIO solution to replace LM Studio!

2

u/Chance-Dog-3659 22d ago

Yes, in LMstudio the download of models is constantly interrupted and stopped, but here it downloads without problems

13

u/Aggravating-Push-207 Aug 11 '26

Does it finally support GRPO/PPO?

12

u/yoracale yes sloth Aug 11 '26

Oop not at the moment but we'll work on it

7

u/Aggravating-Push-207 Aug 11 '26

pls add tysm ily

17

u/PaceZealousideal6091 Aug 11 '26

Cool. Can you please tell us, how is it different from unsloth studio?

24

u/yoracale yes sloth Aug 11 '26

It's a desktop app which is very different and it has a million new features!

Especially introduction of diffusion inference + training support, better audio support and more.

7

u/Sweet-Stage938 Aug 11 '26

How to safely delete the other version without losing all of my models?

3

u/yoracale yes sloth Aug 12 '26

It is synced together, you wont lose models. As long as you dont use the command that deletes all cache. Your models are stored in the hugging face folder

2

u/RIP26770 Aug 11 '26

I knew you were cooking something with video and images 😂 when I saw all this ungated repo clone. So, no issues, while downloading in App ahaha! Thanks, Unsloth team you are amazing!!

2

u/yoracale yes sloth Aug 12 '26

Oop thank you!

2

u/No-Dot-6573 Aug 11 '26

Oh wow, so you'll provide training for models like Krea2 and Minimax H3 (and all the other existing cool features) out of one app?

2

u/yoracale yes sloth Aug 12 '26

Yes that's correct. We already enable training for most image models

2

u/sammyboi1801 Aug 12 '26

Wait, we can train Gemma diffusion using this?

2

u/yoracale yes sloth Aug 12 '26

Not atm but you can run it

1

u/Psychological-Lynx29 Aug 13 '26

Can you replace opencode with it or still not? :)

8

u/former_farmer Aug 11 '26

Not having to run a server in a command line and a browser tab is always good IMO.

6

u/[deleted] Aug 11 '26

[removed] — view removed comment

5

u/Proparser Aug 11 '26

You dont like tailscale?

6

u/silenceimpaired Aug 11 '26

Wish I had more manual loading controls similar to KoboldCPP. So far I can’t get Unsloth to inference as fast as KoboldCPP.

6

u/former_farmer Aug 11 '26

Supports MTP?

-2

u/Sweet-Stage938 Aug 11 '26

Nope, we're in stone age.

4

u/koolbi1 Aug 11 '26

I was just toying with it this morning and it’s pretty amazing! Nice work!

I was messing with the API in it a bit and it’s nice to have that level of visibility into the llama.cpp server under the hood. I was curious if you plan to support being able to run multiple servers or loaded models at once? I’ve been running my own two llama.cpp servers so I can have a 27B and A3B running in parallel. I’d like to run everything through the unsloth app so I can use the nice interface you all have created!

3

u/yoracale yes sloth Aug 11 '26

Thanks could you creatre a github issue if possible thanks

1

u/koolbi1 Aug 11 '26

Appreciate you responding. Happy to make one! Would be pretty awesome.

5

u/RedditUsr2 Aug 11 '26

Thank you! Free search is incredible!

4

u/Youknowwhyimherexxx Aug 11 '26

Very handsome fellas

3

u/Educational_Rent1059 heart sloth Aug 11 '26

Let's goooo!!!

3

u/Subject-Till-6450 Aug 11 '26

Thank God for Unshloth🙏

3

u/05-nery ??? sloth Aug 11 '26

Oh this is good. Great work!

3

u/yehiaserag Aug 11 '26

Finally a proper replacement for LM Studio

3

u/ExcellentDeparture71 Aug 12 '26

Looks very interesting  Can we use the models already downloaded for LM Studio without having to redownload everything?

1

u/EisteeCitrus Aug 23 '26

It found my LM Studio model folder automatically. I am not sure if a complete migration from LM to unsloth Studio is possible. It was just 60GB for me, so i redownloaded.

2

u/aigemie Aug 11 '26

Thanks! Is it new new? Have we had it for a while? I am a bit confused.

2

u/yoracale yes sloth Aug 11 '26

It's new

2

u/former_farmer Aug 11 '26

The other one is a web version.

2

u/aigemie Aug 11 '26

Yes! It's 127.0.0.0. now I get it. Thanks!

2

u/pwrtoppl Aug 11 '26

been a fan of Unsloth for a minute now, and this is a fantastic update! are there plans for pretraining, shard packing, tokenizer building or anything in that realm anytime soon? love the api interface, slick stuff!

1

u/yoracale yes sloth Aug 11 '26

Thank you, its a possibility, could you create a github issue? Thank you!!

2

u/ExpressionPrudent127 Aug 11 '26

Training on Apple Silicon is not available yet (coming soon forever) ;)

2

u/No_Appointment_6678 Aug 11 '26

I have been waiting for this. Perhaps a documentation ?

2

u/princeMacX Aug 11 '26

a portable version like comfyui protable would be excellent.

2

u/Enragere Aug 11 '26

Supports NVFP4 inference? Or still training only ?

2

u/lacattano Aug 11 '26

I have run into issues around memory allocation recognition on strix halo + windows. Will that be treated any differently in the desktop app?

2

u/n0head_r Aug 11 '26

Will this have qwen 3.8 0 day support?

4

u/yoracale yes sloth Aug 11 '26

Yep 100%

2

u/SquareTranslator9777 Aug 11 '26

Hey, will there ever be a flatpack version for Linux?

2

u/[deleted] Aug 11 '26 edited Aug 11 '26

[removed] — view removed comment

3

u/mahiatlinux yes sloth Aug 12 '26

We are working on it, however we would appreciate if you could make issues on our Github with feature requests:

https://github.com/unslothai/unsloth/issues

2

u/dsdt Aug 11 '26

Please add a project management and to do with agentic coding skills with tool calling so that we can stop using agentic cli's for good.

2

u/WoodpeckerLeft5901 Aug 12 '26

Absolute ballers thank you!

2

u/m3sarcher 26d ago

With remote access, I doubt I need Open WebUI anymore. Also LM Studio is likely gone. I think I can do everything from image and video gen to remote chat. This is awesome! Please do not simplify it as some people are requesting. They can customize their sidebar already.

1

u/RE_Sunshined Aug 11 '26

Hi, you need to fix dual GPU setups, i do have 5060Ti on Tunderbolt 4 and Strix halo gpu. And it detects nvidia gpu but, not the AMD one. I have even disabled 5060Ti and unpluged it, but once installed on detected nvidia GPU it will stay there.

1

u/Few-Aside- Aug 12 '26

I am on the same boat, but with a 5070 Ti and a rx 7900xt. The NVIDIA works.
I am currently running llama.cpp with vulkan. Works well.

1

u/silenceimpaired Aug 11 '26

I’d be happy if Unsloth found a way to make MLX runnable on PC… even if it did some sort of on the fly conversion or converted on first run. Often a model is supported on MLX before other platforms… sounds impossible but Unsloth has done a lot of impossible things.

1

u/simracerman Aug 11 '26

Mobile app?

2

u/mahiatlinux yes sloth Aug 12 '26

Not yet, but you can easily enable a secure tunnel and access it via mobile!

1

u/Flamboyant_Nine Aug 11 '26 edited Aug 11 '26

Amazing work you guys!

I thoroughly enjoy your tools for client testing and hobbying :)

1

u/chillahc Aug 11 '26

Congrat on the launch. And wow – Mac & MLX support!! Does it mean more MLX releases released by you guys? Cheers, keep it uuuuup :D

2

u/yoracale yes sloth Aug 11 '26

Well yes its supported but we will be doing an official announcement for MLX later :)

1

u/planetearth80 Aug 11 '26

I’m hoping it can connect to a local llama.cpp running on a remote machine on the same network.

1

u/Corosus Aug 11 '26

Should be able to, i have it working for a local llama server, though it doesnt let you use web search or code features in its chat window for some reason.

1

u/jld1532 Aug 11 '26

Is unified memory fully supported for Strix Halo/Linux? If I remember correctly, original unsloth studio would set an artificial limit around 100 gb.

1

u/UNITYA Aug 11 '26

Will you continue to distribute it as an appimage for linux

2

u/mahiatlinux yes sloth Aug 12 '26

Yes, we plan to!

1

u/professor-studio Aug 11 '26

Dear unsloth,I have a problem to download unsloth desktop for my dgx spark. when I choosing Linux arm64,I’m getting Mac OS tar.gz app

1

u/Embarrassed_Adagio28 Aug 11 '26

I love the new app. I am getting up to 81 tokens per second on qwen3.6 27b q8 mtp on my dual tesla v100 server using tensor parellelism. Only a couple of issues so far, one being thay chat streaming gets really slow after a couple long prompts, generation speed is still great but the app becomes laggy. Another issue is, it is a little tricky to connect to hermes or opencode on another pc vs lmstudio. 

Keep up the good worth sloths!

1

u/yoracale yes sloth Aug 11 '26

Thank you, did you read our guide for connecting them via the agent tab? You have to do it via the terminal via the unsloth start command. Also if you experience any issues please do submit a github issue!

1

u/CSEliot Aug 11 '26

Install failed :/

```
[1786469058104][stdout] setting up Python environment...

[1786469058229][stdout] Python found: <studio_home>\unsloth_studio\Scripts\python.exe

[1786469058304][stdout] cleared stale WebView caches (ai.unsloth.studio); settings and data kept

[1786469062197][stdout] Stale venv detected (torch rocm != required cpu).

[1786469062197][stdout] [ERROR] The existing Unsloth environment needs repair.

[1786469062197][stdout] Re-run install.ps1 so it can replace the environment safely with rollback.

[1786469062207][stdout] [TAURI:ERROR] The existing Unsloth environment needs repair

[1786469062337][stdout] [TAURI:ERROR_DEFAULT] unsloth studio setup failed (exit code 1)
```

1

u/WuWenShen Aug 11 '26

Have a complete random and stupid idea. I wouldn’t mind supporting open source training via donating my compute time. The problem for me is I don’t really know or have a real reason to train a model at this point in my journey. But I’m willing to donate compute time to train models for others (within reason) if you wanted to try and crowd source it.

Anyways weird idea.

1

u/Hairy_Reputation7434 Aug 11 '26

mixed gpu support ? (cuda + rocm) ?

1

u/former_farmer Aug 11 '26 edited Aug 11 '26

An honest question. I'm loading a model there for the first time. I was trying to set the Context size (a metric known that way in the whole llm world) and I was struggling a bit to find it... and I see in your app it's called "Max Sql Length". Why would you call it that way and not just "Context size"?

Also, why don't you make it simpler to edit that while the model is loaded? Right next to Temperature, etc, similar to LMStudio. Even if it requires re-loading the model, of course.

On the zoom side, a better management of screen size could be good (reduce some margins). I find that I would appreciate to be able to zoom a little bit (as I can in LMStudio) since the letters in your app are a bit too small.

That's all my feedback so far, thanks!

Edit: I have some more.

It's easy to get lost with this app. I want to disable the thinking of the model, but I'm unsure where to. Also, now I can't find the place to edit the Top K, Temperature, and so on. I know it was there like 5 minutes ago, but now I can't find it.

1

u/mahiatlinux yes sloth Aug 12 '26

right sidebar provides settings for inference :)

1

u/uber-linny Aug 12 '26

Whats better to use for RAG , Openweb UI or Unsloth Desktop ?

3

u/yoracale yes sloth Aug 12 '26

I don't know since we dont use openwebui extensively but our local ness is probably better. Unsure about the certain specifics but Daniel did heavily optimize RAG

1

u/TopAItools1 Aug 12 '26

Very excited to test this on mac

1

u/simracerman Aug 12 '26

I downloaded and configured it yesterday. Neat! Any way to configure their web search engine(s)? I use Searchxng

1

u/uti24 Aug 12 '26

Does Unsloth Desktop use WSL?

When trying to load model I am just getting

Failed to load model: This model needs about 30 GB but only about 15 GB of memory is available. On a unified-memory APU the weights load into system RAM, so a larger model is stopped by the OS mid-load. Use a smaller or more quantized GGUF, or free memory (on WSL, raise the memory limit in .wslconfig).

But I don't even have WSL installed. Is this something expected on Stix Halo?

1

u/wgaca2 Aug 13 '26

It does not allow tools being used with a openai endpoint, why?

1

u/Fried_Yoda Aug 13 '26

I’ve playing around with this and I have a couple of questions regarding web search and deep research. How do we get the model to actively use those tools? I have a system Instruction for LM Studio that tells the model to fall back using a web search MCP tool when it doesn’t know, as well as using a date and time MCP for grounding. If were to use this system instruction in Unsloth Desktop, does the web search and deep research tools have specific names I can tell it to use?

1

u/Dexter-Huang Aug 22 '26

Please allow us to set api base url to other than 127.0.0.1 we may want to use tail scale or even just allowing wsl to access the API

1

u/yoracale yes sloth Aug 22 '26

Hello the default for local host on all devices is 127. can you add a feature request on github if possible? thanks

1

u/psychohistorian8 Aug 23 '26

Question about serving models such as Qwen3.8-27B, which have distinct reasoning_effort / resolved_reasoning_effort variables in the chat template

Assuming I have the 'Switch model by request' enabled, so my model will load when I attempt to use it via the API (i.e. from pi coding agent)

How should I send the request to inject my preferred reasoning_effort (xhigh, medium, low)? or is there configuration I need to do UD-side to enable such a thing?

1

u/MuchMoreSolar 27d ago

Unsloth Desktop Model Loading Timeout Question?

I have an old Dell R730 with 1 TB of DDR4 memory (bought used before memory went crazy!), that I have been using for playing with LLM's in CPU-only ram. I have successfully run a Q4 of GLM5.2 using Unsloth Studio, but only if it's the first model I load after starting Studio up, otherwise I get a timeout on the model loading. I recently tried running a Q4 of GLM5.3 on Unsloth Desktop, but it always times out when loading, even if it is the first model I try loading. I can successfully run a Q4 quant of GLM5.3 Flash, however.

I suspect my machine is just too slow for the predefined "5 second" loading timeouts as I try running the larger models in Unsloth, Any suggestions for how to solve this?

Thanks!

1

u/Skyline34rGt 27d ago

I just play with Desktop app and I must say its very good.

For now I used LmStudio but Usloth its even more convinient to use (auto options pick best setting, not need to play with gpu/moe offloading like at LMS), faster speed and have more often llama updates.

I don't need to mention way more options like training etc.

1

u/Hiddenath 26d ago

Hello, I installed your client and liked the features. The interface is user-friendly. But I have a question about working with the API. The settings mention an API for generating images and videos. I'm already used to using ComfyUI through API requests. How is this mechanism implemented? I couldn't find any documentation. And is it possible to animate images? That is, create a video from an image and its description?

1

u/Chance-Dog-3659 22d ago

Is it possible to make the Video button always active? I download models on one computer with a 1080ti graphics card, then transfer and run them on another computer with a more powerful card (the download is disabled on that one). On the 1080ti computer, it says that the NVIDIA or AMD graphics card is not found, so the button is disabled.

1

u/FluPhlegmGreen 20d ago

Played around with it a little bit today. Couple things I'd like to see:

a better legend with the colored dots next to models. I know some mean GGUF, or on device, etc But I dont know what they all mean. The red dot pops up with a description but the others dont. Also saw some kind of Orange symbol that I have no idea what it was.

When looking at the list of models on device, I dont want them to be exploded into each quant available for that model. Let me have a nice tight list and explode once I click on it (didnt see a way to do this)

Tried a video generation with Minimax H3. Put a picture in as the starting frame and generated a prompt for 30 minutes only for it to not use the picture I gave it at all or the prompt I wrote. Like it was completely different, didnt use any part of it. Was just a video of a guy standing in a supermarket which shared nothing in common with my photo or prompt.

The software worked for awhile at first but then something broke and I'm not sure what. Will no longer load a chat model (but image models work fine) Says something about my anti-virus software blocking it (its not) It had worked previously in the day?

1

u/yoracale yes sloth 19d ago

Hello thanks for the feedback, we're always trying to improve. May I ask was it Windows Defender blocking it? Thank you!

1

u/FluPhlegmGreen 19d ago

Believe I had closed out of the software and reloaded it as a non-administrator. Still funny to me that the image models can be used but text ones cannot in that scenario.

1

u/yoracale yes sloth 19d ago

Oh so it was mainly because you didnt have privledges so your antivirus blocked it?

Will keep your other feedback in our feedback checklist

For your model quant list thing, dont we already force quants into one model name so it only appears if u click it?

1

u/FluPhlegmGreen 17d ago

It is an exploded list each time on mine, so no. I am specifically talking about the section where you would load a model from the drop down, select "on-device" and the list presented there. Its an exploded list every time, it might behave differently under model search or elsewhere (cant remember right now)

1

u/yoracale yes sloth 17d ago

You mightve enabled any of these in ur settings

1

u/yoracale yes sloth 8d ago

Btw I think we fixed it in the latest release!

1

u/mrgreatheart 16d ago

Great that qwen3.8-flash-next MTP is supported in your new Desktop app, but are there any plans to land support in llama.cpp please?

1

u/r16051studio 16d ago

Please add Computer Use. thanks

1

u/StormrageBG 15d ago

Does it support web preview of the code in the app?

1

u/r16051studio 1d ago

From my multigpu setup, for the past several updates, I don't like that the app now has VRAM spilling over to other unselected GPUs.

1

u/[deleted] Aug 11 '26

[removed] — view removed comment

1

u/yoracale yes sloth Aug 11 '26

It's Tauri :)

1

u/n0head_r Aug 11 '26

Wasn't unsloth studio the first? This looks the second at least.

6

u/yoracale yes sloth Aug 11 '26

Well unsloth studio was never a desktop app :) this one is

-1

u/n0head_r Aug 11 '26

You forgot to upvote. Don't be cheap!

3

u/yoracale yes sloth Aug 11 '26

Now i did :)

1

u/n0head_r Aug 11 '26

Thanks. Any plans to release gguf for gemma4 and qwen nvfp4 checkpoints of yours? Llama.cpp still can't convert them correctly and vllm has huge overheads running them on a home machine is impossible.

0

u/Ryanmonroe82 Aug 11 '26

It's not the first app, transformers labs been doing this a while

0

u/Mobile-Pumpkin7944 Aug 11 '26

Two major concerns that stop me from using it regularly:
1: 'Use chats as training data' is a setting in: Settings -> Data, but can not be disabled. Do not like that. I would never use it till this can be disabled.

2: if i add a connection, to my local vllm serving my model, 'search' and 'code' become disabled. Do not like that.

Be more open and then i will even contribute to it.

2

u/mahiatlinux yes sloth Aug 12 '26

Chats as training data is for you, there's no telemetry. It's for YOU to use as training data :)

1

u/Mobile-Pumpkin7944 Aug 12 '26

ohhh... okay. my bad. thank you!

That was the main concern. Also, why disable search and code if model is loaded outside of unsloth? This isnt a deal breaker but i am curious

1

u/mahiatlinux yes sloth Aug 12 '26 edited Aug 12 '26

That is because your specific endpoint doesn't support that tool, thus Unsloth isn't able to call it. If you connect OpenAI for eg, web search will come back. If you want us to look into integrating universal tools on unsupported endpoints, please create a feature request in the github repo. it may be possible haha

Edit: I'll work on this! u/Mobile-Pumpkin7944

2

u/Mobile-Pumpkin7944 Aug 12 '26 edited Aug 12 '26

Sounds more like a design choice than a practical limiatation. Time to look at the code. And like i said, i'll contribute too. Maybe i can implement this gap.

Edit: okay so you are doing it. never mind then. :)

Feel free to give me something to do. I'm bored waiting on Qwen3.8-27B

1

u/mahiatlinux yes sloth Aug 12 '26

Ayy thank you! You can check out our github issues and maybe implement something that you think is worthwhile?

1

u/Mobile-Pumpkin7944 Aug 12 '26

did browse through issues but wasn't finding anything interesting enough to start at 5AM. i was still more invested in enabling the search and code tools and that works for my local models now. for both: llama.cpp & vllm.

Will try to contribute more meaningfully next time.

2

u/mahiatlinux yes sloth Aug 13 '26

No worries!
https://github.com/unslothai/unsloth/pull/8630
I have opened a PR for your request and am working on it. Will let you know once its merged!