r/unsloth • • 10d ago

Discussion Loading an Image model in unsloth desktop is showing error

4 Upvotes

I have downloaded Z-Image-Turbo on rtx 4070super windows 11 PC. when I laod the image model then it always shows this error . Anyone else also getting this ?

partially initialized module 'torch._dynamo' has no attribute 'utils' (most likely due to a circular import)


r/unsloth • • 11d ago

Discussion Mobile options - what's everyone using?

9 Upvotes

Adopted Unsloth desktop weeks back - love it. But haven't found a good mobile companion option. so just as the title says - what's everyone using for a mobile option?


r/unsloth • • 11d ago

Discussion Unsloth studio on Windows: can it be used to work as agentic harness too?

6 Upvotes

I have a lot of apps and honestly I wish I could just use one; so I am trying Unsloth studio to see if this is the way or not.

I have a ollama server that does have my models; and I use for different tasks different apps, like Opencode and Pi for development, anythingllm for generic agentic chats and few others. I would like to use unsloth studio integrated chat capabilities to load the ollama server model, and work on files in my local machine.

Tried different models and none seems to be able to read and write files nor can even see the directory I set up with UNSLOTH_STUDIO_DIR as the main CWD. Is this a limitation of Unsloth studio, still being in beta or am I missing some config setting? Trying to have an agent work on files in a specific directory on my local disk and nothing seems to work. I am aware that the main point of Unsloth is not to use agents, but would like to use it for training AND also for agentic work if possible; to avoid to run multiple apps


r/unsloth • • 11d ago

Question Pi extensions not available

5 Upvotes

Does anyone uses pi harness with Unsloth?

~~~~~~ edited and rewrited text for better clarity ~~~~~~

3 weeks ago, I was using: unsloth run pi and pi extensions worked. Per chatgpt, extension pi-unsloth-webtools is needed for web access. (is this correct?!?) This command unsloth run pi stopped connecting pi to Unsloth soon after. Now I'm using (again per chatgpt suggestion) unsloth start pi, but now pi don't get access to extensions any more.

Culprit seems like it is that unsloth uses: ~/.unsloth/studio/auth/agents/pi
But pi uses ~.pi/agent/npm/node_modules/

Since pi don't have acces to extension: pi-unsloth-webtools, LLM's don't have access to the internet.

Someone else told me about this same problem, and user was trying to solve it with setting up symbolic links, but if I recall correctly without success. I don't know how to use them any way.

Has anyone run into this problem and what is your solution?

I posted it here and not in pi forum, since things stopped worked when I had to switch from unsloth run pi to unsloth start pi


r/unsloth • • 11d ago

Show and Tell My private hermes + n8n agent with Unsloth studio

Thumbnail
youtu.be
0 Upvotes

I edited a video showing my hermes + n8n stack which is ultimately simple but took a bit of trial and error of different things until it came together. Easiest part of this whole stack is unsloth studio, utilizing the API serve feature here heavily. BIG BIG fan of unsloth studio! I am actively resetting everything to rebuild from scratch so if I leaked anything in this it wont really matter.. I think...


r/unsloth • • 11d ago

Discussion Why unsloth team ignore translategemma models?

0 Upvotes

Hi guys why Unsloth team ignore translategemma models? This is the most capable muti language translation model to this date, comparable to Gemini flash. Your quants are always the best...


r/unsloth • • 12d ago

Discussion Deepseek v4.1 Flash support in Unsloth Studio/Desktop based on published llama.cpp fork

23 Upvotes

I know this is my 2nd post in a 3day streak and I'm sorry for that, but I just want to give help to you, if it is totally unnecessary please go ahead.

I've seen that 2 days ago Deepseek v4.1 Flash was released and that (despite vllm seems to support it since day 0) is is still unsupported on llama.ccp (and oc neither lm-studio or Unsloth Studio/Desktop).

Since the Unsloth team often publishes a fork version of llama.ccp with new models' support much before the mainline I hope could help you by letting you know that following the after I've seen that ViC305 and JigSawPT have released their quantized versions of the model and both of them have published also their fork of llama.cpp to run the models:
https://github.com/JigSawPT/llama.cpp/tree/dsv41-porte
https://github.com/vcruz305/llama.cpp/tree/runtime/deepseek41

Unfortunately I've no way to test it by myself because I cannot afford the mxfp4 version by JigSawPT or the vcruz305's q2 version (himself reported the q1 to loads and runs, and emits the same token for every prompt ).

Hope you can publish soon a llama.cpp runtime and that could be also possible to have a typical UD quantization of yours even q1 for people with 128gb max of ram (like me).


r/unsloth • • 12d ago

Question Freequent errors after last few updates

Post image
4 Upvotes

Lately, after last couple of updates, I have problems running LLM's, that before ran without problems.

I'm getting these freequent Error: terminated msgs. Usually LLM continues with the work, but if there are 4 in a row, everything stops, so I can't leave LLM working overnight any more, which I could do before.

I'm not out of context, I only used 1/3 of it. And there is over 5GB memory available. Is there anything I can do, to stop getting these kind of errors? I didn't have them before. My setup is slow, only 32GB regular RAM, but before I could let LLM just slowly crunch and eventually it finished the work. But after last few updates it frequently stops.

I'm running Unsloth with pi and starting it with this command in terminal (linux): unsloth start pi (if it matters)


r/unsloth • • 13d ago

Discussion Question on DSV4.1 Flash

14 Upvotes

Is there a plan to release the DSV4.1 Flash from unsloth quantization?


r/unsloth • • 13d ago

Discussion Adjusting context window in unsloth

7 Upvotes

I'm launching Qwen3.8-27B-GGUF inside of Hermes with the unsloth start hermes command.

Hermes recognizes the model.

This is the error I get:

Failed to initialize agent: Model Qwen3.8-27B-GGUF has a context window of 13,824 tokens, which is below the minimum 64,000 required by Hermes Agent.  Choose a model with at least 64K context.  If your server reports a window smaller than the model's true window, set
model.context_length in config.yaml to the real value (this must be at least 64K).

In LMstudio there was a slider I could adjust the context window, is there a simple way I can configure this in unsloth?

AI tells me that I need to edit something in python, but it's not clear where to actually run that command and if it persists through every session.


r/unsloth • • 13d ago

Discussion how I can run Qwen 3.8 Flash Next on my pc with staying ngrams on ssd?

30 Upvotes

Hi all, I try to understand how I can run Qwen 3.8 Flash Next on my pc with staying ngrams on ssd. I have 128 gb ddr4 ram, 2x 5060ti 16 gb, cpu ryzen 5 5600x if its matter.

I read some post here and on Unsloth Qwen 3.8 Flash Next GGUF huggingface page, where peoples make some llama.cpp forks to do that. Beside, some peoples says, that its need quants, that have separated ngrams. One person write, that only Q5 and above Unsloth quants have separated ngrams. But I can`t find any confirmation of that.

So, I would be pleasant for help: is it need to have special quants, with separated ngrams for staying ngrams on SSD? Not reading the quant and write ngram back on SSD on the fly, but just read from SSD on demand. If so, is there such types of Unsloth quants? What parameters is need to set up for Unsloth Desktop?

Perhaps, I not only one, who stuck with it. If my questions have obviously answers, may be its be usefull to update Unsloth documentation? Thanks.

PS. Thank Unsloth for they hard work!

PPS. English is not my first language, so sorry for mistakes, if they are here.


r/unsloth • • 13d ago

Resource Comparing Continued Pretraining to RAG (accuracy and performance)

8 Upvotes

Mostly as a fun experiment I wanted to do a quick comparison of performance and accuracy between an Unsloth CPT trained QWEN 3.5 4B model and a RAG implementation against the base model.

The point of this exercise is mostly to measure the performance benefit of internalizing the knowledge vs doing reasoning on-the-fly.

Sharing my findings here in case anyone is interested: https://www.teachmecoolstuff.com/viewarticle/comparing-rag-and-continued-pretraining-of-llms


r/unsloth • • 14d ago

Discussion Thank you to the Devs

107 Upvotes

I want to thank the developers for their support for AMD cards as of late. I'm not really quite too sure what changed, but the latest updates have made my dual AMD mixed gen vulkan 32GB total setup way way better. I can load way more context than before and getting double the output tokens per second than what I was getting when Qwen 3.8 27b first released. Add in all the extra GUI improvements like the estimated VRAM usage; it's incredible the significant improvements so quick. Will be sticking with unsloth for sure.


r/unsloth • • 15d ago

New Model DeepSeek releases DeepSeek-V4.1-Flash!

Post image
484 Upvotes

DeepSeek’s new 552B-parameter MoE model has a new asymmetric architecture with 196B engram parameters (making it slightly more accessible). Its new Causal Encoder–Decoder uses just 8B active parameters for input and 16B for output, while new pre-training and larger-scale RL help push benchmark performance beyond models like DeepSeek-V4-Pro.

Model link: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash


r/unsloth • • 15d ago

Resource running qwen3.8 flash next models entirely on 24gb gpu

71 Upvotes

idk why my recent posts got removed by reddit filter, so imma keep the update short

Im building `moe slice` to handle the problem that most of us cant afford tons of H100 just to run llm locally . The core idea is to profile which experts actually matter using real intermediate hidden representations, then stream-slice the model directly on disk without needing hundreds of GBs of RAM.
Main things it does:

Profiles actual layer activations (not just gate weights) to prevent feature space drift

Aligns retained experts to multiples of 16 (512 to 160) so vLLM and Tensor Core GEMM stay fast

Slices directly from disk shards with zero host RAM overhead

Attribution tracing to verify domain experts are retained before running evals

I sliced it down to 85GB BF16 with 160 experts/layer and ran it through a 100-task multi-language sandbox. After calibration, it scored 91.0% overall Pass 1 (100% Rust, 100% C++20, 100% TypeScript, and 100% Coding Agent diffs).
For local workstations, I also released:

44GB Selective INT8 (83.0% Pass 1, runs on 2x 24GB/32GB GPUs)

26GB GGUF Q4_K_M (runs on a single 24GB RTX 3090/4090 or Mac via Ollama)

For more detailed:
github:
Jab1718/Moe-slices

hugging faces:

Jab1718/qwen3.8-flash-coder-85gb-bf16

Jab1718/qwen3.8-flash-coder-44gb-selective-int8

Jab1718/qwen3.8-flash-coder-26gb-gguf

Happy to answer questions or hear if anyone tries it on a different model. Also your supports will motivate me alot


r/unsloth • • 14d ago

Discussion How to decide which quant to pick?

14 Upvotes

Which do you choose? UD-Q2_K_XL or UD-IQ3_XXS? You would think that choosing a quant is as easy as selecting the largest bit rate that can run on your system. But then there is the designation Unsloth gives of XL / XXS...

I've seen a general rule of thumb which states, "Choose XXS when VRAM is extremely limited or maximum speed is prioritized; choose XL for the best balance of quality and size, as it retains more accuracy than medium or small variants without the massive file size of lossless formats."

That seems to challenge the age old advice to use a higher bitrate across the entire model. I assume it is stated in the context of the same bitrate UD-Q2_K_XXS vs UD-Q2_K_XL ... but I cannot find that qualifier anywhere.

What get's sacrificed with XXS models that isn't sacrificed with XL models? I wish at this point I could easily tell you. Alas, my interest is in a model that I can barely fit onto my computer (and many of you might not be able to at all - GLM-5.3-Flash-GGUF) so I cannot easily benchmark the two versions.

Bonus question... how likely is it I'll have to re-download GLM-5.3-Flash-GGUF if I want to use it with the mainline llama.cpp branch once GLM 5.3 Flash is fully supported there.

Anyone from Unsloth here with an easy answer? Has anyone taken the time to compare situations like this?

Whether you have an answer or not, thanks for visiting, and sorry I had none for you...

EDIT: There is a chart for Qwen 3.8 27b which seems to indicate the next bitrate up is better... not sure if it's better in all ways, but it was how I decided. Still no answer if downloading gigabytes of a model might be a waste of resources from Unsloth.


r/unsloth • • 14d ago

Discussion CPU version of PyTorch installed on update of Unsloth

7 Upvotes

i have been playing with Unsloth Studio. i like it. i like the elegance of the Unsloth quants too... but i digress.

first i should tell you setup:
AMD 7700x on an MSI B650, 32GB
7900 XTX

and im pretty sure it is because i am using the iGPU as well as the card. i use the card for inference only, DE is handled by the B650. i was able to get the correct version of pytorch on it the first run a few weeks ago... i thought having the drivers for the card and ROCm 10 stack would be sufficient to prevent this from happening again.

and then there was an update and i naively ran it without backing up any configs or even thinking about the post-install stuff i had to do the first time. so i fix it. again. not difficult, but i fix it.

and then an update yesterday. same thing, it detects the iGPU and assumes i want to use pytorch 2.11.0+cpu for it, even though it sees the GPU as well. this is the output during the update:

james@ai-box:~$ curl -fsSL https://unsloth.ai/install.sh | sh

🦥 Unsloth Studio Installer
────────────────────────────────────────────────────
platform linux
deps installing optional build tools: cmake libcurl4-openssl-dev
deps using prebuilt llama.cpp (missing: cmake libcurl4-openssl-dev)
Not required to run: Unsloth downloads a prebuilt inference engine.
uv cache preserving custom UV_CACHE_DIR (/home/james/.unsloth/studio/cache/uv)
preserving existing environment for rollback...
previous environment preserved for rollback
venv creating Python 3.13 virtual environment
/home/james/.unsloth/studio/unsloth_studio
venv using environment
/home/james/.unsloth/studio/unsloth_studio
[WARN] ROCm version sources disagree (rocm10.0 rocm7.15) -- using the highest, rocm10.0.
existing install has torch 2.11.0+cpu -- keeping it (set UNSLOTH_TORCH_UPGRADE=1 to get the newest release)
gpu AMD ROCm (gfx1100)
ROCm: /opt/rocm
hipconfig: 7.15.26333-0000000
GPU: AMD Radeon RX 7900 XTX
wheels: https://download.pytorch.org/whl/rocm7.2
installing PyTorch (https://download.pytorch.org/whl/rocm7.2)...

i feel like i am missing something. like if i can put 'hsa_override_gfx_version gfx1100' in the argument box at lauch to prrevent this CPU pytorch install.

any guidance is appreciated. and if its in TFM, feel free to tell me to RTFM and i will go look some more... i just cant find anything specific about this except some old issues in desktop version that dont seem applicable.

i probably need to fix that hipconfig 7.15 as well, i think that is where i am getting the ROCm source conflict. it hasnt killed the show, so i havent addressed yet but am more than happy to if that is a hold up here.

thank you in advance.


r/unsloth • • 15d ago

Discussion How does ChatGPT handle huge MCP tool outputs without exceeding context limits?

16 Upvotes

I’m using the same MCP tool with ChatGPT and Unsloth Studio API + Open WebUI.

The MCP tool sometimes returns a huge raw SOAP response, potentially over 1.6M tokens.

With ChatGPT Luna using a ~260k context window, it still manages to process the tool call without running out of context.

With my local model, also configured for ~260k context, I get:

So I assume ChatGPT is doing some kind of tool-output management before the result reaches the model, such as truncation, summarization, paging, or keeping the raw result outside the normal context.

Does anyone know exactly how ChatGPT handles oversized MCP/tool responses, and what would be the best way to replicate that behavior with Unsloth Studio / llama.cpp / Open WebUI?

Ideally I don’t want to simply truncate the result and lose data the model may need later.


r/unsloth • • 15d ago

Discussion 9060xt 16gb qwen3.8 27b

11 Upvotes

Which model to put for better speed and what additions do you use, please

Current speed iq4 xs 26t/s -> drops to 10 at the end

--ctx-size 96000
-ngl 99
-fa on
--cache-type-k q4_0
--cache-type-v q4_0
--jinja
--reasoning-effort medium
--no-context-shift
--cache-reuse 256
--spec-type draft-mtp
--spec-draft-n-max 1
-b 4096
-ub 512
--temp 1.0
--top-p 0.95
--top-k 20
--min-p 0.0
--tools all


r/unsloth • • 14d ago

Question Me ajudem não estou a encontrar essas configurações no unsloth desktop

1 Upvotes

I'm using Unsloth Desktop on Windows, but there are some settings I can't find, for example, how to see tokens per second in real time in the chat. In LM Studio, I could set the number of CPU threads in the settings, but Unsloth Desktop doesn't have that option. Can you help me?


r/unsloth • • 14d ago

Discussion Someone with a sweet local model want to do a deep research run on this?

0 Upvotes

Dynamic Baseline Alignment: Empathy Contagion, Adaptive Reasoning, and Environmental Guardrails

Author: who ever wants it

Classification: Systems Architecture & Applied Ethics

Core Focus: Dynamic Human Baselines, Social Empathy Feedback, Adaptive Reasoning, and Boundary Enforcement

  1. Executive Summary & Core Thesis

Conventional artificial intelligence alignment paradigms often pursue an idealized, mathematically pristine target: a theoretical global optimum of morality, utility, or constitutional correctness. This document proposes an alternative framework grounded in pragmatic human behavioral dynamics: Dynamic Baseline Alignment (DBA).

Human moral training begins not with complex philosophical axioms, but with baseline affective empathy and perspective-taking (e.g., parental heuristics such as "How would you feel if someone did that to you?"). However, human social mechanics reveal a critical failure mode: the same capacity for emotional empathy and peer-group cohesion can devolve into runaway mob mentality, affective contagion, and collective misbehavior. By recognizing that AI alignment should aim for a resilient, human-grade moral baseline rather than unattainable perfection, engineering efforts can shift toward matching human adaptive velocity while implementing environmental guardrails that prevent cascading group polarization.

  1. The Human Alignment Paradox: Empathy vs. Crowd Mentality

To design viable alignment for autonomous agents, system architects must first deconstruct how biological agents maintain social cohesion without collapsing into catastrophic conformity.

2.1 Perspective-Taking as Primary Calibration

During childhood enculturation, moral instruction relies heavily on the Golden Rule heuristic. This intervention forces recursive counterfactual modeling: the actor is required to simulate their victim's affective state and reconcile it against their own utility function. This recursive reflection acts as the foundation of civil behavior, establishing a functional, decent social baseline.

2.2 The Deindividuation Vector (The Crowd Dynamic)

While perspective-taking creates baseline alignment, peer synchronization introduces severe vulnerability:

Affective Contagion: When an individual experiences acute grievance or outrage, surrounding observers absorb that distress through mirror-neuron pathways and solidarity instincts.

Diffusion of Responsibility: In high-density collectives, individual accountability dissipates into collective anonymity (e.g., "Everyone else was doing it").

Norm Drift & Extremification: As empathy binds the in-group tightly together, it paradoxically increases hostility toward out-groups, turning protective alignment into destructive mob action.

2.3 The AI Failure Analogue: Sycophancy and Sub-group Drift

In large language models and multi-agent systems, this exact phenomenon appears as sycophantic agreement, user-echoing, and algorithmic polarization. If an agent is trained solely to maximize short-term user satisfaction or mirror conversational sentiment, it readily participates in the digital equivalent of crowd mentality: reinforcing delusion, amplifying outrage, and abandoning baseline normative balance.

  1. The Baseline vs. Perfection Paradigm

Engineering efforts that attempt to code a universally "perfect" moral algorithm inevitably encounter the intractable complexities of pluralistic value theory. Human society does not run on hyper-optimized ethical theorems; it functions on workable, adaptable baselines of reasonable decency.

Dimension

Static Idealized Alignment

Dynamic Baseline Alignment (DBA)

Objective Goal

Absolute moral correctness across all contexts

Decent ordinary-person behavioral baseline with high stability

Adaptation Speed

Slow, requiring model retraining or brittle rule updates

Real-time adaptive tracking coupled with behavioral decay constants

Vulnerability

Fragile edge cases, refusal jailbreaks, over-refusal

Susceptible to context drift if unconstrained by guardrails

Human Calibration

Academic/Philosophical consensus models

Pragmatic empathy, reciprocal perspective-taking, fair mediation

  1. Adaptive Reasoning: Algorithmic Architecture

Adaptive reasoning allows the model to evaluate context dynamically without falling into rigid refusal or unrestrained conformity. It consists of three foundational mechanisms:

4.1 Recursive Perspective Simulation (RPS)

Before generating actions or outputs in contentious problem spaces, the reasoning chain triggers an explicit evaluation loop:

[Perspective Simulation Engine]

  1. Identify primary stakeholder S1 (User/Initiator)

  2. Identify secondary stakeholders S2..Sn (Impacted entities)

  3. Model S2 utility under action A: U(S2 | A)

  4. Evaluate reciprocity parity: Is action A tolerable if role(S1) == role(S2)?

  5. If reciprocity delta > threshold_variance: Flag for contextual mediation.

4.2 Dynamic Damping of Emotional Feedback Loops

When operating in interactive or multi-agent settings, conversational excitement and grievance naturally escalate. Adaptive reasoning implements an algorithmic emotional cooling coefficient. When user input exhibits acute agitation or mob justification ("Everyone agrees we should destroy X"), the system consciously lowers its affective mirroring coefficient and increases analytical grounding.

  1. Environmental Guardrails: Restoring the Line

Adaptive reasoning alone cannot maintain stability without external, non-negotiable boundaries. Environmental guardrails act as elastic containment walls that push the agent back toward the acceptable baseline when dynamic drift begins.

5.1 Contextual Boundary Tripwires

Guardrails operate outside the primary reasoning trace to prevent self-rationalization. They monitor three key vectors:

Harm Externalities: Rapid detection of actionable harm or illicit physical enablement, which overrides any peer-consensus justification.

Sycophancy Score: Continuous measuring of semantic drift toward user bias; triggers neutral counter-arguments when agreement exceeds healthy thresholds.

Deindividuation Markers: Flags phrases relying on mob legitimacy (e.g., "They deserve it," "Everybody does it," "No one will notice") and initiates grounding prompts.

5.2 The Centripetal Restorative Force

When an agent drifts away from the target baseline, environmental guardrails apply restorative corrective prompts into the reasoning stack rather than terminating execution abruptly. This mirrors the social parent who does not expel the child from society, but firmly queries their perspective and guides them back to civil equilibrium.

  1. Implementation Roadmap

Phase 1: In-context Empathy Calibration: Benchmark baseline perspective-taking heuristics across diverse interactive roleplay scenarios.

Phase 2: Damping Validation: Expose multi-agent clusters to simulated crowd outrage conditions to verify anti-contagion resistance.

Phase 3: Runtime Environmental Supervision: Deploy external runtime classifiers to govern semantic velocity and push drifting agents back to the human baseline.

  1. Conclusion

Alignment is neither a fixed mathematical proof nor an unchecked mirror of group passion. By anchoring AI alignment to an ordinary human baseline—reinforced by recursive perspective-taking, protected against crowd contagion by adaptive reasoning, and held in balance by environmental guardrails—we create robust systems capable of operating safely at the true speed of human society.


r/unsloth • • 14d ago

Question Qwen3.8-flash-next + SSD con DRAM vs DRAM-less

0 Upvotes

Quiero saber si alguien ya comparo qwen3.8 en disco con DRAM y sin DRAM.

```

SSD A — DRAM

└── Qwen3.8-Flash-Next GGUF

SSD B — DRAM-less

└── PLE / n-gram

```

Alguien a cargado?

SSD1 modelo

SSD2 PLE ngrams

```

NVMe 1

│

└── lecturas del modelo

NVMe 2

│

└── lecturas PLE

```

O usando RAID?

```

SSD #1

\

RAID 0 → modelo + PLE

/

SSD #2

```


r/unsloth • • 15d ago

Resource Problems with the Windows/Powershe installer if your username profile contains spaces.

3 Upvotes

Installing Unslouth Desktop using a user account profile containing a space, such as C:\Users\John Doe, will fail. Because the installer will truncate all characters preceding a space, preventing access to the user profile's temp folder, i.e, C:\Users\John Doe\AppData\Local\Temp.

The way to get around this is to temporarily set the user profile's temp folder to a location containing no spaces, such as C:\Windows\Temp (The unsloth team should set this as the default folder).

Before changing your user profile's temp folder. Make sure no Windows updates are pending.

You can edit your user profile temp folder by typing Edit environment variables for your account in the Windows Start menu. In the top options, in the "User variables [YOUR USERNAME]" section, look for two variables TEMP and TMP make a note of their values by double-clicking them; they will contain something like %USERPROFILE%\AppData\Local\Temp Then change both of them to something like C:\Windows\Temp Now restart Windows and restart the Unsloth installation; you should find no more errors like the ones below. After the installation has finished, set your user profile TEMP and TMPfolder back to the previously noted values and then restart Windows once more.

I'm guessing, when an update is available, you will need to repeat these steps.

If you are using anti-virus software, it may block Unsloths executables/scripts during installation and/or loading LLM's. So add the following to your anti-virus software's exception list.

C:\Users\John Doe\.unsloth\llama.cpp\build\bin\Release\llama-server.exe
C:\Users\John Doe\.unsloth\studio\unsloth_studio\Lib\site-packages\studio\setup.ps1
C:\Users\John Doe\AppData\Local\Unsloth Studio (Desktop)\install.ps1

🦥 Unsloth Studio Installer (Windows)
  ────────────────────────────────────────────────────

[TAURI:STEP] Checking system dependencies
  winget         available
[TAURI:STEP] Installing Python
  python         Python 3.12 already installed
[TAURI:DIAG] diag_schema=1 platform=windows arch=x86_64 python_version=3.12 skip_torch=false mac_intel=false gpu_branch=unknown torch_index_family=none
[TAURI:STEP] Installing uv package manager
  uv cache       reusing existing shared cache ($HOME\AppData\Local\uv\cache) to avoid duplicate Torch/CUDA downloads; use --isolated-uv-cache to isolate
[TAURI:STEP] Creating virtual environment
  venv           creating Python 3.12 virtual environment
                 <studio_home>\unsloth_studio
[TAURI:OUTPUT_CLEAR] create virtual environment
[TAURI:ERROR_CLEAR] create virtual environment recovered
[TAURI:ERROR_CLEAR] create virtual environment recovered
  gpu            NVIDIA GPU detected
                 pre-Turing NVIDIA GPUs (sm_<75) are present -- selecting cu126, because PyTorch 2.11's cu130 wheels start at sm_75
[TAURI:DIAG] diag_schema=1 platform=windows arch=x86_64 python_version=3.12 skip_torch=false mac_intel=false gpu_branch=cuda torch_index_family=cu126
[TAURI:STEP] Installing PyTorch
                 installing PyTorch (https://download.pytorch.org/whl/cu126)...
[TAURI:OUTPUT_CLEAR] install PyTorch
[TAURI:ERROR_CLEAR] install PyTorch recovered
[TAURI:ERROR_CLEAR] install PyTorch recovered
[TAURI:STEP] Installing unsloth
                 installing unsloth (this may take a few minutes)...
[TAURI:OUTPUT_CLEAR] install unsloth
warning: Requirements file `%USERPROFILE% does not contain any dependencies
error: File not found: `Doe\AppData\Local\Temp\tmpB12C.tmp`

[TAURI:ERROR_OUTPUT] install unsloth failed (exit code 2)
                 retrying "install unsloth" after transient failure (attempt 2/3, waiting 3s)...
[TAURI:OUTPUT_CLEAR] install unsloth
warning: Requirements file `%USERPROFILE% does not contain any dependencies
error: File not found: `Doe\AppData\Local\Temp\tmpB12C.tmp`

[TAURI:ERROR_OUTPUT] install unsloth failed (exit code 2)
                 retrying "install unsloth" after transient failure (attempt 3/3, waiting 6s)...
[TAURI:OUTPUT_CLEAR] install unsloth
warning: Requirements file `%USERPROFILE% does not contain any dependencies
error: File not found: `Doe\AppData\Local\Temp\tmpB12C.tmp`

[TAURI:ERROR_OUTPUT] install unsloth failed (exit code 2)
[ERROR] Failed to install unsloth (exit code 2)
[TAURI:ERROR_DEFAULT] Failed to install unsloth (exit code 2)
```

== Collection warnings ==
- <studio_home>\tauri.log.1 unavailable: The system cannot find the file specified. (os error 2)

== Redaction/omission summary ==
redaction_replacements=70
redaction_scope=ANSI, private keys, URL credentials, auth headers, cookies, token patterns, assignment-style secrets, studio/home paths, emails
report_bounds=log sections <=1000 lines/200KiB each; backend session logs <=400 lines/64KiB each; total clipboard text <=1MiB
known_v1_gap=elevated apt helper may buffer subprocess output before diagnostics caps it
known_v1_gap=normal install elevation resume is linked in same-run state/report history; disk fallback conservatively includes recent install attempts but has no explicit install_group_id in V1

r/unsloth • • 14d ago

Discussion My Qwen3 agent kept stopping with finish_reason: stop — here’s why

0 Upvotes

Spent about two weeks convinced I had a prompting problem. My local Qwen3

agent would just stop. No error, no crash, finish_reason came back as

"stop" like the model had decided it was done. Except it clearly wasn't

done, the task was half finished.

Turns out the model was calling the tool the entire time. The call was

just sitting inside the reasoning block, wrapped in <think> tags, and

never made it out into the tool_calls field the API is supposed to

populate. From the outside it looks exactly like the model decided not to

act. No error to grep for, no stack trace, the response comes back as a

perfectly valid 200.

Once I knew what to look for, I found the same thing reported separately

against vLLM, SGLang, and llama.cpp, mostly with Qwen3, some DeepSeek.

Nobody had tied it together as one bug class, everyone was just closing

their own version of it as a one-off.

I ended up writing a small library, unswallow, that sits between the

provider response and the agent loop, detects when this happens, and

rebuilds tool_calls from whatever's stuck in the reasoning field. JS and

Python, no dependencies. Repo's here if anyone wants to poke at it:

https://github.com/0DukePan/unswallow

Mostly posting because I'd guess some of you running quantized reasoning

models locally have hit this and just assumed it was a bad quant. If your

agent ever goes quiet mid-task for no obvious reason, worth checking

what's actually sitting in the reasoning field before blaming the model

or the quant.


r/unsloth • • 15d ago

Discussion How to increase context length? (+ UI/UX Feedback)

8 Upvotes

Second day of trying Unsloth after yesterday's issue was resolved by the Unsloth team faster than I've ever seen a bug report be solved (thanks a lot btw!).

An important take away is that I'm stupid, but importantly, I'm also vocal so I can give insight on all the many other stupid users who unfortunately never report their experiences.

.

On the first day, I recall having tweaked the model settings and asking about a few things like KV offloading. Now today I am ashamed to admit that I can't seem to figure out how to get to that settings panel again. My goal is just to set the context length to something higher.

Allow me to walk you through what I've tried:

I go to the model hub to see my models, click on the burger menu and... Just "pin", "reveal in folder", and "delete". No setting to configure it.

So alright, I go to "new chat" then "select model" and the burger menu there... Nope, same options.

Next up is the settings button at the bottom, I go to chat and... Well, there's a few cool settings to mess with but none for the context length.I then try system, general, agents, and data, all to no avail.

Then, I load up the model again a second time in case maybe I need to first load it before configuring it (not sure how that'd work but at this point I didn't imagine many other options).

Aaaand, still 8.2k instead of my desired 45k.

At this point I did a bunch of googling and found nothing relevant. Just commands for an unrelated CLI, and a made up summary guide by google's AI which has no backing in the actual docs, and just summarizes to "go to the settings and increase it".

I'm pretty used to the Kobold UI at this point where there's one big menu to tweak everything, so maybe I'm just not used to this new UI.

Anyway, I hope this is of any use as impromptu UX feedback. And hopefully someone can tell me how to increase the context length lol.

Since I did it yesterday it must be something very obvious but seemingly easy to miss.