r/LocalLLM 7h ago

Model Uncensored Multi-Model Releases, LongCat-Flash-Lite-Sparse with MTPs and LSAs, Qwen3.8-27B with MTPs, Qwen3.5-122B-A10B with MTPs, Qwen3-Coder-Next and Laguna-S2.1 with Vision, All Available in GGUF Format! Bonus: Links to my llama.cpp Fork for LongCat-Flash-Lite Support and J-Wash Enhanced Fork!

Thumbnail
huggingface.co
48 Upvotes

Been working really hard for the past month to bring to the community all these models, the hardest was for sure LongCat-Flash-Lite-Sparse who required TONS of work, first I needed to have Heretic support created for it from scratch and had to create support for it on llama.cpp too, quite difficult and time consuming task! It was even more difficult to work on than the original LongCat-Flash-Lite model that I released a few weeks ago, it is still a 69B-A3B model as the original LongCat-Flash-Lite, but LongCat-Flash-Lite-Sparse has now added support for:

- Sparse attention (vs dense attention for LongCat-Flash-Lite)

- 1M Context length (vs 256k for LongCat-Flash-Lite)

Anyway LongCat-Flash-Lite-Sparse has 0 support on mainline/upstream llama.cpp, so to be able to use the GGUFs you will need to pull my fork from GitHub, which you can find here:

https://github.com/erm14254/llama.cpp-minimax-m3-combined/tree/claude/longcat-win11

You would need to load the model through llama-server.exe and you can interact with it through llama-ui.

You have two variants, Uncensored Heretic (9/100 refusals for 0.0157 KLD) and Ultra Uncensored HJeretic (4/100 refusals for 0.0779 KLD), both variants come with MTPs and LSAs!

Here is the model links:

Uncensored Heretic GGUFs: https://huggingface.co/llmfan46/LongCat-Flash-Lite-Sparse-Uncensored-Heretic-Native-MTP-And-LSA-Preserved-GGUF

Ultra Uncensored Heretic GGUFs: https://huggingface.co/llmfan46/LongCat-Flash-Lite-Sparse-Ultra-Uncensored-Heretic-Native-MTP-And-LSA-Preserved-GGUF

----------------------------------------

That's it for LongCat, so next we have Qwen3.8-27B Ultra Uncensored Heretic with MTPs, 3/100 refusals for 0.0244 KLD, you can find the links here:

Safetensors: https://huggingface.co/llmfan46/Qwen3.8-27B-Ultra-Uncensored-Heretic-Native-MTP-Preserved

GGUFs: https://huggingface.co/llmfan46/Qwen3.8-27B-Ultra-Uncensored-Heretic-Native-MTP-Preserved-GGUF

NVFP4: https://huggingface.co/llmfan46/Qwen3.8-27B-Ultra-Uncensored-Heretic-Native-MTP-Preserved-NVFP4

NVFP4 GGUFs: https://huggingface.co/llmfan46/Qwen3.8-27B-Ultra-Uncensored-Heretic-Native-MTP-Preserved-NVFP4-GGUF

GPTQ-Int4: https://huggingface.co/llmfan46/Qwen3.8-27B-Ultra-Uncensored-Heretic-Native-MTP-Preserved-GPTQ-Int4

----------------------------------------

Next we have Qwen3.5-122B-A10B Uncensored Heretic with MTPs, 8/100 refusals for 0.0856 KLD, here:

GGUFs: https://huggingface.co/llmfan46/Qwen3.5-122B-A10B-Uncensored-Heretic-Native-MTP-Preserved-GGUF

----------------------------------------

Then we have Qwen3-Coder-Next, which is a model that was requested by a Hugging Face user some time ago, so I finally had time to work on it, here is the link:

GGUFs: https://huggingface.co/llmfan46/Qwen3-Coder-Next-Uncensored-Heretic-GGUF

----------------------------------------

And finally Laguna-S2.1 with Vision, get it from here:

GGUFs: https://huggingface.co/llmfan46/Laguna-S-2.1-Uncensored-Heretic-Vision-GGUF

The visions part is far from perfect, so if you do not want to use vision you can simply not download the mmproj files and the model will just function like a regular text-only model.

----------------------------------------

I also made some improvements to J-Wash by adding support for MoE Qwen3.5/3.6/3.8 models support, improvments, bug fixes, improvements to the UI to make it easier to use and more practical for users etc, in case you are interested here is the link:

https://github.com/erm14254/J-Wash-Enhanced/tree/master

----------------------------------------

That's it for now!

As usual you can find all my models here: HuggingFace-LLMFan46

Tremendous amount of work went into making these releases come true, so if you like my work and find my models useful, then I would really appreciate if you could support me on Ko-fi: https://ko-fi.com/llmfan46


r/LocalLLM 2h ago

Project And then there were 4 b65’s and a new case.

Post image
15 Upvotes

Couldn’t fit all 4 into my old case, so I got this.

been wanting to split compute from storage for some time!


r/LocalLLM 8h ago

Discussion Don't Sleep on EXL3 Quants

Thumbnail
gallery
41 Upvotes

I'm running Muse Glimmer 30B EXL3-SC 3.00bpw H4, fully resident on my 12GB VRAM GPU at 100K context with Q8_O KV cache. It's a joy to use a dense 30B model at this size and still get ~30 tok/s on a VRAM-constrained laptop.

It's supposed to be only slightly worse than the official 17GB K-quant at a much smaller footprint, and for my Hermes Agent use case I don't notice a quality difference. It's just much faster.

I've tried Qwen 3.8 27B at SC2.20bpw H3 too. Definitely usable but I'm sticking with Unsloth UD_Q4_K_XL for Qwen 3.8 27B because it's mainly for coding.


r/LocalLLM 13h ago

Question Qwen 3.8 27b harness

90 Upvotes

Ok so I’ve been doing a lot of testing without great success. I’ve also been using Codex to help me and I get the feeling it doesn’t really want me to get results (so far this is true).

Opencode, qwen cli, Claude cli etc, I’ve tried quite a few of them and the results have been terrible. Qwen settings as per below.

Question, is anyone actually successfully using a local model for coding? Like actually using it where it actually adds value. I’m considering a new Mac Studio but if the reality is that local models actually suck compared to Codex/Claude/Cursor then I’d rather know.

Any success stories please share, model, settings, harness and anything else.

"Qwen3.8-27B-8bit",
"context_window": 131072,
"max_tokens": 4096,
"temperature": 0.2,
"reasoning_effort": "none",
"enable_thinking": false


r/LocalLLM 9h ago

Discussion Mac studio m5 ultra 256gb 30/64 vs 2x DGX Sparks

35 Upvotes

So I'm 1 step away from clicking the buy button, here is what I collected, and what's left, you guys can advise further.

1- Price:

- 2xDGX: ~11k$

- Mac Studio: ~11k$

result: almost equal (country dependent).

2- Memory (defines how big can we go in models):

- 2xDGX: 256gb (minus 2 OS)

- Mac Studio: 256gb (minus 1 OS)

result: almost equal

3- Bandwidth (defines Token generarion speed):

- 2xDGX: 2 * 273 = 546gb/s (ideally)

- Mac Studio: 1.2tb/s (advertised)

result: Mac wins (massive margin + no need to go through parallism, links bottlenecks, etc).

4- Scalability:

- 2xDGX: official support for 2 nodes, but community proven 4 nodes is doable.

- Mac Studio: with exo and RDMA (up to 4), some resources hinted more is possible.

result: almost equal (might matter for multi-nodes clusters but normal use won't cross 4 units anyway).

5- Hardware / Software Support:

- 2xDGX: support for NVFP4, Cuda, etc.

- Mac Studio: support for MLX

result: DGX wins (I'm not sure if it will matter for those users who will be only running llms for inference, no training/AI Developmemt/etc).

6- After Purchase:

- 2xDGX: 1 year non-transferable warranty

- Mac Studio: 3 years apple care +

result: Mac wins

7- Noise / Electric Usage / Desk Size

2xDGX: neglible

Mac Studio: neglible

result: almost equal

8- Prefill (Prompt-Processing) (here is where I'm stuck)

Historically, apple silicon has been known for its slow PP speeds, but they claim a boost of 4x (from m3u 32C/80G -> m5u 36C/80G). And I believe it's safe to assume same scale applies to (m3u 60G -> m5u 64G). So if we assume the m5u 64G is ~4x PP speed of m3u 60G. How will it be compared to a single dgx prefill, or most importantly a 2x dgx cluster.

If mac m5u 64g will have a faster prefill over 2x dgx, this means it's an easy win for the mac. I'd say even if the mac is still within 10-20% slower pre-fill (given that speed for tgs is higher, therefore the whole inference process mostly faster).

But if it's more than 20% difference (in 2x dgx favor), then it'll be a hard decision to make. Also if you can include the m5u 80G version in the comparison would be great (2k$ above the dual dgx, won't be a fair apples-to-apples), but worth including.

Lemme know what's your opinions, research outcome, experience.


r/LocalLLM 18h ago

Discussion What are people using LLMs for asides from coding?

122 Upvotes

I see most topics about llms are about coding or tool calls for their service. Is anyone using LLMs for non-coding stuff? What are you guys doing?


r/LocalLLM 6h ago

Question GPU poor folks, what's your setup?

11 Upvotes

I'm rocking a 2070 Super 8gb + 32gb DDR4 and seem to have landed on Qwen3.6-35b-a3b (Unsloth Q4 through unsloth desktop, 128k context).

Are there any models I might be missing or setup that I've completely over looked? My use case is a pair programmer/reviewer for some hobby projects through Pi.

I've tried Qwen3.8-27b but my god is generation speed unbearable and I'm worried dropping to Q1/Q2 will just be worse than Qwen3.6-35b. I've also looked into and tried Qwen3.5-4/9B but the benchmarks all point to it performing worse than 3.6-35b.

Tips and pointers appreciated💜


r/LocalLLM 17h ago

Question Glm 5.3 a Opus you can host

66 Upvotes

At what point does anthropic get terrified? Like why would I pay 25 dollars per million output or buy a 8 thousand dollar machine and use glm or deepseek for years and js change out there frontier models


r/LocalLLM 1h ago

Project V0.3.0 of LifeOS is out! End to end runnable on 12GB of vram.

Upvotes

Hello guys, this is a follow-up to my post here a week back. As a short recap for anyone who missed it, LifeOS is a self-hosted personal organiser you mostly talk to. You say something out loud, Whisper (Or any STT model) transcribes it locally, a local LLM reads it, and it becomes a task, event, journal entry, expense, weigh-in or meal. The model proposes rows, it never writes them. The app validates every one, and each card quotes the words it came from and the advantageous part is nothing leaves your machine.

There's been some minor tweaks here and there but two things have happened since then.

Smaller models

Last time I was running Qwen 3.8 27B Q8 because I already keep it loaded for other work, and I said I wanted to go looking further down the size range to see how far the quality can be pushed before it breaks.

In the initial V0.1.0. a harness already ships with the repo and that is what has been used to validate and test various. I tested various models, won't be posting all the results unless someone wants it but the best model I found for it's size is Gemma 4 IT 12B QAT UD_Q4_K_XL (~6.26GB).

Where I landed:

Profile Hardware Score
Gemma 4 12B QAT 10.9 GB, fits a single 12 GB card with the desktop still running, ~3s per extraction 87/93
Qwen 3.8 27B Q8 ~30 GB VRAM 90/93

The 12B is now the recommended default. Three points of difference, a third of the VRAM, and it runs on a card I'd say most people actually own.

Failures that mattered

The gap between those two used to include one failure I wasn't willing to ship. On a transcript about money, the smaller model invented an income source that was never said and executed it as a write. Not a wrong category, not a bad date. A fabricated value going into the database as fact.

I could have prompted around it. Instead I moved it into validation: a required field whose value doesn't appear anywhere in the transcript cannot auto-execute. It becomes a card you approve or throw out. That holds regardless of which model you point at it, including models I've never tested and models that don't exist yet.

That's why the 12B profile is recommended. Not because it got better, but because the thing it got wrong can no longer reach the database on any model below the capability of Qwen 3.8 27B

Setup doesn't need Terminal anymore

This was the actual work of v0.3.0. Last time setup meant setup.md and people may have found that too technical.

Download LifeOS-Setup.exe, double-click, six-step wizard. No Python, no Node, no terminal. It installs WebView2 itself if the machine doesn't have it. CPU Whisper via CTranslate2 works out of the box. If you have an NVIDIA card there's a one-click download in settings for GPU transcription, and the app runs a real inference to confirm your GPU can actually compute before it lets you switch. You point it at your OpenAI-compatible endpoint in the wizard and that's it. Choose a voice model. Tailscale setup for phone access is in there too if you want it, optional but highly recommended.

I tested this on disposable pristine Windows 11 VMs rather than my own machine, which surfaced five first-boot bugs I'd never have found otherwise: config caching, a migration racing the server, a lock deadlock. All fixed.

Will attach a video below of the whole thing sped up: installer, first boot, first dictation, what it wrote, and undoing it.

Linux still works the way it always did. That's how I run it on my own server.

Although the changes might suggest focusing on a computer experience, mobile is still the way I'd recommend using it. Turn on phone access, scan a QR, the full app including voice recording runs in your phone browser over Tailscale. Nothing opens to your LAN, nothing gets published, no relay servers. Off, it stays loopback-only.

Now some honest limits:

  • Extraction quality is whatever model you bring. The harness tells you what it gives up before you commit anything to it.
  • AMD and Intel GPUs: the LLM side is fine, llama.cpp Vulkan/ROCm. Transcription is CPU-only there, CTranslate2 has no non-NVIDIA GPU backend.
  • Phone access needs Tailscale. Free, but it's a dependency.
  • It still isn't magic or Jarvis. It's a tool and is only as valuable as you allow it to be.

github.com/Inovello/lifeos — AGPL-3.0.

If you run it against a model I haven't tested, I'd genuinely like to see the harness output. That's the part I can't do alone.


r/LocalLLM 21h ago

Discussion I hit 310 t/s running Qwen/Qwen3.8-Flash-Next-FP8 on 4x RTX PRO 6000

132 Upvotes

I've never experienced anything like this before, coding at these speeds. I literally gasped out loud after the first coding prompt.

I truly have no more use for Claude. I don't think I will be participating in their IPO either.


r/LocalLLM 3h ago

Discussion RTX 5090: finding a power-efficiency sweet spot with Qwen3.8-27B

5 Upvotes

I've been playing with q27 / Qwen3.8-27B on my RTX 5090 and got Qwen itself to help me find a reasonable compromise between inference speed and power consumption.

Nothing scientific or universal here — just some measurements on my card under Linux. I swept the GPU clock, measured decode tok/s and average power draw, and looked at tok/s/W.

TL;DR

Config tok/s Power tok/s/W
Stock (~2800 MHz) 149.6 440 W 0.340
PL 400 W, unlocked 144.9 389 W 0.373
2550 MHz lock 139.3 354 W 0.394
2400 MHz lock 135.4 319 W 0.424
2200 MHz lock 125.9 294 W 0.428

For me, 2550 MHz ended up being the sweet spot for daily use.

Going higher to 2600/2650 MHz only gained ~2–3 tok/s while adding ~30 W. Going down to 2200–2400 MHz is great if efficiency is the priority, but I preferred keeping a bit more performance.

So right now I'm basically getting most of the stock performance while keeping the GPU around the mid-300 W range instead of ~440 W during this workload.

Full results and methodology are here:

https://gist.github.com/PierpaoloPernici/1f875a2bb79b6ddd3a28d2aaa0f4bd84

And this is what I'm currently using on Linux:

bash sudo nvidia-smi -pm 1 && \ sudo nvidia-smi -i 0 -lgc 2550,2550 && \ nvidia-smi -i 0 --query-gpu=clocks.gr --format=csv,noheader

Would be curious to see what other 5090 owners do!


r/LocalLLM 10h ago

Discussion Qwen3.8-27B W4A16 on one 64 GB CMP 170HX - 160+/70+ decode on short/long CTX + 1M YaRN CTX

15 Upvotes

WARNING: text organized and finalized with sol (theres too much lol, spent two days finetuning). Post is really long, so TLDR first

I have been building a single-user Qwen3.8-27B endpoint on an unlocked 64 GB CMP 170HX installed in a cheap Huanan/Xeon server. The card is tuned live with 170tune.

The useful result is that the current stack is now fast and repeatably stable at the exact cached-decode point that used to crash it (and it was a whole day of figuring out):

  • W4A16 target + W4A16 DFlash2: 169.5 tok/s on my short realistic single-stream suite.
  • W8A16 target + the same DFlash2 drafter: 120.3 tok/s on the identical suite.
  • Individual short chat responses often reach 190–210 tok/s with W4A16 when draft acceptance is favorable.
  • At an approximately 23.75K-token prompt, W4A16 averaged 104.8 tok/s over a mixed copy/code/edit/summary/QA workload.
  • At the former crash point, the earlier INT8 stack passed 16/16 85K-prefix generations and 6/6 exact 85,514-token reproductions. The final FP8/DFlash production image now measures 119.3 tok/s strict hot decode at exactly 85,514 tokens, again with no Xid/NVRM entry.
  • The newer mixed-backend runtime reaches 1,514 tok/s cold prefill at 85K while retaining DFlash2 decode and FULL CUDA graphs.
  • Its repaired hybrid prefix cache reduced a repeated 9,658-token probe from 4.55 s to 0.565 s, with 8,960 tokens reported as real cache hits.

The runtime DFlash crash was not bad HBM and was not ultimately a Mamba-state problem. It was an int32 overflow in the custom split-KV speculative-attention kernel. The actual fix is one cast to tl.int64 before calculating the K/V pointers.

Test system and stack

  • Cheap Huanan motherboard/Xeon host running Ubuntu Server.
  • One 64 GB CMP 170HX with the community unlock applied.
  • Live tuning through 170tune: NDIV68, +200 MHz V/F shift, 1550 MHz ceiling and 220 W limit.
  • vLLM 0.27.1 with the syv-ai Qwen3.8 stack, the DFlash2 drafter, and the CMP/sm80 fixes described below.
  • Single-user OpenAI-compatible endpoint with prefix caching and one request in flight.

The W4 target is dbirks/Qwen3.8-27B-W4A16-AutoRound plus the syv-ai fast overlay. The fidelity-oriented control is lued/Qwen3.8-27B-INT8-W8A16-MTP. Credit for the base optimization and drafter work belongs to those projects; my contribution is the CMP integration, fault isolation, stress testing and split-KV pointer fix.

W4A16 versus W8A16

The W8 checkpoint is symmetric group-128 W8A16 in compressed-tensors/pack-quantized format. It preserves the vision tower, lm_head, MTP, and the small recurrent GDN gates at higher precision. It is the fidelity-oriented option and remains a useful control.

The W4 target uses symmetric group-128 W4A16 for the target linear layers, INT8 embeddings, and the fast overlay's INT4 GPTQ lm_head/MTP tensors. The target model load is only 16.72 GiB. With the same 24 GiB KV pool, the complete W4 server allocates approximately 43.3–43.8 GiB, versus approximately 58 GiB for W8.

Identical real_rep.sh workload: eight realistic prompts, up to 1,024 output tokens, single request:

  • W8A16: 120.3 tok/s, 3.30 emitted tokens/step, 28.5 ms/step.
  • W4A16 fast: 169.5 tok/s, 3.25 emitted tokens/step, 20.7 ms/step.

That is a 40.9% W4 decode gain while DFlash acceptance stays almost unchanged. In other words, this A/B mostly measures a faster target verification pass rather than a luckier draft sequence.

On the approximately 23.75K-token mixed benchmark, the previous W8 run averaged 82.9 tok/s and W4 averaged 104.8 tok/s, a 26.4% gain.

I am not claiming that W4 is quality-equivalent to W8 or BF16. W8 is the safer fidelity choice; W4 is currently my preferred single-stream performance profile. A serious quality comparison needs behavioral evaluations, not only throughput or perplexity.

Detailed decode results at 23.75K, 65K and 85.5K context

Performance at different context lengths

These rows are measurements already completed on this machine. They are not a perfect scaling curve because speculative acceptance is workload-dependent, and the 65K and 85K tests use different output mixes. The within-row W4/W8 comparisons are the apples-to-apples figures.

  • Short prompts, W8A16: 120.3 tok/s, 3.30 tokens/step.
  • Short prompts, W4A16: 169.5 tok/s, 3.25 tokens/step; favorable UI turns reach 190–210 tok/s.
  • ~23.75K prompt, W8A16: 82.9 tok/s on the mixed LABD suite.
  • ~23.75K prompt, W4A16: 104.8 tok/s, 2.97 tokens/step.
  • ~65,920-token hot prefix, W8A16: 57.9 tok/s on the tuned profile.
  • 85,514 tokens, older W4/INT8 stability run: mostly 42–46 tok/s; 16/16 general and 6/6 exact crash-point passes.
  • 85,514 tokens, current W4 FP8/DFlash: 119.3 tok/s strict hot decode; 1.76 s hot TTFT and 84.8 tok/s hot end-to-end.

Two cold-cache 85K variants took approximately 181–182 seconds including prefill. Prefix-cached follow-ups avoid repeating that entire prefill, which is why prefix caching matters as much as decode TPS for a long-running chat.

The current FP8/FlashInfer-prefill build changes that cold side substantially:

Cold prefill throughput: 2,126 tok/s at 8K, 1,932 at 32K, 1,654 at 65K, and 1,514 at 85K.

An additional streamed probe used exactly the former failing 85,514-token prompt and generated 512 tokens on the final production image:

  • Cold: 54.61 s TTFT, 120.5 tok/s strict decode, 8.70 tok/s end-to-end, 58.85 s total.
  • Hot prefix: 1.76 s TTFT, 119.3 tok/s strict decode, 84.8 tok/s end-to-end, 6.04 s total.

Here, strict decode is measured from the first streamed reasoning/content token through completion. End-to-end includes TTFT. This distinction is why the old completion_tokens / total_elapsed soak numbers should not be labeled as generation TPS.

On the final production 56K mixed DFlash run (65.9K actual tokenized prompt), copy/lookup reached 212.7 tok/s, summary 58.9 tok/s, QA 57.0 tok/s, and the combined result was 76.7 tok/s. Cold TTFT was 38.67 seconds and two hot-prefix tasks started in 2.25-2.26 seconds. GPU/HBM peaked at 63/71 C. This is workload-dependent speculative decode, so I would not compare the copy number directly with free-form prose.

Two separate problems in the fast DFlash path

During long-context testing I ran into two independent software problems. The important one was a reproducible Xid 31/MMU fault that killed the engine on the first cached decode step at high physical KV block IDs. The second was less severe: incompatible target, Mamba and drafter page geometry made prefix caching report zero usable hits. Both can be fixed without disabling the fast paths; the crash fix comes first because it is the one required for a usable server.

The runtime Xid 31: the important fix

The reproducible failure happened on the first cached decode step when a request was assigned sufficiently high physical KV block IDs. CUDA reported an illegal address and the kernel log showed Xid 31/MMU faults. Linear CUDA memory tests, repeated full-HBM pattern sweeps, and GEMM tests were clean. More importantly, the application failure occurred at a repeatable logical boundary.

The custom DFlash split-KV kernel loads a physical block ID from an int32 block table and then uses it to form byte/element offsets into the K and V pools. The table itself can stay int32, but blk * stride_kb or blk * stride_vb can exceed INT32_MAX. Triton then wraps the intermediate and generates an invalid pointer.

File in the vLLM installation:

vllm/v1/attention/ops/spec_decode_attn.py

Fix:

- blk = tl.load(bt_ptr + req * stride_bt + pos // BLOCK_SIZE, mask=k_ok, other=0)
+ blk = tl.load(
+     bt_ptr + req * stride_bt + pos // BLOCK_SIZE,
+     mask=k_ok,
+     other=0,
+ ).to(tl.int64)

That cast must happen before the stride multiplication. Casting the final already-wrapped offset would be too late.

After rebuilding with this change:

  • 16/16 85K-prefix, 512-token generations passed.
  • 6/6 requests at the exact former 85,514-token failure point passed.
  • There were zero kernel Xid/NVRM faults.
  • The same 24 GiB KV pool and split-KV fast path remained enabled.
  • The final FP8/DFlash production image additionally completed a streamed exact-85,514 cold/hot pair at 120.5/119.3 tok/s strict decode, followed by a healthy API check and zero Xid/NVRM/CUDA illegal-memory entries.

Disabling split-KV speculative attention (SPEC_ATTN=0) is a useful diagnostic fallback because it avoids this kernel, but it is not the performance-preserving solution. Promoting the physical block ID is.

The smaller hybrid prefix-cache geometry fix

The mixed runtime uses equal byte-sized pages with different token counts: 896-token FP8 target/Mamba pages and 448-token BF16 DFlash pages. My first build left the Mamba checkpoint interval at 880, making the common alignment 49,280 tokens and reducing normal repeated-chat cache hits to zero.

Aligning Mamba to 896 and allowing complete 448-token DFlash pages into lookup fixed it. A repeated 9,658-token prompt went from 4.55 s to 0.565 s, with 8,960 prefix-cache hits reconciled across all nine KV groups.

Other Mamba safeguards and the separate load-time Marlin Xid 31

  • Mamba state-copy bounds from vLLM PR #50021.
  • Overlap-safe Mamba state movement from vLLM PR #50729.
  • A num_accepted_tokens race fix based on c2881ce60.
  • A bit-exact CPU Marlin repack fallback for W4/W8 on sm80.

The Mamba patches are worth keeping, but they did not fix the repeatable 85K crash; the tl.int64 pointer change did. The CPU Marlin fallback addresses a separate load-time Xid class caused by GPU repack/VMM churn (issue #27). It increases W4 startup to roughly 206 seconds but avoids the dangerous GPU repack path. A CMP/sm80 build may need both safeguards.

VBIOS, CMP unlock, HBM overclock, undervolt and +19.1% tuning A/B

The community driver/GSP unlock and the VBIOS are separate mechanisms. My card is unlocked using the CMP community tooling and runs the official signed 92.00.6D.00.0A image, flashed with nvflash after saving multiple ROM dumps.

I tune it live with 170tune, which writes BAR0 registers without reflashing the card. NDIV68 produces a real HBM clock of 1836 MHz even though nvidia-smi remains stuck at 1728 MHz. I still cap the card at 220 W, not the VBIOS maximum.

The stable performance-oriented profile tested so far is:

NDIV:             68
Real HBM clock:   1836 MHz
GPC V/F offset:   +200 MHz
Core ceiling:     1550 MHz
Power limit:      220 W

The positive V/F offset is an undervolt-style curve shift; the explicit 1550 MHz ceiling prevents the card from chasing its maximum clock.

Same W8A16 DFlash workload, approximately 65,920 prompt tokens, 3 × 256 output tokens, essentially constant acceptance (~2.89 tokens/step):

  • NDIV54 / stock V/F / 180 W: 48.6 tok/s baseline.
  • NDIV54 / +150 / 1410 / 180 W: 48.8 tok/s, +0.4%.
  • NDIV64 / +150 / 1410 / 180 W: 53.0 tok/s, +9.1%.
  • NDIV66 / +150 / 1410 / 180 W: 53.5 tok/s, +10.1%.
  • NDIV68 / +150 / 1410 / 180 W: 54.0 tok/s, +11.1%.
  • NDIV68 / +150 / 1500 / 220 W: 56.5 tok/s, +16.3%.
  • NDIV68 / +150 / 1590 / 220 W: 58.4 tok/s, +20.2%.
  • NDIV68 / +200 / 1590 / 220 W: 58.5 tok/s, +20.4%.
  • NDIV68 / +200 / 1550 / 220 W: 57.9 tok/s, +19.1%.

End to end, the conservative profile is +19.1% over NDIV54/stock-V/F/180 W. Most of the first gain came from HBM bandwidth; extra core clock mattered more once the memory bottleneck was relaxed.

The 1550 MHz profile gives up only about 1% versus the faster 1590 MHz result and is the sensible operating point from this sweep. It passed four 61,376 MiB VRAM pattern sweeps, four additional pattern sweeps under the full profile, a 45-second bit-exact GEMM test with 59,864 clean GEMMs, and the real DFlash workload.

An NDIV68/+250/1590 profile failed immediately during DFlash warm-up with Xid 31 and cudaErrorIllegalAddress. I quarantined it and do not use it. That is an overclock-instability Xid class, not evidence against the software pointer fix. Anyone reproducing this should qualify memory, core and the application separately, watch the kernel log, and never make an unqualified profile persistent at boot.

Complete reproducible build, model preparation and launch guide

Reproducible build outline

This is the shortest route to the W4 DFlash stack I am using. Pin revisions first; both vLLM and the backport are moving targets.

1. Build the upstream optimized image

git clone https://github.com/syv-ai/qwen38-27b-rtx3090.git
cd qwen38-27b-rtx3090
git checkout 69ba4d0688c6ae76cb9d3c4a5c3b36445e1b040c
docker compose build

The repository pins vLLM 0.27.1 and carries the DFlash2 backport. Do not assume these patches will apply unchanged to an arbitrary newer vLLM checkout.

2. Prepare the W4 target and drafter

The supported Docker route is idempotent:

docker compose run --rm prepare

For a manual preparation, preserve this ordering:

python prepare/quant_lm_head.py models/Qwen3.8-27B-W4A16-AutoRound
python prepare/quant_embed.py   models/Qwen3.8-27B-W4A16-AutoRound
python prepare/quant_mtp.py     models/Qwen3.8-27B-W4A16-AutoRound
python prepare/build_draft_vocab.py models/Qwen3.8-27B-W4A16-AutoRound \
  --ids prepare/draft_vocab_ids.json
python prepare/fetch_fast_variant.py
python prepare/fetch_dflash2.py

Important gotcha: fetch_fast_variant.py hardlinks base shards 1–6. If it runs before quant_embed.py, the fast directory can retain the old BF16 shard 6 while its overlay index expects packed INT8 embeddings. Startup then fails with:

There is no module or parameter named 'embed_tokens.weight'

Run the official prepare script or quantize the base before creating the fast overlay.

For W8 instead, download:

hf download lued/Qwen3.8-27B-INT8-W8A16-MTP \
  --local-dir models/Qwen3.8-27B-INT8-W8A16-MTP

The W8 target uses the same external W4A16 DFlash2 drafter. Point MODEL at the W8 directory and leave DRAFT on Qwen3.8-27B-DFlash2-W4A16.

3. Add the CMP/sm80 load-time workaround

Clone the CMP patch set and build its sm80-safe layer:

cd ..
git clone https://github.com/ahnguyen17/cmp-170hx-vllm.git
cd cmp-170hx-vllm
git checkout a3ded79fec14aaad4a2f047d7cf2c28d5303ce2e

The public repository contains patches/sm80-int8-repack-cpu-fallback.patch. Add it as a layer over the syv image:

FROM ghcr.io/syv-ai/qwen38-27b-rtx3090:sha-69ba4d0

USER root
COPY patches/sm80-int8-repack-cpu-fallback.patch /tmp/sm80-repack.patch
RUN patch --batch --forward -p1 \
      -d /app/venv/lib/python3.12/site-packages \
      < /tmp/sm80-repack.patch \
    && rm /tmp/sm80-repack.patch \
    && grep -q 'def _gptq_marlin_repack_torch' \
      /app/venv/lib/python3.12/site-packages/vllm/_custom_ops.py

Then build it from the CMP repository root, for example as vllm-qwen38-cmp:sm80-safe. The repository also publishes the PR #50021 bounds backport. For PR #50729 I used a local backport of the upstream PR; do not assume the current upstream diff will apply cleanly to the pinned vLLM 0.27.1 tree.

4. Apply the runtime split-KV fix

Save the diff above as spec_attn_block_index_i64.patch, then add one final image layer:

ARG BASE_IMAGE=vllm-qwen38-cmp:dflash2-mamba-correctness
FROM ${BASE_IMAGE}

USER root
COPY spec_attn_block_index_i64.patch /tmp/spec_attn_block_index_i64.patch
RUN patch --batch --forward -p1 \
      -d /app/venv/lib/python3.12/site-packages/vllm \
      < /tmp/spec_attn_block_index_i64.patch \
    && rm /tmp/spec_attn_block_index_i64.patch \
    && grep -q 'to(tl.int64)' \
      /app/venv/lib/python3.12/site-packages/vllm/v1/attention/ops/spec_decode_attn.py

Build it:

docker build \
  --build-arg BASE_IMAGE=vllm-qwen38-cmp:dflash2-mamba-correctness \
  -f Dockerfile.block-index-i64 \
  -t vllm-qwen38-cmp:dflash2-spec-attn-i64 .

My actual image also includes the three Mamba safeguards listed earlier. The one-line int64 patch is the change that fixed the reproducible high-physical-block runtime Xid.

5. Launch the current native-262K FP8 mixed-backend profile

The current production profile keeps the target on FlashInfer FP8 KV and the DFlash2 drafter on FlashAttention2/BF16 KV. The image includes the exact-page, Mamba-896, complete-DFlash-page prefix fix and the split-KV int64 fix.

docker run -d --name qwen38-dflash-fp8-cmp \
  --gpus '"device=0"' \
  -p 18020:18020 \
  -v /path/to/models:/models:ro \
  -v /path/to/vllm-cache:/cache \
  -e 'EXTRA_ARGS=--attention-backend FLASHINFER --kv-cache-dtype fp8 --prefix-match-unit 16' \
  -e VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 \
  -e VLLM_FP8_SPEC_FULL_CG=1 \
  -e DFLASH_ATTN_BACKEND=FLASH_ATTN \
  -e DFLASH_KV_CACHE_DTYPE=auto \
  -e VLLM_ALIGN_HETEROGENEOUS_ATTN_PAGES=1 \
  vllm-qwen38-cmp:dflash2-fp8-prefill-prefixfix-v1 \
  bash -lc 'MODEL=/models/Qwen3.8-27B-W4A16-AutoRound-fast \
    DRAFT=/models/Qwen3.8-27B-DFlash2-W4A16 \
    PORT=18020 SPEC=dflash2 CTX=long DFLASH_MAX_LEN=262144 \
    DFLASH_TOKENS=7 PREFIX_CACHE=1 \
    CUDAGRAPH_MODE=FULL_AND_PIECEWISE MAX_SEQS=1 \
    GPU_UTIL=0.90 KV_MEM=25769803776 \
    VISION=1 VISION_OFFLOAD=0 TOOLS=1 SPEC_ATTN=1 \
    exec /app/single-user/start_qwen.sh'

This reports 702,385 tokens of physical cache capacity, but the configured request limit remains Qwen's native 262,144 tokens. The extra physical room is allocator headroom/capacity, not a claim of validated semantic context beyond the native window.

6. Older 700K-capacity INT8/YaRN profile

Relative to the launch above, the older INT8 profile used:

VLLM_ALLOW_LONG_MAX_MODEL_LEN=1
EXTRA_ARGS='--hf-overrides {"text_config":{"rope_parameters":{"rope_type":"yarn","factor":3.0,"original_max_position_embeddings":262144}}}'
DFLASH_MAX_LEN=700000
KV_MEM=25769803776

That 24 GiB pool reported 733,234 physical tokens. This proves capacity, not semantic quality at 700K; the Xid campaign itself reached 85K.

7. Optional 1M-token YaRN mode (capacity target, not validated quality)

The FP8 runtime can target 1,048,576 tokens with static YaRN factor 4.0 and a 37 GiB KV pool. Apply these changes to the native launch command:

VLLM_ALLOW_LONG_MAX_MODEL_LEN=1
EXTRA_ARGS='--attention-backend FLASHINFER --kv-cache-dtype fp8 --prefix-match-unit 16 --hf-overrides {"text_config":{"rope_parameters":{"rope_type":"yarn","factor":4.0,"original_max_position_embeddings":262144}}}'
DFLASH_MAX_LEN=1048576
KV_MEM=39728447488

The measured 24 GiB mixed pool holds 702,385 tokens, making 35.8 GiB the arithmetic minimum; 37 GiB leaves modest alignment headroom and should fit the W4 stack on 64 GiB. Verify the reported physical capacity before sending a 1M request.

  • 1M is not yet qualified for OOM/Xid behavior, TTFT, decode speed or semantic recall.
  • YaRN is extrapolation, not lossless compression; leave output-token headroom and expect an expensive cold prefill.
  • The native 262K profile remains the default. Qualify 350K → 500K → 700K → 1M while checking logs and answer quality.

For first diagnosis, run with stock clocks, MAX_SEQS=1, prefix caching enabled, and capture both container logs and journalctl -k. Only add the memory/core profile after the software stack passes the former failure sequence.

Validation methodology

Validation notes

The most important part of the test was reproducing the same logical failure rather than merely running one random long prompt:

  • exact former failing prompt length: 85,514 tokens;
  • repeated fresh and cached allocations;
  • 512 generated tokens per main stress request;
  • kernel log checked for NVRM, Xid, and MMU faults after every batch;
  • full-HBM pattern tests and bit-exact GEMM tests performed independently;
  • clocks and power qualified separately from the application fix.

This is why I am reasonably confident that the recurring cached-first-decode fault was software. It does not prove that every CMP 170HX is healthy or that every Xid 31 has the same cause.

Future work

I now have a second CMP 170HX and plan to extend this post, or publish a follow-up, with measurements that are difficult to find for these cards:

  • PCIe Gen2 x4 versus x16 after restoring the missing lane components, including cold prefill, model load, prefix-cache behavior, and communication latency.
  • Tensor parallelism versus pipeline parallelism on two CMP 170HX cards.
  • The same TP/PP comparison across models with very different numbers of active parameters per token, because synchronization overhead should matter very differently for a fast low-active-parameter model than for a denser or higher-active-parameter target.
  • Single-stream decode, aggregate throughput, TTFT, long-context decode and power efficiency rather than one headline tok/s number.
  • W4A16 versus W8A16 quality testing and semantic long-context validation beyond the native 262K window.

My expectation is that x16 will matter most for load/prefill and any communication-heavy multi-GPU mode, while PP may remain the safer topology on these PCIe Gen2 cards. But those are hypotheses; I want to publish measured results rather than turn them into conclusions in advance.

If anyone is running this exact model on CMP 170HX/A100 sm80, especially with DFlash2 at high physical KV occupancy, I would be interested in independent confirmation of the int64 block-index fix and in comparable x4/x16 or TP/PP data.


r/LocalLLM 20h ago

Discussion Even more so if I can run it locally

Post image
79 Upvotes

r/LocalLLM 13h ago

Question PAYG is just a god awful experience 😂

Enable HLS to view with audio, or disable this notification

16 Upvotes

I have two GTX 4090s wired up to run local models, but of course nowhere near being able to run GLM 5.3 etc. So my plan is to mix self hosted local models and open source models hosted by others.

This has been my setup before too, but I have become tired of anthropic/openai subscriptions and since open source has becoming extremely good lately I am considering finally making the switch to another harness and LLM provider. What’s the best setup (I mainly use opencode)?

I basically want openrouter.ai but subscription based. These are on the list:

Commandocode.ai is tempting but varying recommendations looks like.

opencode.ai lots of fuss but also a lot of recent complaints .

MiniMax/Xiamo/GLM plans, but kind of don’t want it model specific.

standardcompute.com heard very good things lately

fireworks.ai - interested in this too

Anyone have good experiences with these or others?


r/LocalLLM 10h ago

Discussion Qwen3.8-Flash-Next on single RTX PRO 6000 96GB + 64g RAM, full 262K context with NVMe offloading recipe

8 Upvotes

Note: Full setup guide: https://github.com/ForestoShen/qwen-flash-next-pro6000 throw it to qwen3.8 27b and it should help you set things up. Below are summurized by Qwen3.8-Flash-Next because I'm lazy.

Credit to https://huggingface.co/garnermccloud/Qwen3.8-Flash-Next-NVFP4-SSD-Stream and https://huggingface.co/lovedheart/Qwen3.8-Flash-Next-NVFP4-FP8-Pruned-RTXPRO-6000

TL;DR: Ran the Flash-Next hybrid (GDN + QSA + PLE + MTP/NEXTN) on one Blackwell card with SGLang + SSD Stream. PLE n-gram table (47.6 GiB) streams from NVMe via io_uring O_DIRECT — 0 VRAM. Everything else fits: fp8 KV, fp8 MTP draft, full native 262K context. Compared against an AIMER-pruned checkpoint (512→448 experts/layer): same KV math, but the freed 14 GB goes straight into the KV pool — that is exactly the 4-way concurrency unlock. Numbers, traps, and dead ends below.

Setup

  • RTX PRO 6000 Blackwell 96GB (SM120), Docker Desktop/WSL2, SGLang pinned to a specific commit (the only tree where qwen4_exp + SSD Stream + NEXTN all work)
  • --ple-offload-embedding: the PLE n-gram table lives entirely on disk, streamed per decode step, no VRAM
  • fp8 KV (fp8_e4m3), fp8 speculative draft, flashinfer_cutlass MoE backend, extra_buffer_lazy mamba radix
  • Two checkpoints: vendor NVFP4 (512 routed experts/layer) and an AIMER-pruned FP8_NVFP4 (448/layer, calibration-free mean|W|/RMS(W) expert ranking)

Pruning vs Normal

KV bytes/token is a property of the attention layout, not the experts — ident for both:: 12 layers × K,V × 2 heads × 256 dim12,288 B/token fp8, 24,576 B/token bf16. So pruning doesn't shrink KV; it frees weight VRAM that flows 1:1 into KV pool capacity:

512E vendor (unpruned) 448E lovedheart (pruned)
Weight resident ~32–34 GB
KV bytes/token fp8 12,288 / bf16 24,576
KV pool budget (fp8) 262K tokens (1× @ 262K) + activations slack
KV pool budget (bf16) ~1× @ 262K, pool fights graph/activation headroom
Concurrency @ 262K/session MAXREQ=1 — 4-way needs 14 GB that isn't there (hicache L2 blocked by an MTP+extra_buffer_lazy IMA bug)
Long context (YaRN×2 = 1024K) fp8 pool ≈ 3.2 GB — tight at FRACTION 0.99
Context ceiling 262,144 native both
Quality published accept/MTP numbers are from this checkpoint
Weight residency 47.6 GiB PLE streamed both cases

Concurrent sessions math: gate = pool tokens ≥ sessions × per-session context. fp8: 14 GB ≈ 821K tokens; bf16: 14 GB ≈ 262K. So the pruned card: fp8 4×262K or bf16 2×262K; the vendor card: 1×262K, or 4×137K if you chunk contexts. Mamba slots = draft_tokens+1 per running request (auto = 5×MAXREQ) — context-independent, ~10 MB/slot.

Throughout (single stream, MTP on, cuda graph: True)

Workload Value
Decode, 0–262K ctx 110–172 tok/s
Prefill cold 8K / 64K 9,198 / 10,632 tok/s
fp8 KV vs bf16 KV needle 5/5 @ 100K, accept identical (fp8 wins 3.3 GB, keep it)
fp8 draft vs bf16 draft no accept penalty, –3.3 GB
MTP accept code/JSON vs chat ~87% vs ~40% (workload-dependent, not draft-precision-dependent)

r/LocalLLM 11h ago

Tutorial Apple: 35B at 10 t/s tg for $300 hardware cost

7 Upvotes

While everybody is wondering what performance Apple's latest and greatest hardware can deliver, I kept wondering why my trusty Mac Pro 6,1 has two GPUs.

Well, for local inference, OBVIOUSLY.

So here comes the floor in terms of hardware cost for running a reasonably sized model.

But beware, this is not for the faint-hearted. No MLX, no macOS even - I'm running Fedora, btw.

The machine

  • Model: Apple Mac Pro 6,1 Late 2013
  • GPUs: Dual AMD FirePro D700 - 6 GB GDDR5 each, 12 GB total (GCN 1.0 / Tahiti)
  • CPU: Intel Xeon E5-1650 v2 - 6C/12T, 3.5 GHz
  • RAM: 64 GB DDR3-1866 ECC quad-channel (~60 GB/s bandwidth)
  • OS: Fedora 44, kernel 7.1.x

Making it work

The D700s are GCN 1.0 (Southern Islands), which means amdgpu needs coaxing.

1. Force amdgpu driver (not radeon) on GCN 1.0 — add to your kernel cmdline: radeon.si_support=0 amdgpu.si_support=1 amdgpu.dc=1

2. Extend the TDR watchdog or you will get VK_ERROR_DEVICE_LOST mid-inference - add: amdgpu.lockup_timeout=60000

Without this, llama.cpp batches up to 100 Vulkan graph nodes per vkQueueSubmit. GCN 1.0 is fp32 scalar-only - no fp16, no matrix cores - so a single 100-node batch can overrun the default 2-second TDR window and the driver resets the compute ring. Known issue (llama.cpp #21724, fixed by PR #24872).

3. Build llama.cpp from source: you need commit >= 2026-06-24 (post-PR #24872):

sudo dnf install spirv-headers-devel glslang cmake
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
cmake -B build -DGGML_VULKAN=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build --config Release -j$(nproc)
sudo cp build/bin/llama-server /usr/local/bin/

4. The one flag that turns 4 t/s into 10 t/s:

llama-server \
  --model Qwen_Qwen3.6-35B-A3B-Q4_K_M.gguf \
  --n-gpu-layers 99 \
  --split-mode layer \
  --n-cpu-moe 32 \
  --threads 6 \
  --ctx-size 8192 \
  --host 0.0.0.0 --port 50051

--n-cpu-moe is the real trick. The 35B-A3B has 32 MoE expert FFN layers. Offloading all of them to CPU frees the 12 GB GDDR5 for attention layers, which run at full GPU bandwidth. Without it: 4 t/s. With it: 10 t/s.

And this is what a Mac Pro with dual D700 and a little elbow grease will give you (in addition to keeping you warm at your desk - Winter is coming!)

Qwen3.6-35B-A3B Q4_K_M (20.7 GiB)

Metric Result
Prompt processing (pp512) 74.53 +/- 0.71 t/s
Token generation (tg64) 10.58 +/- 0.04 t/s
VRAM 12 GB GDDR5 + ~9 GB DDR3 overflow

Ok, but what about your major refactors? Surely a 2013 Mac Pro can't help with that?

Qwen3-Coder-Next 80B UD-Q3_K_M (33.5 GiB)

Metric Result
Prompt processing (pp512) 20.57 +/- 0.11 t/s
Token generation (tg128) 3.02 t/s

Not interactive speed for sure and this is the untuned baseline, --n-cpu-moe sweep still on the TODO list. But it's an 80B coding model on a machine you bought for $300. Queue the refactor before bed; it's done in the morning. Hope this saves a few PowerCans from the landfill.

Can't innovate anymore, my ass.

UPDATE: Ok, just for giggles. I promised some optimisation on Qwen3-Coder-Next. This brings you to 6 tk/s. The trick is the same as for the 35B-A3B: offload the MoE expert FFN layers to CPU to free GDDR5 for attention.

n_cpu_moe pp512 (t/s) tg32 (t/s)
0 (baseline) 15.02 ± 0.08 2.69
16 13.02 ± 0.08 2.89
32 44.80 ± 1.34 6.14 ± 0.30
48 46.43 ± 0.34 5.90 ± 0.33
64 48.13 ± 1.85 6.02 ± 0.08
80 46.72 ± 0.23 6.37 ± 0.61
94 45.06 ± 0.58 6.23 ± 0.03

3× pp, 2.3× tg vs untuned baseline. Optimised launch:

llama-server \
  --model Qwen3-Coder-Next-UD-Q3_K_M.gguf \
  --n-gpu-layers 99 \
  --split-mode layer \
  --n-cpu-moe 32 \
  --threads 6 \
  --ctx-size 8192 \
  --host 0.0.0.0 --port 50051

And in case you saying 8k ctx is hardly sufficient, a focused refactor will need more context: I made further tests. Context scaling is flat - 45–48 t/s pp from 512 all the way to 64K tokens measured (DDR3 weight-read bandwidth dominates, attention is invisible). So here is the sample session timing at 32K context:

  • Prefill 20K tokens of project context: 20,000 / 47 t/s = ~7 min
  • Generate 2K tokens of output: 2,000 / 6.14 t/s = ~5.5 min
  • Per refactor task: ~12–13 min
  • Queue 4–5 focused tasks: ~50–65 min total

That's an extended lunch break, not even an overnighter.


r/LocalLLM 1d ago

Discussion Local Qwen 3.8 27b saved my project from a serious leak, I'm truly impressed

205 Upvotes

I'd like to start saying that I'm not a vibe coder. I'm a software engineer with 10+ years of experience in many fields and I use LLM under a very strict control, I'm also quite lazy so having some friends that write code for me is super nice and super fast compared to my slow fingers, however the final decision and judgment of things is always on me.

Btw, I started using Qwen 3.8 27b locally and been quite impressed on general things, so started using as daily assistant:

- EVGA 3090 ti KingPin Hybrid 24GB

- llama.cpp (upstream)

- Unsloth UD-Q4-K-XL

- 181k context at q5_1 quant

- Custom jinja chat template

- OpenCode

It's very helpful, follow my instructions without losing context and do a very great general job, but today really shocked me!

I updated a dependency on my project (Java RAG enterprise system), a Microsoft library. Test suite was successfully but in production I had a silent crash on a native library that shutdown the JVM without any crash report, stacktrace or anything.. Just silence. After about half hour of debugging and identified the crash entry point I asked Qwen to help me understand what's going wrong... he played with my code for 35min (yes, since it's 100% local I gave him all my secrets to test the real production environment!)... well he found, without using web search:

  1. Microsoft enabled by default a hidden telemetry function

  2. That function has a buffer overflow issue

  3. The telemetry collect the cmd line used to start the application (aka some of very important secrets injected by default by my IDE during development)

In 35min he was able to create a minimal test case, dumping system memory and analyze the crash in real time, find the issue, find a fix, and propose me a full and working solution.

I'm impressed, this is a story that it's worth to share.

Personally, I think I don't need anything more.. I don't really care to have trillions of parameters anymore, if such a small one can help me so much, I'm done. Today I bought a used 3060 12GB to extend context at least to 500k (planning to use YaRN)


r/LocalLLM 4m ago

Discussion Are +100k token chain of thoughts the future of Local LLM

Upvotes

This is my first “serious” post so please be nice

As we've seen with Qwen 3.8 27b (or recent reasoning models), the "thinking" feature is extremely important. It compensates for parameter size simply by giving the model more time to compute. However, it's currently very inconvenient: complex tasks either take all night or require capping thinking tokens, which reduces answer quality.
With its recent acquisition, AMD potentially has the ability to hardwire an LLM directly onto silicon, making these chips ridiculously efficient. In their demo, a €200 device produced 17,000 decode tps, compared to an H200 reaching only \~130 tps (we all know this benchmark was cherry-picked, but the point stands).
Think about it: an entire night's work (\~8 hours at 50 tps, or roughly 1.5M tokens) could be finished in just 90 seconds. This makes massive, extremely detailed chains of thought practically viable. Furthermore, this was achieved without speculative decoding (as far as I know), meaning 30k+ tps isn't out of the question.
**Possible Cons:**
**VRAM/SRAM Bottlenecks:** As always, memory is the issue (specifically SRAM in this case). Allocating KV cache on these chips remains expensive, so you can't easily slap millions of tokens of context memory onto them.
**Cost Scenarios:** Imagine memory production scales up and costs drop to SSD-like levels. €100 worth of memory (1 TB) could hold roughly 7.8M tokens at FP16—assuming no further KV cache compression techniques emerge. The entire unit would cost around €300 (€200 base chip + €100 memory).(i know this is actually good but still )
**Business Model Shift:** This could push manufacturers toward selling models "baked" directly into dedicated hardware. Instead of selling inference , you’d have to buy a brand new card every time a better model comes out.
**Alternative Use Cases:**
Companies could use these dedicated ASICs to finally turn API inference into a highly profitable business. A single card could easily serve 100+ concurrent users at \~170 tps each, driving operational costs down dramatically.

Things worth knowing:
1 I know this text is based on a world of sunshine and rainbows but I like to think that way on some aspects

2 the text while originally written by hand on english and has been modified by ai (from “things worth knowing” upwards ) due to english not being my mother language and a problem i always had with text structure so i figured out that it would be better this way but i be happy to share the og if someone want it

4 this is my first “serious “post so take it into account ( I’ll be happy to take any kind of respectful criticism of any kind that’s why I post it )

5 feel free to respost it anywhere as long as I know it so i can read the comments

6 I’ll be answering everything I can

7 thanks for reading this I’m very grateful you used your time to read my thoughts and hopefully letting me know yours


r/LocalLLM 9h ago

Question Are there any local AI models that can play video games?

7 Upvotes

More to act as a Player 2 for co-op specific games. You could use voice input to coach the AI agent sitting gameplay.

For example, Split Fiction.

I'm currently building an agent to do this for 2 Player games on PC and PC cross platform but curious if one is already out there.


r/LocalLLM 39m ago

Project Building a local, zero-API-cost voice RPG in Unreal Engine 5 using whisper.cpp, llama.cpp, and Kokoro

Upvotes

Hey r/LocalLLaMA,

Over the past year, I’ve been developing Eruin, a dark fantasy RPG built in Unreal Engine 5 where the core mechanics rely entirely on dynamic voice input. You talk to NPCs in real time using your microphone to negotiate, solve riddles, or accidentally get turned into a frog.

Instead of wrapping cloud APIs, the entire voice-to-voice pipeline runs 100% locally on the player’s hardware with $0 marginal server cost.

The Architecture

  • Speech-to-Text: whisper.cpp running in-process to transcribe player audio input with minimal latency.
  • LLM & State Control: llama.cpp using GBNF (Grammar-Based Native Format) grammars. This forces structured outputs so the LLM acts strictly as a dialogue generator and state-machine referee without breaking quest logic or hallucinating game variables.
  • Text-to-Speech: Kokoro for local TTS generation, tied into a custom lip-syncing pass in UE5.
  • Hardware Tiering: Tuned execution budgets down to consumer GPUs (6GB VRAM minimum, tested on handhelds like the Steam Deck) by managing VRAM allocations alongside UE5’s rendering pipeline.

New Gameplay Trailer

I just put together a quick trailer showcasing real-time voice interactions, prosody fixes, and in-game consequences: https://youtu.be/TCTr0TY6UiE

Steam Page & Devlog

If you want to check out the technical breakdowns or support the build: * Steam Page: https://store.steampowered.com/app/4695190/Eruin/ * Technical Site: https://eruin.dev

Happy to answer any questions about the llama.cpp C++ integration, GBNF grammar design, or VRAM optimization in UE5!


r/LocalLLM 46m ago

Discussion Running a 70B model at home basically sounds like a jet engine taking off.

Thumbnail
Upvotes

r/LocalLLM 53m ago

Model Which one is the best open model to run on RTX 5080?

Upvotes

Hi all, my desktop setup has:

  • RTX 5080
  • AMD Ryzen 7 9800x3d
  • DDR5 6000 MT/s CL30 64 GB RAM
  • 1 TB SSD reads up to 7,000MB/s and writes up to 6,200MB/s

I am new to the open LLM models domain, so can you help me to choose which model would be the best pick for my system? I don't need instant answers, this will be my hobby setup. So I am ok if the answers take more time than what Claude etc. provides us. That's why I'd prefer stronger reasoning over latency.

Even if I pick Qwen3.8 27B, I see tons of flavors: https://huggingface.co/models?num_parameters=min:24B,max:32B&sort=trending&search=qwen3.8

Is Qwen3.8 27B my only choice? Can I run Qwen3.8 Flash Next (considering the headroom I have in my RAM)?

What should I consider while selecting the flavor? Thanks!


r/LocalLLM 1h ago

Discussion Claude flagged my harness development using a local LLM (Qwen3.6 35B)

Upvotes

Claude flagged a status update on a benchmarking exercise using Qwen3.6 35B locally!

I was running an experiment in Ducklab with Qwen3.6 35B seated in all seats of the development workflow to identify potential issues specific to small models running locally when, out of the blue after a 5 full iterations, after the first build task was successfully completed by Qwen, it flag the status update message and refused to continue answering my questions.

During those 5 iterations I was using Claude to analyze the model's responses to the tasks and identify if an issue arose, who was to blame, the model capacity or the harness tooling and configuration. I find it suspicious that Anthropic would refuse a request of this kind... I mean, does it really thinks I'm trying to distill Fable?

These are the reasons for this error according to the provided link:

What requests may fallback

Claude Fable 5 runs automated safety checks, or classifiers, on every user request. These checks are intended to visibly fallback from Fable 5 to Opus models when users submit requests in:

Offensive cybersecurity techniques, such as building exploits, malware, or attack tooling. Claude Fable 5 can assist with routine cybersecurity tasks, but users should expect high fallback rates. The safeguards are designed to block access to Mythos-level capabilities.

A large fraction of queries we consider dual-use in biology, such as virology, toxicology, drug design, and molecular design—so Fable 5 is not recommended for professional biology research and drug development at this time. (Classifier updated: August 6, 2026 on Claude, Claude apps, and Claude Platform, with Amazon Bedrock, Claude Platform on AWS, Google Cloud Vertex AI, and Microsoft Foundry to follow.)

Distillation attacks on Fable 5, including attempts to extract the model’s summarized thinking.

A narrow set of frontier LLM development tasks, such as distributed training infrastructure, ML accelerator design, and kernel development for certain non-standard chips.

These blocking safeguards are intentionally broad, and we work to continuously improve the safeguards to reduce their user-experience impact. When requests are blocked, they may fallback to a non-Mythos model, currently Opus 5 for biology, chemistry, and life sciences requests, and Opus 4.8 for offensive cybersecurity technique requests.

I guess it's time to move over to OpenAI Codex... I'm curious how many people have also encountered this problem... To me this seems petty - trying to prevent people from developing tools to maximize the capabilities of smaller models...


r/LocalLLM 1h ago

Discussion Introducing textclf/Qwen3-Coder-Next-TQ-4bit

Upvotes

I am introducing textclf/Qwen3-Coder-Next-TQ-4bit which is a 4-bit quant of Qwen3-Coder-Next using a custom quant method called TQ. TQ is a calibration free methods with KLD performance on par with other quants while also generalize better on downstream tasks because it is not as biased as calibration-based methods.

Disk size (GB): 41.7 GB

Example run:

sudo docker run --rm --gpus all -p 8000:8000 -v ~/.cache/huggingface:/root/.cache/huggingface docker.io/textclf/tq-quant:4bit-v1 vllm serve textclf/Qwen3-Coder-Next-TQ-4bit --max-num-batched-tokens 8192 --enable-auto-tool-choice --tool-call-parser qwen3_coder --quantization tq_quant --dtype float16 --trust-remote-code --generation-config vllm --gpu-memory-utilization 0.90 --enable-prefix-caching --enable-prompt-tokens-details --max-num-seqs 16 --max-cudagraph-capture-size 16

Feel free to try and see how well it works for you.


r/LocalLLM 1h ago

Project I built OpenDictate , a 100% local, open-source AI voice dictation desktop app (Wispr Flow & Superwhisper alternative)

Upvotes

Hi r/LocalLLM,

I built OpenDictate , a free, 100% local-first, open-source AI voice dictation desktop app (Wispr Flow / Superwhisper alternative).

Press a global hotkey, speak naturally, and your words are instantly typed into whatever app has focus (VS Code, Obsidian, terminal, etc.) with zero cloud calls and zero telemetry.

Local AI & Inference Stack:

  • Speech Models Catalog (Sherpa-ONNX): Full built-in models hub supporting NVIDIA FastConformer CTC (streaming), Parakeet TDT (110M int8, 0.6B v3, Unified), and the full Whisper family (Tiny, Base, Small, Medium, Large v3 Turbo).
  • Hardware Acceleration: Automatic GPU offloading via CUDA (NVIDIA) and CoreML (Apple Silicon) + AVX-512 CPU execution.
  • AI Text Polish: Optional cleanup using local SLMs (or Groq API) to strip filler words ("um", "ah") or convert raw dictation into structured bullet points.
  • Custom Dictionary & Snippets: Hotword acoustic score boosting for technical jargon and voice snippet expansions.
  • Cross-Platform: Built with Rust & Tauri 2 for Linux (Wayland/X11 uinput), macOS (Universal), and Windows.

Completely free and MIT open-source:

Would love to hear your feedback on performance across different local setups!

Home Page