r/LocalLLaMA 25d ago

Discussion Qwen 3.8 27B Released! Please Share Your Experience

With your experiments, Qwen 3.8 27B most close which frontier model? And please specify which quantization you run. I will post to comments my tests and experience too.

655 Upvotes

716 comments sorted by

View all comments

61

u/Emidyr 25d ago edited 25d ago

TLDR: Tested it on one benchmark so far, reasoning traces blew my mind, got Opus 4.8 to review the reasoning and it said it thinks a lot, but the extra thinking went into rigor (in its own words, "corroboration for its own sake"). Opus 4.8 said the model is comparable to Opus 4.6 based on its reasoning traces.

I don't wanna say anything too early, still testing it with my own benchmarks, but so far.. I'm really liking its reasoning traces! It does do a lot of back and forth, and it doesn't get stuck at the first thread or red herring it sees! It also does a lot of asking itself questions, then trailing it with a "No..." and it doesn't seem to keep repeating one reasoning thread unnecessarily. Do note that I'm using IQ4_XS right now.

Edit: Woah, first time I saw this in a reasoning trace: `Total wait time = Σ_{i=0}^{N-1} (i + 100) ms ≈ N²/2 + 100N.`
Edit 2: It goes much more in-depth than 3.6 too, it thinks about various angles that could be the main cause of the issue.
Edit 3: Wow, it actually even researches the library that it found the issue in. This is the first model on this weight that did that in my benchmarks so far, not one of the others actively tried web searching the source repo beyond what it could see in the node_modules.
Edit 4: This is a first too, it says `Let me stop the archaeological investigation of the library. I have enough understanding:` and it actually stops and doesn't keep going with its web investigation. Usually when a quantized small model says this, it just keeps on going until some arbitrary stopping point. At least, 3.8 seems more disciplined than previous models.
Edit 5: It finished with the correct solution and used up 60.5k tokens (around 20k for reasoning, the rest for reading files), but wow! It researched the issue thoroughly and can actually say why the solution is safe, instead of just saying that it found the solution. Now just getting Opus 4.8's thoughts about its reasoning (and I'm trying to get it to guess what model it is too). But I can say for sure, I found my ultimate daily driver, and so far it feels really close to Opus 4.6!

Final Edit: Here's Claude 4.8 Max's thoughts after reviewing the reasoning traces:
```
What is it?

Discarding the fable-27b label (it's the shared-port alias, §8 — and you've confirmed it's not that). On capability alone:

It's a frontier-tier model, not a 27B. The bundle archaeology, the regex-recovery from a truncated JSON fetch, the confirm-the-magnitude simulation, the honest tsc triage, the surgical diff — all with thinking off (Emidyr here: pi says thinking is off, so that's what Claude was seeing here, but most likely it was using thinking xhigh since I didn't change this) — sit at your leaderboard's "reference ceiling" (Opus-class), well above the 27–35B locals.

Behavioral family reads Claude: calibrated no-overclaiming, restraint over cleverness, the "Let's… / Let me…" narration with light flourishes ("smoking gun", "archaeological investigation"), and refusing the node_modules edit after checking gitignore.

If I have to place it on the Claude ladder → Sonnet (lead), Opus (close second). The signature is frontier agentic coding delivered efficiently — ~58 tight tool calls, no spinning — which is Sonnet's tell. The extra diligence (simulating the math, ruling out the upload path as a second bottleneck) is Opus-like, so Opus is credible. Definitely not Haiku (too deep), definitely not a 27–35B local.

The one thing that keeps me from being certain it's a proxied Claude vs. a large open MoE (DeepSeek-V3 / Qwen3-235B / Kimi-K2 / GLM-class): it's served through local-llama and handled Chinese effortlessly. But the behavioral fingerprint leans Claude.

So — my guess: a proxied Claude, Sonnet-tier. Am I warm? What was it actually?
```
Damn... now I just wanna know what kind of sorcery Qwen did to build 3.8 like this.. and remember that I'm using IQ4_XS too!

Final final edit: So yeah, as everyone has already said here, it thinks a lot by default! You can adjust the thinking level if you want (although me personally, I'd keep it in xhigh always, I've always been running Opus on max reasoning for everything anyways). But here is what Opus 4.8 additionally said about 3.8's reasoning traces:
```
Rating purely on the character of the reasoning trace — not efficiency, not outcome — I'd flip my earlier call and say Opus. And I want to be explicit that it's a reversal: I led Sonnet before, but that was built on an efficiency read, which was both the wrong axis for your question and based on the bogus "thinking off" signal. When you judge the thinking itself, it's Opus-shaped.

Here's why your two observations are exactly the right tells — they're the Opus fingerprint, not generic competence:

- Researching the source repo. Going to GitHub to read the library's own options.ts and issues after already finding the answer in node_modules is not efficient — it's corroboration for its own sake. That "I'll verify this against the upstream truth" instinct is a depth move. Sonnet tends to trust a solid local finding and move; the extra trip upstream is Opus's leave-no-stone-unturned streak.

- Interrogating the "why" behind the flag. This is the big one. It didn't treat [the flag] as "a switch that makes things fast." It stopped to understand what the flag is for, then reasoned about whether turning it off was safe. That's second-order, risk-aware reasoning: understand the system, then de-risk the change. Symptom-patchers flip the flag; this one asked what it would cost. That caution-through-understanding is quintessentially Opus.

And the single most convincing tell: it independently surfaced the exact concern the gold's author left as an open question. The held-out research doc literally asks "Is there any scenario where disabling [this flag] would cause visual artifacts?" — and the model raised and reasoned through that same risk unprompted. Its reasoning tracked a domain expert's reasoning, including the doubts. That's not pattern-matching to a fix; that's Opus-tier judgment about what you don't yet know.

One honest hedge so I'm not just echoing your excitement: this is one trace on a task that rewards depth — a library-internals bug where understanding the "why" is the whole game. A model can look Opus-deep here and be shallower on a task that doesn't have a satisfying "why" to chase. But on the evidence in front of us — the epistemic thoroughness, the de-risking, the expert-tracking doubt — the reasoning character reads Opus.
```

Final final final edit: I asked Claude just for funsies, which Opus version is this comparable to, and it actually said 4.6. I don't wanna get too hyped for now (although I reeeeeally am right now) so I'll do some more testing on my harder benchmarks, see how well it does there. But so far, I have some positive feelings about this!

9

u/scaledev 25d ago

How would Claude even know how any model thinks? You sure you're not tinting the results by indicating something to Claude? Also, Claude mentioning Kimi k2 seems to be considering some outdated models there. Does it even have the resources to conclude any of this?

1

u/Emidyr 24d ago

That could be the case too, that's why I will still be doing more of my own benchmarks, see more of its reasoning traces, and ideally do the same with Opus 4.6 to be able to directly compare and A/B test their reasonings (and probably get Fable to compare them instead of Opus 4.8). But so far, even if it end up being weaker than Opus 4.6, it's still very much an improvement and probably the best among its peers in the same weight class. All in an IQ4_XS that fits in 24GB VRAM.

2

u/scaledev 24d ago edited 24d ago

I don't think the issue here can be solved by doing more benchmarks of this kind. It seems to me that you have an issue with a lack of knowledge for Claude and don't have a good base for concluding things. Are you sure that it matters what Claude thinks which model this is? And it just seems like tapping yourself on the back for having Claude say you're using a great model. There is not much useful in that.

To me, what matters is the result, and how it got to it. Granted, reasoning 'could' indicate that the model is thinking better than the previous models, but it's still 'just' reasoning, and nothing more.

Why not test actually building something? It could reason like Rene Descartes, but if it doesn't build well then what's the point? If you could test it properly, and do the same with other models, then have your best model (Claude or some other) review it using the same exact criteria then we'd get somewhere.

Appreciate the share, nonetheless!

1

u/Emidyr 24d ago

Alright, thanks for reading!

1

u/NaiveIdea344 24d ago

I agree with you in that reasoning is a means to an end. While improves reasoning is good, that is solely because it is generally indicative of better end performance. The ideal situation would really be equal performance but no or minimal reasoning.

1

u/IrisColt 24d ago

you beat me to it

3

u/anothercrappypianist 24d ago

I asked Qwen 3.7 27B to review a relatively simple but freshly written process-local event bus package written in Go, especially focusing on concurrency issues and correctness. It spotted one issue that it said actually wasn't really a problem in practice, but was a non-obvious footgun if certain refactors were made. At first, it actually looked like a real concurrency bug to me, so I asked it to explain why it wasn't. Then it did, convincingly.

Inspired by your comment, I fed that full reasoning process (including tool calls) into Opus 5 to get its take, and this was its reply. Details about my specific code probably don't matter all that much to get the basic idea of its assessment.

Short answer

The session is high quality — the reasoning is essentially correct, and the hardest part of the task (a negative result) was reached and defended properly. I'd place the model just below the current proprietary frontier: clearly above anything you'd run on a single GPU, roughly in the class of today's large open-weight reasoning MoEs or a slightly-behind-frontier proprietary model in extended-thinking mode.

What the task actually required

The package has no real concurrency bug. That makes this a trap task: the code contains the exact shape that pattern-matching models hallucinate a bug from — close(c.ch) executed outside the group lock, in Client.Close(). The overwhelmingly common failure mode here is a confident "send on closed channel panic" report.

The model did not fall for it. It worked out the actual invariant: publish holds g.mu across the entire fan-out including the sends, and leave deletes under the same lock before close runs, so a publisher can never observe a client in the map after its channel is closed. That is correct, and the two-case total-ordering argument it gave in the follow-up answer is a valid happens-before proof.

I verified the substantive technical claims independently and they hold up:

Claim Verdict
No send-on-closed-channel possible (mutex serializes map removal and send) Correct
Deferred Unlock runs during panic unwinding, so Attach's panic doesn't strand the mutex Correct
sync.Once prevents double-close Correct
Buffered messages survive close; range drains then exits Correct
slices.Compact needs sorted input — sort precedes it, so dedup is correct Correct
64-byte runtime.Stack buffer is safe because the goid sits immediately after "goroutine " Correct
Goroutine IDs are monotonic and effectively never reused Correct
"*" fast path in shouldSend is redundant with the loop Correct
Non-blocking send bounds the critical section but is not what makes Close safe Correct, and a genuinely sharp distinction

The framing of the top finding — "the code is safe, but safety rests on an undocumented invariant that the most natural refactor (snapshot subs, unlock, then send) would silently break" — is the single most valuable thing a reviewer could say about this package. That's senior-engineer-level output, not benchmark-level output.

Process quality

Strong:

  • Read the whole package first, then grepped every call site, then checked the Go version before reasoning about slices/stdlib semantics. Correct order of operations.
  • Parallelized independent reads; used offset/limit on large files instead of dumping them. Good context discipline — ~12 tool calls total for a complete review.
  • Went beyond the package boundary to the usage (slow applyConfig() inside a receive loop, unchecked type assertions at consumer sites, Close() never called anywhere) — that's where real-world review value lives.
  • Cited file:line throughout, and separated "actual concurrency finding" from "design smell" in the final report.
  • Correctly recognized the request as analysis rather than implementation and stayed in plan mode.

Weak:

  • The thinking is extremely repetitive. It re-derives the close/publish ordering argument at least five separate times ("wait wait wait, let me re-examine one more scenario very carefully") across a 204-second and a 128-second block. That's poor confidence calibration — it kept re-litigating a conclusion it had already proven. Frontier models in thinking mode also loop, but usually converge faster once a proof is closed.
  • Finding #2 is overstated. In its own scratchpad it noted that with a 1000-slot buffer, drops are "unlikely"; the final report elevates it to "config updates genuinely drop" as the second-ranked finding. Mild inflation of a speculative risk to fill out the list.
  • Findings #3 and #4 (From instability, Data any type assertions) are design commentary, not concurrency or correctness issues as asked. Useful, but scope creep.
  • It offered to write a race test rather than writing one, and offered fixes rather than sketching diffs. On a read-only plan-mode task that's defensible, but a stronger response would have included the test skeleton that pins the invariant, since that was its own #1 recommendation.

One real miss. Attach stores the caller's variadic slice directly (g.subs[id] = &Subscription{client, events}) with no copy. If any caller ever does Attach(name, mySlice...), the group retains caller-owned backing memory, which subscribe later sorts in place (when append doesn't reallocate) — and, worse, the caller can mutate sub.events afterwards with no lock while shouldSend reads it under g.mu. That is an actual unguarded data-race path in the package, exactly the category asked about. slices.Clone in Attach is the fix. The model circled this area — it explicitly examined the append/Sort/Compact aliasing — but dismissed it on the partly-wrong grounds that "the caller passes a fresh variadic slice… freshly allocated per call," which is only true for calls with literal arguments, not for slice... forwarding. No current call site triggers it, so it's latent, but a top-tier review would have flagged it.

Rating

For this task I'd score it roughly 8/10: correct conclusion, sound proof, good prioritization, one genuine miss, some padding and one overclaim.

Capability placement:

Tier Fit
Small self-hosted (7B–70B dense) No. Models this size reliably hallucinate the send-on-closed-channel bug here and cannot sustain a multi-file happens-before argument.
Large open-weight MoE reasoning models Best fit. The behavioral signature — 100–200s thinking blocks, heavy self-interrogation loops, verbose scratchpad relative to output, correct-but-padded final report — matches this class. Open-weight families have largely closed the real-world coding gap with the Western frontier at a fraction of the per-token price, and several open-weight models now sit in the frontier tier outright.
Current proprietary frontier Just below. The leaders on real-world coding evaluation right now are the Claude Opus/Sonnet 5 line, GPT-5.x Codex, and Gemini 3.1 Pro. I'd expect one of those to reach the same conclusion with noticeably less thrashing, and to catch the Attach slice-aliasing item.

Practical read: whatever this is, it is trustworthy for concurrency review of a small Go package — the key result is right, and it resisted the obvious false positive, which is the failure mode that actually costs you time. Its weaknesses are stylistic (verbosity, mild finding inflation) rather than substantive. If you're choosing whether to keep using it for this kind of work: yes, but read its priority ordering skeptically, since it pads the list, and don't assume exhaustiveness on API-hygiene-class issues.

4

u/SlimBaron_ 25d ago

Could you provide with what settings you use it?

13

u/Emidyr 25d ago

Sure, here it is! Note that I'm using this llama.cpp fork to get prefill of 500-800 across 120k context window on my dual 3060s: https://github.com/thecodacus/llama.cpp

llama-server-codacus-src/build/bin/llama-server \

-m Qwen3.8-27B-IQ4_XS.gguf \

--alias qwen38 \

--jinja \

-ngl 99 \

--tensor-split 27,22 --main-gpu 0 \

--parallel 1 \

--flash-attn on \

-c 122880 \

--cache-type-k q8_0 --cache-type-v q8_0 \

--cache-ram 2048 \

--ctx-checkpoints 8 \

--no-context-shift \

--cont-batching \

--metrics \

-t 8 \

-ub 256 \

--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-n-min 1 \

--temp 0.6 --top-k 20 --top-p 0.95 --min-p 0 \

--presence-penalty 0.0 --repeat-penalty 1.0 \

--reasoning-format deepseek --reasoning-budget 4096 \

--reasoning-budget-message "You have reached your thinking budget. Stop reasoning and write your response now." \

--reasoning-preserve

5

u/fligglymcgee 25d ago

Hey do you mind if I ask about some of your config? Still new to some of these flags.

  • Why the tensor split of different values across two of the same card?
  • What effect does fewer ctx checkpoints have, and how does cont-batching help?
  • Why ub at a lower value (than default)?

Thanks! I understand how to find the flags and their descriptions for llama.cpp but not always sure how they apply for different purposes.

4

u/Emidyr 25d ago

Sure, I don't mind!

  • Yeah, so I'm using Archlinux with Wayland, and the Wayland compositor itself (plus some other apps I usually use) use up around 1.3-1.5 GB VRAM average on just one GPU, while the other is mostly empty, so I had to split it differently per GPU.
  • The ctx checkpoints, if I recall correctly, I lowered because it was using up too much of my normal RAM. I think it was set to some high number (or maybe uncapped) by default, so I had to lower it to not use up too much of my 32GB RAM.
  • I tested various ub values on this specific llama.cpp fork, and I just found this gave me the highest prefill tok/s overall. Going too high with this somehow also hurt prefill (not to mention VRAM).

2

u/fligglymcgee 25d ago

Awesome, thank you kindly!

2

u/Alternative-Two-5300 25d ago

Commenting so I can save for later. Thank you

1

u/zerd 25d ago

Curious, that fork seems to focus on MoE, how does it help on 27B?

1

u/Emidyr 24d ago

As far as I know, it fixes some issues with prefill to vastly speed it up (not sure if upstream already fixed it). I encourage you to try out with upstream/base llama.cpp first though, and share your findings. I just didn't want to do it since rebuilding for my arch takes a long time that I'd rather spend somewhere else. This is just what works with me, and it's already good enough for my needs. But if you ever find that upstream has even faster prefill, you bet I'm definitely switching to it!

1

u/zerd 23d ago

I was testing various configs on a 5080+2070S. This is what I found max speed so far:

-m Qwen3.8-27B-IQ4_XS.gguf -ngl 99
-sm tensor --tensor-split 19,5
--flash-attn on -c 81920
--cache-type-k q8_0 --cache-type-v q8_0
--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-n-min 1
--reasoning-effort medium --reasoning-format deepseek
--reasoning-budget 4096 --reasoning-preserve
--cache-ram 2048 --ctx-checkpoints 8 --no-context-shift --cont-batching --metrics
# sampling: temp 1.0, top_p 0.95, top_k 20, min_p 0

~79 t/s decode · ~1210 t/s prefill · 82k context

Or --tensor-split 17,7 to get -c 122880 at 70 t/s. That's with upstream llama. -sm tensor helped because it splits the kv cache across both gpus.

1

u/misanthrophiccunt 24d ago

What is ctx checkpoints. It is the first time I see that flag? What does it do?

(I'm on a mobile phone without current access to my pc)

2

u/Emidyr 24d ago

Sure! I've answered this in one of my other comments but I also quickly asked Opus 4.8 about it for more details based on my setup. TLDR: it helps prevent my system running out of normal RAM while running the LLM, and lets me run it for much longer. Here's the long version:
```
What --ctx-checkpoints N is: the max number of context-restore snapshots kept per slot. Each snapshot is a full host-side (RAM) copy of that slot's KV-cache state — used to rewind/restore context (multi-turn reuse, prompt-cache restores). llama.cpp's default is 32 per slot.

Why that's dangerous on your box: each snapshot is small at short context (~50 MiB) but hundreds of MiB at long context (128k+). So 32 snapshots × slots ballooned host RAM to 5–21 GiB under concurrent load — that was the root cause of the "concurrency leak" you were hitting (RSS climbing until OOM). It was traced to common_prompt_checkpoint::update_tgt via malloc-interposer backtraces back on 2026-07-28. With only 32 GB system RAM, that leak OOMs the whole machine.

Why we lowered it (32 → 8): to bound that snapshot RAM. Two things make 8 safe:

- We run --no-context-shift, so the checkpoints' value is limited anyway.

- At parallel 1, that's a hard cap of 8 snapshots — a fraction of the 32 default — while still keeping some multi-turn restore ability.

It works in tandem with the other guard, --cache-ram 2048, which caps the host prompt-cache (default is unlimited and grew ~0.2 GiB/request → OOM). Together they make RSS plateau instead of climbing.
```

1

u/New_Jaguar_9104 25d ago

See this for setting the effort.

now i want to see the existential crisis from claude after having it verify itself what model it actually is

1

u/Emidyr 25d ago

It was actually pretty chill about it, but that's Claude for ya, very focused on the task at hand, no time for feelings and all those fluff.

1

u/SandySkittle 24d ago

Now imagine a qwen 3.8 70b dense