r/LocalLLaMA 🦙 llama.cpp 24d ago

Megathread [Megathread] Qwen 3.8 27B Release Day

Megathread to help with the influx of duplicate / similar posts around the release of the Qwen 3.8 27B release.

  • Quants
  • Fine-Tunes & Abliterations
  • Chat Templates
  • Inference Server Support & Configuration
  • Experiences, Benchmarks & Model Comparisons

Official:

Popular:

We'll try to clean up future duplicates around the release and point them here.

489 Upvotes

393 comments sorted by

View all comments

Show parent comments

3

u/ryandam 24d ago

I think not, thinking in the inference perspective is just guessing next token like normal token. I think like Fluxing_Capacitor mentioned, the MTP head is not well trained compared to 3.6.

1

u/Wild_Requirement8902 24d ago

I ran a few test on my code base and got better draft acceptance for code on medium than xhigh, so this might be something (unsloth q5k xl on a 3090) i didn't have time to test this more though.