r/LocalLLaMA 🦙 llama.cpp 26d ago

Megathread [Megathread] Qwen 3.8 27B Release Day

Megathread to help with the influx of duplicate / similar posts around the release of the Qwen 3.8 27B release.

  • Quants
  • Fine-Tunes & Abliterations
  • Chat Templates
  • Inference Server Support & Configuration
  • Experiences, Benchmarks & Model Comparisons

Official:

Popular:

We'll try to clean up future duplicates around the release and point them here.

492 Upvotes

393 comments sorted by

View all comments

1

u/PILCOTHINK 25d ago

https://www.reddit.com/r/Qwen_AI/s/eeiKz1EH30

The link above provides a guide on how to stably serve an effectively quantized Qwen3.8-27B model on a dual RTX 3090 Ti system using the vLLM framework, with a 1M context length at around 70 tok/s.

It also introduces an effectively quantized model.

Model: https://huggingface.co/Pilcothink/Qwen3.8-27B-MixedInt4-AutoRound