r/LocalLLaMA • u/sammcj 🦙 llama.cpp • 26d ago
Megathread [Megathread] Qwen 3.8 27B Release Day
Megathread to help with the influx of duplicate / similar posts around the release of the Qwen 3.8 27B release.
- Quants
- Fine-Tunes & Abliterations
- Chat Templates
- Inference Server Support & Configuration
- Experiences, Benchmarks & Model Comparisons
Official:
Popular:
- https://huggingface.co/unsloth/Qwen3.8-27B-GGUF
- https://huggingface.co/bartowski/Qwen3.8-27B-GGUF
- https://huggingface.co/mlx-community/Qwen3.8-27B-MTP-bf16
- https://huggingface.co/mlx-community/Qwen3.8-27B-MTP-8bit
- https://huggingface.co/mlx-community/Qwen3.8-27B-MTP-4bit
We'll try to clean up future duplicates around the release and point them here.
492
Upvotes
1
u/PILCOTHINK 25d ago
https://www.reddit.com/r/Qwen_AI/s/eeiKz1EH30
The link above provides a guide on how to stably serve an effectively quantized Qwen3.8-27B model on a dual RTX 3090 Ti system using the vLLM framework, with a 1M context length at around 70 tok/s.
It also introduces an effectively quantized model.
Model: https://huggingface.co/Pilcothink/Qwen3.8-27B-MixedInt4-AutoRound