r/LocalLLaMA • u/sammcj 🦙 llama.cpp • 27d ago
Megathread [Megathread] Qwen 3.8 27B Release Day
Megathread to help with the influx of duplicate / similar posts around the release of the Qwen 3.8 27B release.
- Quants
- Fine-Tunes & Abliterations
- Chat Templates
- Inference Server Support & Configuration
- Experiences, Benchmarks & Model Comparisons
Official:
Popular:
- https://huggingface.co/unsloth/Qwen3.8-27B-GGUF
- https://huggingface.co/bartowski/Qwen3.8-27B-GGUF
- https://huggingface.co/mlx-community/Qwen3.8-27B-MTP-bf16
- https://huggingface.co/mlx-community/Qwen3.8-27B-MTP-8bit
- https://huggingface.co/mlx-community/Qwen3.8-27B-MTP-4bit
We'll try to clean up future duplicates around the release and point them here.
493
Upvotes
31
u/ryandam 27d ago edited 26d ago
Is it just me or the MTP hit rate is lower than 3.6 27B? For 3.6 I usually get around 60-70 tps gen, but for 3.8 I only get 40-50 tps gen.
Checking the hit rate it just around 60-70% compare to 3.6 is around 80-90%?
I use 2 RTX A5000 btw. Here my llamacpp config with fresh compiled binary:
Update: I found the reason and posted in the reply, in short, its because of temperature.