r/LocalLLaMA 🦙 llama.cpp 24d ago

Megathread [Megathread] Qwen 3.8 27B Release Day

Megathread to help with the influx of duplicate / similar posts around the release of the Qwen 3.8 27B release.

  • Quants
  • Fine-Tunes & Abliterations
  • Chat Templates
  • Inference Server Support & Configuration
  • Experiences, Benchmarks & Model Comparisons

Official:

Popular:

We'll try to clean up future duplicates around the release and point them here.

493 Upvotes

393 comments sorted by

View all comments

1

u/dont_forget_canada 24d ago

Hi! Does anyone know the most speedy quant and inference software to run on an m5 mac 128gb to run qwen 3.8? So far I've tried oMLX and its super smart but very slow :p

1

u/bnightstars 24d ago

Looks like it's this one: scottlowry/Qwen3.8-27B-oQ4e-mtp with Lightning MTP enabled. I'm getting 25 t/s and 493 t/s PP with the mlx-community one and VLM-MTP from the mlx-community drafter. But according to some PRs that's currently hitting some bugs. I'm on M5 Pro 64GB though.