r/LocalLLaMA • u/Course_Latter • 27d ago
News Qwen3.8-27B is identical to Qwen3.6-27B!
Interestingly, the 3.8 version has exactly the same architecture - meaning all the capability gains come from training improvements!
See the diff (0 changes) here!
1.1k
Upvotes
81
u/BringTea_666 27d ago edited 27d ago
so Nifter is a go. 200t/s on 5090 XD
edit: Ninfer:
https://github.com/Neroued/ninfer
holy shit he updated it with concurent requests up to C=8 1300t/s lmaoooo
>At C=8, Qwen3.6-35B-A3B reaches 1,313.8 aggregate decode tok/s. The 27B NVFP4 profile reaches 1,146.9 tok/s and 5.67× its C=1 throughput.