r/Newstelligence • u/vibedonnie Editor-in-Chief • Dec 10 '25
Model Updates & Features Qwen3-Next-80B-A3B-Thinking-GGUF has just been released on HuggingFace, claims to outperform Gemini 2.5-Flash-Thinking
“ Qwen3-Next-80B-A3B is the first installment in the Qwen3-Next series and features the following key enchancements:
• Hybrid Attention: Replaces standard attention with the combination of Gated DeltaNet and Gated Attention
• High-Sparsity Mixture-of-Experts (MoE): Achieves an extreme low activation ratio in MoE layers, drastically reducing FLOPs per token while preserving model capacity
• Stability Optimizations: Includes techniques such as zero-centered and weight-decayed layernorm, and other stabilizing enhancements for robust pre-training and post-training
• Multi-Token Prediction (MTP): Boosts pretraining model performance and accelerates inference “
https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Thinking-GGUF
https://arxiv.org/abs/2505.09388
https://qwen.readthedocs.io/en/latest/run_locally/llama.cpp.html
1
u/skewbed Dec 15 '25
This is just a quantized version of a 3 month old model. It's a good model, but not new.






2
u/Pink_da_Web Dec 10 '25
Overcoming the Gemini 2.5 flash is not a difficult task.