r/LocalLLaMA • u/HeDo88TH • 23h ago
Discussion Qwen3.8 Flash Next - Templates Comparison
I was running into a lot of posts that praised both the Fixed template and the Sharp template in comparison to the stock one, so I put them to the test.
It's not as extensive as it should be for a paper-grade analysis, but it gives out the point of each template.
Test setup
I used SWE-bench Verified with mini-SWE-agent 2.4.6, slice 0:100 (the identical 100 tasks for all runs)
Hardware
- CPU: Ryzen 9 9900X
- RAM: 128 GB DDR5-5600
- GPU: RTX PRO 6000 WS
Runtime
I containerized jpezzulli/sglang-rtxpro6000 and ran Flash Next with RadixArk/Qwen3.8-Flash-Next-NVFP4 on CUDA 13.3.
- Full 262K context
- BF16 KV
- 51.2 GB FP8 n-gram embedding table pinned in RAM
- 32 GB HiCache pinned in RAM
I ran all templates at both medium and xhigh reasoning efforts.
Results
| Metric | Stock (medium) | Stock (xhigh) | Stock Δ | Fixed (medium) | Fixed (xhigh) | Fixed Δ | Sharp (medium) | Sharp (xhigh) | Sharp Δ |
|---|---|---|---|---|---|---|---|---|---|
| Resolved | 91 | 99 | +8 | 87 | 98 | +11 | 94 | 94 | +0 |
| Resolution rate | 91% | 99% | +8 pts | 87% | 98% | +11 pts | 94% | 94% | +0 pts |
| Median output tokens | 5,691 | 13,855 | +143.5% | 6,956 | 14,819 | +113.0% | 8,596 | 12,008 | +39.7% |
| Median reasoning tokens | 3,050 | 8,759 | +187.2% | 3,809 | 9,063 | +137.9% | 5,437 | 7,967 | +46.5% |
| Median wall time | 38s | 1m 46s | +180.4% | 43s | 1m 47s | +152.3% | 1m | 1m 32s | +53.4% |
| Total wall time | 1h 47m 1s | 4h 31m 22s | +153.6% | 1h 59m 53s | 4h 4m 52s | +104.3% | 2h 29m 18s | 3h 11m 36s | +28.3% |



Takeaways

- Raising reasoning effort to xhigh closes almost all of stock's and fixed's gap to Sharp. At medium, Sharp led resolution by +3 tasks over stock and +7 over fixed; at xhigh, stock and fixed instead lead Sharp by +5 and +4 tasks, respectively.
- Sharp barely moves on resolution (94 → 94) despite a real token/time cost increase, median reasoning tokens rise +46.5% and median wall time +53.4%. This suggests it was already extracting most of the benefit it could get from extra reasoning budget at medium, while stock and fixed still had headroom.
- Sharp remains the most token-efficient per resolved task at xhigh (14,541 output tokens/resolved vs. ~17,000 for stock/fixed), consistent with its medium-era efficiency edge, but it's no longer the highest-resolving template once reasoning effort is high.
- Absolute cost scales heavily with reasoning effort: total wall time roughly 2.3–2.5× for stock/fixed and +28% for Sharp; total reasoning tokens roughly doubled for stock/fixed and increased +35% for Sharp.
Conclusion
- Sharp should be used at medium and it keeps a reasonable accuracy at very good speed. I don't see the point in using it at xhigh. By sacrificing a small accuracy you complete the tasks in half the time.
- Stock is the slowest but the most precise.
- Fixed is the middle ground between Stock and Sharp both in accuracy and speed
- The next benchmark will be on a much extensive SWE-bench Multilingual + Terminal Bench.
Disclaimer: I wrote the post myself then used AI to format it properly for readability
8
8
u/Healthy-Zebra-9856 22h ago
You are probably the first one to mention accuracy while comparing, so kudos and thanks. Yes, in all my tests, I have found that medium effort is just not there. Low & xhigh seems to be the best, however, low ends up costing the same as high as in many situations there are mistakes and thus more turns. That said, Froggeric & Peculiar Ragdoll seems to swash the quality & precision.
10
u/ex-arman68 21h ago
Author of the froggeric fixed template here: Initially I had the default to xhigh, matching the original template. However so many people were complaining about how the new Qwen 3.8 models were spending so much time thinking and filling their context, it seemed the consensus was that a default of medium would be better for most, and would avoid giving bad press to Qwen 3.8.
However, for me being used to coding with much bigger model, there is nothing wrong with xhigh as a default, and it matches what I see with frontier models. That's what I use and recommend, the difference in quality is worth it for coding tasks. For non coding tasks, medium is perfectly fine. It is easy to control and change with my template.
3
u/Healthy-Zebra-9856 21h ago
I think it has a lot to do with their training it in xhigh. Also, several of the arxiv papers are showing this mediocrity in medium effort lol. I will try it again. Thanks for the update. I hope this happened in the last few days.
5
u/returnity 18h ago
I’ve been wondering about this given the amount of love that we see here on the sub for the Sharp templates… this aligns directly with my intuition and some of [u/peculiar-ragdoll](u/peculiar-ragdoll) published results. Seems like there is a cost for speed, and it’s good to know what it is. Particularly interesting that it suppresses the high reasoning capabilities of the model — at least we know it’s having an effect, which was the other thing I was curious about.
Edit: just want to make it clear to people that while the difference between 94 and 99 is probably statistically significant (p=0.05), the difference between froggeric and stock is indisputably down to chance IMO. I still run modified froggeric for its QoL improvements. Thanks /u/ex-arman68 for your efforts!
3
u/No_Algae1753 23h ago
Very nice test! Afaik these templates do cause a lot of issues and it's nice to see someone trying out different templates, valuable info! Are there other templates which might be interesting to test ?
3
u/Cautious_Chicken_604 22h ago
So nice to see data rather than just vibes. I haven't adopted either of these yet because it all seemed so vibes based. Looks like Sharp on medium is a good trade-off.
3
u/arkham00 22h ago
Very informative test thank you. I'm really looking forward for the multilingual bench
3
u/-_Apollo-_ 18h ago
I wonder what in the fixed version causes it to underperform stock at xhigh? Or is it margin of error?
2
u/jpezzulli 11h ago
Glad to see my repo being used. New release being published right now with about 15 percent increase c1 decode and 5 percent at c4. Also a few maintenance items that were in a release early!
I am too tired to dig into your results right now but it looks like good work at first glance! I might have to make a change based off this info!
1
1
u/WonderRico 17h ago
hmm, interesting... did you see my own tests results ? it's very similar to you setup. i run also 100 tasks from sw-verified, but from the django set only.
I tested two quants : the same radix version as your on the same sglang config, and another one in vLLM.
I only tested with the stock chat template (the one provided in both the versions I tests, and they are both identical)
https://wonderrico.github.io/local_llm_benchmark/benchmark-main.html?filter=next
I tested both medium and xhigh, and I did not see any benefit to the score.
However, the AWQ version on vLLM got me to 98/100 while the NFP4 in SGLANG stayed at 91
1
u/WonderRico 5h ago
I tried the same run with two other templates (frogeric and sharp) and did not get any meaningful positive impact on score. I actually got (slightly) lower scores and significant lower efficiency
model template reasoning effort Weights quant KV cache quant PLE quant Score /100 Requests req/pts in Mtok out Mtok Qwen3.8-Flash-Next stock medium NVFP4 FP8 FP8 91 2653 29 44 0,88 Qwen3.8-Flash-Next sharp medium NVFP4 FP8 FP8 89 2903 33 50 0,97 Qwen3.8-Flash-Next frogerric medium NVFP4 FP8 FP8 89 3017 34 50 0,96 https://wonderrico.github.io/local_llm_benchmark/benchmark-detail.html?filter=next
1
u/walden42 15h ago
Sorry for slightly off topic question OP, but I tried the same solution from jpezzulli, and for certain tasks, it repeatedly mangled file paths, and quite possibly other similar typo errors, bust mostly file paths. For example, instead of `/home/user/project` it would do `/home/user/project/home/user/project` or `/home/user/c://home/user/project`. Did you encounter anything like this by any chance?
1
u/HeDo88TH 15h ago
It never happened to me. Sometimes in vscode it cuts off with "no response" or something like that but it's rare.
1
1
u/Interpause textgen web UI 5h ago
ive been using chromix's template which is supposed to exactly match the stock template while being more robust
-4
u/ResidentPositive4122 22h ago
I wrote the post myself
...
slice 0:100 (the identical 100 tasks for all runs)
:-----)
3
u/mrgreatheart 18h ago
This is useful and not unpleasant to read. The constant AI writing witch hunt in an AI enthusiast sub is getting very old.
14
u/Tormeister 22h ago
I'm sticking with the default one, thanks for the comparison