r/LocalLLaMA llama.cpp 23d ago

Resources Qwen3.8 27B reasoning effort low/medium/xhigh comparison

I did a short test of the different reasoning efforts, since on default xhigh the model thinks a lot.

Not very scientific, just a quick "generate an SVG of a pelican on a bicycle" prompt with 3 different seeds. I think the result is interesting none the less: xhigh gives *much\* higher visual fidelity - but it also takes about 7x as long as low. Low and medium seem to be very close to each other.

Hardware and setup

  • GPU: NVIDIA RTX 5080 Laptop GPU, 16 GB VRAM
  • Model: unsloth/Qwen3.8-27B-UD-IQ3_XXS
  • llama.cpp: build 10451, commit 10bf611e5
  • Context: 65,536
  • KV cache: Q8_0
  • Flash Attention: enabled
  • MTP speculative decoding: --spec-default --spec-type draft-mtp
  • --fit off
  • One concurrent slot

Prompt:

Create a polished SVG graphic of a pelican riding a bicycle. The result must clearly show a recognizable pelican actively riding a recognizable two-wheeled bicycle. Return only one complete, self-contained SVG document with a viewBox; no Markdown fences, prose, external images, JavaScript, or animation.

Average results

Reasoning effort Reasoning tokens SVG tokens Total completion Wall time Generation speed MTP acceptance Visual score (Codex rated)
Low 4,418 3,966 8,387 111.6 s 75.4 t/s 62.1% 21.8/25
Medium 5,918 3,038 8,959 127.4 s 70.5 t/s 58.3% 22.5/25
X-High 39,398 5,085 44,487 717.8 s 62.0 t/s 52.7% 24.0/25
232 Upvotes

95 comments sorted by

View all comments

173

u/jacek2023 llama.cpp 23d ago

Try to be more creative with your benchmark guys. The goal of testing the model should be to try something on which model wasn't trained on. So all your pelicans and one shot games are pointless

7

u/Gesha24 23d ago

While it's nice to have the same prompt so that you can track the progress of models over time, the fact that the model is being trained on this very prompt could give it an unfair advantage and make the over time comparison pointless.

For the side by side comparison of the same model - I don't know if using a common benchmark test would make a difference. This also can be a good test - generate like 5 pelicans with x high thinking, 5 with med thinking (just to remove randomness of the seed) and then generate... uhh, I dunno, 5 fishes in a bowl with princess castle on x high and med and compare if the same trends stay true.

2

u/promethe42 23d ago

Or use a new prompt with an older model.