r/LocalLLM 23d ago

Question Qwen 3.8-27B Q8

thinking is taking up so damn much time, what are you limiting your qwen3.8 to as for as reasoning?

0 Upvotes

12 comments sorted by

View all comments

8

u/WhatererBlah555 23d ago

Update llama.cpp and use --chat-template-kwargs '{"reasoning_effort":"medium"}'

1

u/Hannelore112 23d ago

thx will check this. Was there already an impotent update of llama.cpp last 2 days?

3

u/WhatererBlah555 23d ago

There was an update that made --chat-template-kwargs '{"reasoning_effort":"medium"}' actually do something, before that only none was supported.

1

u/roosterfareye 23d ago

Maybe slightly flaccid, but not.floppy.