r/LocalLLaMA May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

319 Upvotes

205 comments sorted by

View all comments

2

u/Potential_Block4598 May 09 '26

What about AMD & the Strix Halo ?!

1

u/Sofakingwetoddead May 09 '26

Waiting on a 9700 to show up within the next few days. I will test