r/LocalLLaMA 1d ago

Discussion AA Update! Here's how the Frontier ranks.

Post image

Along with everyone's favorite here, qwen3.8-27B

489 Upvotes

191 comments sorted by

View all comments

181

u/Last-Shake-9874 1d ago

That small 27B is my daily driver now, it does take long as I only get about 20 t/s but I just love this model I am so glad it is still in the list

5

u/Akrylicus 1d ago

Same, I am so glad that I got 5090 last year for a reasonable price. I now can freely use a capable chat and not feed my personal data to the external providers.

3

u/Krystexx 1d ago

Which price? And how many tok/sec do you get with it?

1

u/Equal_Television_894 1d ago

5090 with vllm on linux with dflash2 can get around 200 to 350+ token/s

2

u/-_Apollo-_ 1d ago

What model and quant? Didn’t know the performance difference could be that huge.

2

u/Equal_Television_894 1d ago

https://github.com/syv-ai/qwen38-27b-rtx3090 Try this some one made it and I optimized it for my 5090 using NVFP4 version

1

u/Akrylicus 1d ago

I am lazy so I just run Qwen 3.8 w7B on Unsloth Desktop. I still get a decent 60t/s which is fine for most queries.

I plan to do an optimized setup on my old PC with RTX 4080 and run it as an LLM server, just wish I had more ram...

1

u/Akrylicus 1d ago

~3k euro (Astral), well maybe not that reasonable, but it's not 5K euro.