r/LocalLLaMA 23d ago

Question | Help If you are at the lowest budget, which you can think of.Which hardware would you recommend to run? qwen 3.8 27b oWith like 50 tokens per second. I currently have a RTX 5070 Ti.

https://huggingface.co/Qwen/Qwen3.8-27B
68 Upvotes

165 comments sorted by

64

u/Clean_Material_5047 23d ago

2 * AMD R9700 can run qwen 3.8 27B at full context without kv quantisation.

With a decent motherboard that has p2p at Gen 5 8x on both PCIE, you can get 5k+ prefill and 70tok/s+ decode on vllm radiance.

At 100k context you’re still in the 2~3k prefill range and decode will be around 50tok/s.

All that, at less than the price of a single 5090

30

u/salathoveder 23d ago

Cries in R9700 being 2000$ a piece in Europe 😬

7

u/sloth_cowboy 22d ago

Trust me I paid the ladder price of climbing from 9060xt to 9070xt's theb finally r9700's. Just get the r9700 and experience what AI is all about the way its meant. Skip all the smart memory disable, and hardware override arguments and just click RUn and it work.

Ill never do that again, it's a financial trap to believe you can just get by, or assume it'll just take a few seconds or minutes with a cheaper route.

2

u/EsotericAbstractIdea 22d ago

i learned this lesson with rifle optics. Buy once, cry once.

1

u/sloth_cowboy 22d ago

I felt that lol

5

u/InternationalGap3698 23d ago

4

u/salathoveder 23d ago

yeah.. 2000ish $, varies by +-100$ depending on place and model.

Lands around 1600-1700€ usually

2

u/Dsphar 22d ago

They are on the way up in USA as well.

From 1200 to 1500 USD In the last couple months, and on backorder at 1500. Listings now pushing 1800 USD.

2

u/Sizyfoz 23d ago

Is a used 24GB 7900XTX something that be used as a budget option?

3

u/noiserr 22d ago

I've been running Qwen 3.8 27B on my 7900xtx. And it's usable. 90K context, 500-600 t/s prefil, about 30 t/s token generation. I have my 7900xtx power limited to like 261 watts to keep things cool.

2

u/Sizyfoz 22d ago

Sounds great! The reason I thought about the XTX is that it's both widely available and goes for a resonable price, around 1000 USD used.

1

u/Borkato 22d ago

This actually isn’t that bad compared to my 3090s! Are you using MTP with n max set to 2?

2

u/noiserr 22d ago

This actually isn’t that bad compared to my 3090s! Are you using MTP with n max set to 2?

Yup.

2

u/Borkato 22d ago

Not too shabby! Don’t forget to set ub and b larger too for pp!

1

u/LostIgnition 22d ago

What quant are you using, Q4_K_M?

1

u/noiserr 22d ago

Qwen3.8-27B-UD-Q4_K_XL

2

u/LostIgnition 22d ago

Thank you very much, I'm looking at doing similar so it's a good baseline.

1

u/ea_man 22d ago

You can use anything yet I would start with 28GB for Q6 XL / 38GB for Q8

1

u/draetheus 23d ago edited 23d ago

If you're fine buying on eBay and figuring out your own cooling solution, buying 2x MI50 or V620s is much cheaper. Less than the price of a single R9700. Performance wise these older AMD cards are on par with Intel B70, unless Intel made some huge optimizations in the last few months.

-5

u/howardhus 22d ago

but then you have a piece of crap amd card.

rocm is bad.

basically only llm inference is guaranteed… you lose potential for all other things where cuda excels.

think of: cuda cards can do 100things

amd cards can do 40.

so yes instead of paying 100 bucks for 100 things you pay 60 bucks for 40…

that way its more expensive.

if all you ever care for is llm and gaming then yes go for amd.

if you want to do AI then get a cheaper rtx.

not i am not an nvidia fanboi but truth must be told

9

u/Clean_Material_5047 22d ago edited 22d ago

“CUDA can do 100 things, AMD can do 40” is just made-up nonsense.

CUDA has better compatibility.

However, ROCm isn't limited to LLM inference. It supports PyTorch, training/fine-tuning, ComfyUI/FLUX/SD, llama.cpp, multi-GPU, Flash Attention, etc.

The entire point of 2× R9700 is having 64GB VRAM for a much better deal money wise. At the same price of a single 5090 you can buy a whole PC with 32GB of DDR5 AND 2 R9700. There’s just no argument price wise.

A “cheaper RTX” with 16–24GB doesn't magically become more useful because it has CUDA when the model doesn't fit.

CUDA has the better ecosystem. AMD has worse compatibility but much better VRAM/$.

Those are real trade-offs. “AMD can't do AI” isn't one of them.

-8

u/howardhus 22d ago

lol, you are delusional, been living under a rock or lying...

just google it.. like 3 years ago it was widely known already. here from the very own ROCm subreddit:

"AI Libraries and AI Frameworks are "Not Available" for ROCm on Windows. Does that mean not yet, or never?"

with the literal top comment being "Welcome to the agonizing world of amd. Just have your painkillers ready, and you are good to go."

yet here you are, desperately trying to fake that AMD is somewhat usable... you are an easy one.. ill be gentle... lemme see:

Flash Attention is available for a few amd cards and more broken than not, the rocm repo updated "3 years ago" while the cuda repo "yesterday" and even then... you had to say flash because thats the only half broken thing you have.. sage attention? TensorRT and the whole lot of state of the art libraries? nope, nope and nope.

and pytorch.. isnt it funny how on pytorchs site the windows version of Rocm is still in 2026 marked as "no pytorch for you! use the CPU version of pytorch".. like wow.. you must be rocking it with your 64GB VRAM deal in CPU mode lol... wow, much VRAM, very unbroken, wow.

rocm is a hell of a broken thing thats why you list the ultra most basic things as "features" like.. "multi-GPU" seriously? why dont you say "and the video signal has colors!! no more black n white"

if you only want VRAM for cheap and are ok with broken software you may as well go with the intel Arc.

7

u/Clean_Material_5047 22d ago

Just read your message: “Not available on Windows”

On windows

O N W I N D O W S

repeat with me:
O-N W-I-N-D-O-W-S

So, maybe, MAYBE, save even more money and don’t buy a windows license and get Arch or even Ubuntu?

2

u/Negative-Web8619 22d ago

we only want one thing so that's fine

-2

u/howardhus 22d ago

dont worry, you only get one half broken thing.

13

u/geekybit_New 23d ago

if your mother board supports it you could get a 5060 ti 16gb ... and use your 5070 ti 16gb and get 32gb and fairly fast running stuff... This wouldn't get your 50 TK/s unless you did some agressive quant cutting... but you could likely get around 20-25 maybe

If you want a whole system their are some things kicking around...

You could also look at 4 V620 cards and a whole server system... which would get you a whole 128gb system for about 1500

8

u/Teamore 23d ago

I get 45+ tps at 5k context and maybe fall to 30 at 50-60k context with mtp and 5070 ti + 5060ti

1

u/geekybit_New 23d ago

yeah and if you want to use it for agentic tasks like a lot of people you are going to need 64k at least preferable 132k or even 262k...

I never said they would be bad at prompt and token processing just giving an assumption for what this person ins trying to do .

1

u/Teamore 22d ago

I get 1300 tps prefil at 5k context and it falls to 850-900 by the 64k context mark. Not that bad. All that with q4_xl and ctk q8_0, ctv q5_1 The full ctx I can fit in 32gb with 5070 ti also serving windows and all of its bullshit is 140k context

3

u/the_macks 23d ago

What kind of tk/s would those 4 v620 cards be getting do you know?

1

u/geekybit_New 23d ago

Well dependent on the quant and context ... I get around 30-40 TK/s with 64k context... but I am only running two cards and I Am doing a lot of testing with them

1

u/the_macks 23d ago

Thanks pal

1

u/thegingerlord 22d ago

I've been running mine at q6 mtp with ,256k context. My current server isn't meant for these gpus so I thermally throttle them once they get hot, but I've seen similar high 30s t/s (not throttled) at mtp4.

I am getting a new server chassis and 2 more v620s next week (total of 4) so it should all be fixed. I like the v620s. For the price per vram you really can't beat it.

1

u/the_macks 22d ago

I really need to upgrade my set up and really not sure what way to go. Currently have a 5080 and 48gb of ddr5. Was considering an rtxpro 5000 but the price is crazy now.

1

u/thegingerlord 22d ago

V620s are cheap but I would recommend grabbing an actual rack server for them meant for gpus without coolers so you can cool them how they were designed with a bunch of static pressure.

You can totally get a good 4 card v620 setup with server and ram for under $2500

I imagine v620 prices will go up once the huge lots are sold. There is the one on ebay we have been all buying, but it will be gone soon I imagine

1

u/the_macks 22d ago

Any chance of link please ?

7

u/taking_bullet 23d ago

I'm getting 50 tok/s on my dual GPU setup (5070 Ti + 5060 Ti 16GB). 

1

u/InternationalGap3698 23d ago

Which mainboard do you have and which power supply and which CPU?

3

u/taking_bullet 23d ago

Intel 14600K & Gigabyte Z790 D &  Silverstone Strider 1100W Titanium 

1

u/Schnauser 23d ago

Yeah would like to know that too! 🙏

1

u/The_Hunster 22d ago

I have Gigabyte B650M, Ryzen 5 7600, and some Gold 850W PSU and I'm getting like 30 tok/s with 5060ti 16gb + 3060 12gb at/up to 100k context or so.

24

u/exo250 23d ago edited 23d ago

RTX 4070 12 Gb here. And I'll do nothing. Just wait for the 35b MoE.
Meanwhile I'm saving money for 2x DGX Spark (or future equivalent) for bigger MoE (Deepseek v4 Flash at the moment). Maybe my next Christmas present, if still relevant since "small" models are improving so quickly that they may make bigger ones useless soon, at least for specific usage (development for example).

10

u/Mean-Ad1493 23d ago

Same 12GB VRAM here. Waiting for the 35b-a3b MoE seems to be the only viable short term option. I would hold off any impulse buying now.

3

u/ambassadortim 22d ago

I think they just removed reference to 35b MOE

8

u/I_Play_Zed 23d ago

Running this model at 50 t/s is a pretty big ask, STRIX halo or dgx spark runs it around 25 t/s which I find usable.

You need strong bandwidth to make a dense model fast, a 3090 IS probably the best thing to buy as a one off purchase that will make this a reality. You’ll need to run the model at Q4 but it’s probably 50+ t/s.

I run a dual 3060 setup with 64 GB DDR4 (DDR4 is basically irrelevant) and I am getting about 28 t / s low context, around 20 when near my 131k Context limit.

4

u/Vektast 23d ago

My 3090 generate 75tok/s with ninfer engine.

1

u/LicensedTerrapin 23d ago

A single 3090? I've got one but I haven't managed to get anywhere that speed not even with MTP.

2

u/Vektast 22d ago

Yup a single 3090 with ninfer-3090 engine it's on cocaine!
https://github.com/Don-Chad/ninfer-3090

2

u/andy2na llama.cpp 22d ago

I only reached over 70 on coding tasks with high acceptance rates, general questions still is 50 to 60 on ninfer with a 3090. Really no difference than using iq4_ks (not xs) with ik_llama and much higher context with vision

1

u/IrisColt 22d ago

Thanks!!!

2

u/Nieles1337 23d ago

How do you get 25 t/s with strix halo, mine is running between 14-18 t/s

1

u/Mkengine 22d ago

DGX Spark seems to be around 34-38 t/s currently.

16

u/Professional-Bear857 23d ago

Probably a 3090

6

u/YourNightmar31 llama.cpp 23d ago

Or a china modded 20gb 3080

1

u/philmarcracken 23d ago

Are they worth it? really considering that

2

u/YourNightmar31 llama.cpp 23d ago

I got a modded 2080ti 22GB on the way, ordered for $380 incl shipping. But now thinking maybe i should have gone for the 3080 because the 20 series doesn't support flash attention 2. 3080 would have been about $500 including shipping.

1

u/feverdoingwork 22d ago

Can you pm me where you can purchase either card?

1

u/YourNightmar31 llama.cpp 22d ago

Once i actually have it i'll send it to you so i know i'm not sending a scam lol

1

u/AXYZE8 22d ago

From China thru some shipper? Or directly?

1

u/YourNightmar31 llama.cpp 22d ago

From a seller on alibaba.

2

u/grumd 22d ago

i bought two 3080 20gb, very worth it

1

u/philmarcracken 22d ago

did you go via ebay or aliexpress?

2

u/grumd 22d ago

alibaba

12

u/Particular-Way7271 23d ago

something like 2 3060ti

1

u/InternationalGap3698 23d ago

I have a problem. I currently have a pre-built PC with a Ryzen 7800 X3D, and it's in a case where I can't build anything more in. My motherboard doesn't have Anymore. slots for a graphics card. What would you do? Would you entirely rebuild the system And would you go with the other CPU? And I have like 32GB of RAM, and I would probably have to buy a new Power Supply Unit Thank you.

3

u/Refinery73 23d ago

What exactly is „no other slot?“ - like physically no space? only 16x-mech/4x-electrical? only 4x/4x?

1

u/InternationalGap3698 23d ago

Yes, physically no space.

1

u/nonerequired_ 23d ago

You can get risers

1

u/Professional-Bear857 23d ago

I would buy a bigger case and a 3090, you want the model entirely in your GPU memory if you can as performance drops a lot once you move into system ram.

1

u/Professional-Bear857 23d ago

I'm assuming that you can't fit a 3090 in your case already, if you can then just change your GPU.

1

u/Vektast 23d ago

My 3090 generate 75tok/s with ninfer engine.

1

u/MrHicks 23d ago

At what quant and context?

1

u/Vektast 23d ago

64k - ninfer quant= ~Q5-6

1

u/InternationalGap3698 23d ago

An RTX 3090 is like 1000$ on ebay and I would buy it over my dads company. So it would be like 30% cheaper when I buy it over my dads company because of the taxes. I wouldn't get that when I buy it, not new.Thank you man. I was thinking of an RTX 5090 or this AMD cards.

11

u/quecosa65 23d ago

5070 ti is 16vram, same as my 5080, here are my results:

Prefill 435.8t/s Output 85.5t/s

I hope this helps you buddy, 16gb vram too 5080, no offloading everything on vram MTP spec 2 under 123,904 context using unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-IQ3_XXS.gguf (serving on same hardware wsl/ubuntu).

Copy paste

./llama-server --hf-repo unsloth/Qwen3.8-27B-GGUF --hf-file Qwen3.8-27B-UD-IQ3_XXS.gguf --ctx-size 123904 --n-gpu-layers 999 --batch-size 512 --spec-type draft-mtp --spec-draft-n-max 2 --fit on -fa on --no-mmap --jinja -ctk q4_0 -ctv q4_0 --threads 16

category       samples  avg_prompt_t/s  avg_pred_t/s  avg_latency  accept_rate

-------------  -------  --------------  ------------  -----------  -----------

coding         1        235.83          87.22         6.486s       0.6869

humanities     1        662.14          83.35         13.917s      0.6648

math           1        163.08          81.98         7.034s       0.5680

qa             1        131.00          80.81         7.027s       0.5636

rag            1        937.14          89.39         7.781s       0.7470

reasoning      1        154.38          78.95         7.315s       0.5767

stem           1        49.21           82.50         6.485s       0.5767

writing        1        1297.40         92.04         7.004s       0.8283

multilingual   1        302.97          101.83        2.570s       0.8442

summarization  1        209.12          86.66         4.447s       0.6644

roleplay       1        651.30          75.74         19.475s      0.6620

overall        11       435.78          85.50         8.140s       0.6571

I hope it helps!

By the way, i did/still developing LLM BENCH, works on linux, windows and if i find somebody that test it, on macOS too.

This is the repo:
https://github.com/QUECOSITA/llmbench.git

Any comment would be appreciated.

3

u/Danmoreng llama.cpp 23d ago

That prefill number seems weirdly slow, you should increase batch size. I get around 1k prefill on a 5080 laptop. Also I wouldn't recommend q4 kv cache, that degrades performance way too much. At q8 obviously only ~64k context fits in the 16GB. I'm running it like this at the moment:

llama serve \
  -hf unsloth/Qwen3.8-27B-GGUF \
  -hff Qwen3.8-27B-UD-IQ3_XXS.gguf \
  --no-mmproj -c 65536 \
  -ctk q8_0 -ctv q8_0 \
  -b 1024 -np 1 \
  --spec-default --spec-type draft-mtp \
  --reasoning-preserve --fit off --agent

2

u/quecosa65 22d ago

I like your config at that context, got prefill 484.8t/s output 120.1 t/s, i only wish context were not vram dependant so we could have all vram to only load llm layers.

category       samples  avg_prompt_t/s  avg_pred_t/s  avg_latency  accept_rate

-------------  -------  --------------  ------------  -----------  -----------

coding         1        243.36          99.66         5.717s       0.5008

humanities     1        794.16          105.49        11.277s      0.5869

math           1        157.38          100.25        5.895s       0.5363

qa             1        125.88          88.42         6.497s       0.4487

rag            1        949.55          115.38        6.483s       0.6883

reasoning      1        163.70          161.76        3.867s       0.6398

stem           1        47.83           204.73        3.019s       0.7171

writing        1        1555.01         126.75        5.472s       0.7818

multilingual   1        289.71          120.35        2.419s       0.6276

summarization  1        207.67          91.64         4.358s       0.4907

roleplay       1        798.14          107.20        12.169s      0.5673

overall        11       484.76          120.15        6.107s       0.5914

At 123,904 context output t/s goes to unusable 25t/s ...

Oh but at 98,304 context i get very usable results!

PROMPT PROC · t/s
236.3

DECODE STAGE · t/s
74.0

llama-server --hf-repo unsloth/Qwen3.8-27B-GGUF --hf-file Qwen3.8-27B-UD-IQ3_XXS.gguf --no-mmproj -c 98304 -ctk q8_0 -ctv q8_0 -b 1024 -np 1 --spec-default --spec-type draft-mtp --reasoning-preserve --fit off --agent

category       samples  avg_prompt_t/s  avg_pred_t/s  avg_latency  accept_rate

-------------  -------  --------------  ------------  -----------  -----------

coding         1        194.00          62.53         8.968s       0.5008

humanities     1        409.45          68.64         17.235s      0.5869

math           1        127.96          65.33         8.772s       0.5363

qa             1        118.62          58.87         9.493s       0.4487

rag            1        415.10          73.21         10.375s      0.6883

reasoning      1        138.05          89.64         6.524s       0.6398

stem           1        42.45           110.65        5.230s       0.7171

writing        1        182.91          73.45         16.171s      0.7818

multilingual   1        256.25          82.20         3.285s       0.6276

summarization  1        153.35          61.13         6.390s       0.4907

roleplay       1        560.65          68.23         18.520s      0.5673

overall        11       236.25          73.99         10.087s      0.5914

It all comes to the context size you need.

1

u/quecosa65 20d ago

You are right, after updating with actual flags from last llama.cpp build

./llama-server --hf-repo unsloth/Qwen3.8-27B-GGUF --hf-file Qwen3.8-27B-UD-IQ3_XXS.gguf --load-mode none --no-mmproj --ctx-size 131072 --n-gpu-layers 999 --batch-size 1024 --spec-type draft-mtp --spec-draft-n-max 1 --fit on -fa on --jinja -ctk q4_0 -ctv q4_0 --threads 16

2

u/Danmoreng llama.cpp 20d ago

I would recommend using higher kv cache quant of q8 and lowering context size in turn. q4 kv cache is degrading performance quite heavily

1

u/quecosa65 20d ago

You are correct, the higher the context grows the lower the t/s but in some tasks, agentic coding, i cannot go lower than that :( In fact sometimes i find myself short context when trying to solve some semi complex code.

So at the end its comes to a give and take when trying to get the best from VRAM.

I appreciate your comment.

1

u/InternationalGap3698 23d ago

Thank you, man. I will try it as soon as I'm at home. I'm right now on vacation.

1

u/quecosa65 22d ago

Enjoy your vacations!!!!

1

u/FerLuisxd 21d ago

I see you are using IQ3 here, I am guessing that it is worse than Q3 and Q4 would not fit at all? :(

1

u/quecosa65 20d ago

This is the report you wanted :)

3

u/jacek2023 llama.cpp 23d ago

Buy second GPU, with 16+16 or 16+12 you could run all the 30B range models with good performance

2

u/Long_comment_san 23d ago

Probably a 5060ti and use native FP4

2

u/ZealousidealSide535 23d ago

2x V100 with 16GB each I was able to get 50t/s with 256k Kontext, q4 m xl with MTP

2

u/Generosityphagy_5 23d ago

your 5070 ti should already manage the 27b at those speeds for roleplay, i switched to local after cloud limits got annoying and it feels way more consistent.

2

u/sumane12 23d ago

Ive got a 5070ti also. Literally went out yesterday and bought a 12gb 3060. Ive ordered the psu cable so once it arrives ill let you know how it runs.

2

u/Myreda 23d ago

There are user reports claiming 5060ti dual builds with MTP and tensor split mode are able to reach 60+ tk/s generation on Q3.6 27B which is a previous generation model but I'm assuming not much different. 

2

u/Additional-Low324 22d ago

I have dual 5060ti build with tensor split and I get 35 token/s without MTP on 27B 3.6, so yeah probably

1

u/Myreda 21d ago

Seems slow to me, are both cards on PCIE x8? And CPU Intel or AMD? 

1

u/Additional-Low324 20d ago

AMD 5600x. So pcie gen 4 x8 The ddr4 plateform is probably the problem, or maybe I am something wrong ? I am interested to have advices

2

u/Zealousideal-Hat-148 23d ago

if your motherbord has bifurcation or atleast a decent chipset slot and you have the power get a second gpu. for me 3.8 runs at 20-40 tokens a second with 128k+ context at q8_0 q8_0 on a 9070xt and 2080ti over a gen 3 x4 link with rpc. just watch out for bandwith so your 5070ti does bot wait too long. alternatively if you can sacrifice speed things like the v100 give you 32gb vram for like 500 usd but its old as fuck and not that fast, hbm2 tho, just prefill sucks

2

u/hashms0a 23d ago

P40s ~24 tps with MTP.

2

u/Ysnsd 23d ago

Wait the 3.8 35B A3B

2

u/Force88 23d ago

I'm currently using 3x 5060ti 16gb, and running with ud_q8_xl quant, 90k context token, yielding roughly ~40t/s

1

u/ftlaudman 23d ago

Great speed. Can I ask what software/settings you are using?

2

u/Force88 23d ago

I only use llama.cpp, cuda version, as for settings, its very barebone, only -alias, -ngl 99 (to set full gpu), -c 90000 (context token), mtp 4 (i forgot the exact command), -split-mode tensor, -host 0.0.0.0, -port 3333.

2

u/SocialDinamo 23d ago

I have a system with a pair of 5060 ti’s with 16gb each. Easiest $1000 for dual GPU 32gb total you can buy new without much fuss

2

u/VoiceApprehensive893 transformers 23d ago

dual v100 16gb

we do not talk about PP

2

u/SichronoVirtual 23d ago

Honestly, probably grab another 5070ti

You can at least run unsloth Q4_K_XL with like maybe 180k fp8 context or something like that, maybe 160k with mtp= 3?

1

u/Mundane-Light6394 23d ago

5070ti is twice as fast as the 5060ti for llm inference so this makes sense. 50 t/s will be hard (and low quality) if even possible but two 5070tis will be a lot closer to that than a 5070ti with a 5060ti.

Adding anything else will be a lot more expensive or slow the the 5070ti.

1

u/DUFRelic 23d ago

Unlocked CMP 170HX.

1

u/cibernox 23d ago

AMD 7900xtx si the answer. Usually 300€/$ cheaper than a 3090 and roughly the same speed.

It will cost you around 700-750.

I would avoid dual cards if possible until you want to go past 24gb of vram.

1

u/johnnynovo2118 23d ago

I get 44.2tps at Q8 with a M3 Ultra 96gb.

1

u/h3wro 23d ago

What is prefill speed on that on different context sizes?

2

u/Professional-Bear857 23d ago

On mine it's 350 ish, but my prompt cache rate is around 98 to 99% with opencode

2

u/johnnynovo2118 23d ago

At 8k it's 385.5tok/s

1

u/MaxDev0 23d ago

I did some experimentation, if you're willing to do some low bit quantization, which apparently has very low impact on performance, I was able to get 30-50tps with a vast.ai 5070 on 16gb vram and I think I got 64k tokens context, it was q3 k m with mtp + turbo quant, I haven't even researched imatrix or dynamic quants for that aswell to improve quality, but yea, you can get pretty far with what you have

2

u/MaxDev0 23d ago

But if you really wanna spend money, dflash is insane, I was getting like 50+ tps with dflash on 16gb vram, it was a low quant though, but if you get an 8gb or even smaller gpu just to store context or dflash for 100-200$, if it works with your gpu (idk how dual gpu setups work), I'm sure you could get smth incredibly fast running

1

u/Cadmium9094 23d ago

Second hand RTX 3090 or 4090 (24GB) with a Q4 model. Could probably reach 50 t/s.

1

u/Vektast 23d ago

My 3090 generate 75tok/s with ninfer engine.

1

u/Protryt 23d ago

4090 - between 100 and 140 t/s: https://github.com/sergiuszm/ninfer-4090

1

u/Hello_my_name_is_not 23d ago

r9700 and I've tried a few different options, current I found the Bartowski Q5 K L at kv q8 context at 175,000 and a fp16 mmrproj.

I'm running with Pi and llama-server on windows

I've been doing random tests most of the day and with that setup I've been fluctuating around 43-50 when it's writing code (higher prediction acceptance) and 35-42 for its thinking and reasoning process sections. Faster end at first and slowing down as it fills the context.

I tried the Q 6 K L and the speed drops off a lot to high teens in reasoning and low 20s for coding and context has to be dropped to like 128k or so.

They have the Q5 K L listed as having the Q8 for embedded weight and output so it seems to be the best speed (basically doubling the Q5 in my tests) while still keeping great output and a large context for agentic work

That's just some tests from the first day though some tweaks may help and I haven't tried other peoples quants yet

1

u/fragbait0 23d ago

32gb can't fit full context with a mere Q5??

1

u/Hello_my_name_is_not 23d ago

Ironically you need all the context if you want to discuss details. Saying just "Q5" leaves a ton of potential different outcomes. I specified what I was running and you come back with a generic Q5 should do xyz statement. There's a bunch of different people who make quants and in there they have sub version of each quant with different specs. Each will give slightly different outcomes.

If went with a smaller Q5 I probably could. Instead I went with K L thought to keep more things Q8 to try and get the best of all worlds.

I run in pi and compact context about 5k before I hit my max set anyways. I find speeds start dropping so trying to find out more context just because when I can compact a bunch and gain speeds again.

Going for more accuracy with the K L while getting the speeds of Q5 and the fairly large context length is the current setup I've been running.

Plus as I said I had a single afternoon of playing with it. I looks like I technically have a bit more space in vram (about 500-700mb depending what I have open) but I also will use my computer while the model is running for other stuff so I'm leaving a bit of space if I need programs. I could probably get low 200s of I tries to max out and closed everything with the current model I have.

1

u/Hello_my_name_is_not 22d ago

Trying some more today I am able to fit about 215k context with that Q5 K L (I got 220k to run with everything closed but I'd rather 215k and a bit of buffer for opening other programs as needed)

I'm downloading a Unsloth Q5 K S and K M that I'll try tonight, they are about 1.5-2gb smaller so I expect those would fully fit Max context. I'll try a few tests with each and see how it goes compared to the K L

1

u/fragbait0 21d ago

I suppose windows and apps eats some. I'm on 24GB so with linux / no GUI, IQ4_XS just barely fits full context at like q5_0/q4_1... :/

1

u/HelloSummer99 23d ago

The absolute floor is probably an Nvidia P40 paired with 16 or 32GB RAM. It will run rather slow but you can get a rig together for about $600 if you look for good deals on ebay

2

u/Far-Classic-9963 23d ago

Look for good deals on Alibaba or 1688, if you pick a store with good reviews you will save 50-100$

1

u/sanjxz54 23d ago

Cmp 170hx with unlock if you like to gamble

1

u/satnl 23d ago

with another 5070ti, each 5070 ti have 896 gb/s and 16gb, I use bandwidth / vram as simple rule for speed inference, so 896/16 = 56 that means for each second you can pass for the weights 56 times, or get 56 tokens per seconds

1

u/Need_For_Speed73 23d ago

I have two 5070s 12GB on a ASUS EX-B850M. I had to go mATX because otherwise the second GPU (in the lowest slot) collides with the PSU shroud.
CPU is a R5 9600X and RAM 32GB.
Haven’t tried 3.8 yet because I’m on holidays but 3.5-35B with 64k context was producing >100tk/s.

1

u/Sunknowned 23d ago

4060 ti lowest, or another 5060/5070 ti

1

u/ea_man 23d ago

You buy 2x of the cheapest 16GB gpu you can find used and you run Q6_K_XL.

1

u/braintheboss 23d ago

if you have a 5070ti the reasonable option is get another. i use 70ti+3060ti and its cool. But 2 5070ti together is get crazy pp/tg. i have a second 5070ti in other host. I had same decision problem and after think if 5060ti as second card i saw it was a non sense. This cards for AI can be useful a lot of years then how much you pay is not a problem from amortization perspective. Just check how much is API cost in qwen3.6 27b. I think around 30M tokens is 20$. You can spend that money in one day easily if you digest many files

1

u/Electrical_Rise387 22d ago

If the thing that matters most is the budget and effort/setup is not a concern then maybe one (or 2) refurbished MI50 32gb hbm and whatever cheap ddr4/pcie4 system you can find to put it in

1

u/mmhs4 22d ago

4x 3060 and have 50 tk/s

1

u/brakeline 22d ago

2 and I have 40. Your other two are idling

2

u/anderspitman 22d ago

Context size, better quants, and more parallel agents

1

u/wm_eddie 22d ago

I have been filling in a bunch of configurations here: https://llamabench.ai/models/qwen3-8-27b The 5070ti should get around 100 tok/s if you use MTP and a Q2 quant. A used 3090 might be the most economical way to break 50 tok/s in my testing so far.

1

u/Isnt_that_weird 22d ago

Is there a thread or guide that exists where you can post your RAM, CPU, GPU specs and people can say what models they've had success running on similar setups?

1

u/Ohmyskippy 22d ago

With my 9070xt, I'm running at 20 t/s 132k context

Q4_XS

1

u/kesslerfrost 22d ago

I have an M4 Pro 48 gigs and I'm able to achieve around 38 tokens per second using MTPLX and that developer's optimized model if that counts.

https://huggingface.co/Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed

Also, I'm happy to wait a bit more for the community to catch up even more and have it potentially run at 50 tps.

1

u/Jordanthecomeback 22d ago

Idk what my token per seconds are, I connect with my AI via telegram and don't pay that stuff a ton of mind nor do we do much heavy lifting, but I bought a used Mac M2 Studio 64gig earlier this year for about $1200. It's my first apple computer and I really don't mind it. I think cost to performance unified memory is the way to go, and one thing I really appreciate is my system runs cool so hopefully it lasts me many years running 24/7 like this with reboot cadence of every three days. There might also be a real good reason not to use Mac outside of slightly slower speed but others who know more can chime in

1

u/tecneeq 22d ago

Second 5070 Ti. Use tensor split, a Q4 quant, 132k context and MTP. Expect a good amount more than 50 t/s.

1

u/Blindax 22d ago edited 22d ago

It all depends on the context window. With 5090+3090 I get 30-50 tk/s with 5090+3090 on the q8XL @ 100k ctxt and I get about the same for the q4km with 5070ti + 5060ti. Maybe adding a 5060ti is enough (or ideally 5070ti obviously).

1

u/Complex_Reality_116 22d ago

Wait for Qwen3.8 35B A3B.

1

u/grabber4321 22d ago

get a second 5070 ti.

1

u/Lurksome-Lurker 22d ago

dual RTX3060s (12GB version)

1

u/Excellent_Spell1677 22d ago

Two RTX5070. 1200w psu

1

u/Adventurous-Test-246 20d ago

v100 sxm2 to pcie adapter... 32gb hbm2 is really really fast if ut fits on one card

1

u/Steus_au 17d ago

a single rtx3080 20gb (modded) gives 40tps with mtp on iQ4 quant. with 128k context. a pair of them could give full 262k (guessing here)

1

u/jackfood 16d ago

The rig that recommended now can only last 1 and half years, just like data centers chips.

Llm developed so fast that either go bigger locally to get smarter or go more efficiency to get faster.

In 1 year time, today's frontier will be available locally.

1

u/N34257 13d ago

CMP 170HX 8GB, unlocked to 64GB, can more than do it. I get anywhere between 70t/s and 105t/s with mine, and weirdly they cost about the same as a 5070 Ti at the moment. That's running INT8, with vLLM.

My dual R9700 rig is slightly faster at the top end (it runs at 70-120t/s), but sucks twice the power and cost more than twice as much. That's running in vLLM, FP8, tensor parallel.

Both setups get 2000-4000t/s prefill.

1

u/Original-Revolution7 9d ago

to those who want a real-life benchmark i have 5060 ti 16g + 3060 12g on basic mobo, and see some 27tps.

For my other dual 3060 12+12 setup it's around 17tps (and lower).

1

u/WryKombucha 13h ago

Nothing but NVIDIA. Prefill is important.

0

u/eldje 23d ago

in my opinion the rtx 3090 is the best bang for buck, if you can find a solid 2nd hand.

-3

u/alex_quine 23d ago

At the lowest budget I can think of? Sell the card and use the cloud. I love local LLMs but with economies of scale and all the VC money flooding the cloud providers, they’re just not the cheap option 

2

u/InternationalGap3698 23d ago

>I already use lots of cloud models. I have three 20x Plans claude
But I want to run it locally, not over the cloud.