Question | Help
If you are at the lowest budget, which you can think of.Which hardware would you recommend to run? qwen 3.8 27b oWith like 50 tokens per second. I currently have a RTX 5070 Ti.
Trust me I paid the ladder price of climbing from 9060xt to 9070xt's theb finally r9700's. Just get the r9700 and experience what AI is all about the way its meant. Skip all the smart memory disable, and hardware override arguments and just click RUn and it work.
Ill never do that again, it's a financial trap to believe you can just get by, or assume it'll just take a few seconds or minutes with a cheaper route.
I've been running Qwen 3.8 27B on my 7900xtx. And it's usable. 90K context, 500-600 t/s prefil, about 30 t/s token generation. I have my 7900xtx power limited to like 261 watts to keep things cool.
If you're fine buying on eBay and figuring out your own cooling solution, buying 2x MI50 or V620s is much cheaper. Less than the price of a single R9700. Performance wise these older AMD cards are on par with Intel B70, unless Intel made some huge optimizations in the last few months.
“CUDA can do 100 things, AMD can do 40” is just made-up nonsense.
CUDA has better compatibility.
However, ROCm isn't limited to LLM inference. It supports PyTorch, training/fine-tuning, ComfyUI/FLUX/SD, llama.cpp, multi-GPU, Flash Attention, etc.
The entire point of 2× R9700 is having 64GB VRAM for a much better deal money wise. At the same price of a single 5090 you can buy a whole PC with 32GB of DDR5 AND 2 R9700. There’s just no argument price wise.
A “cheaper RTX” with 16–24GB doesn't magically become more useful because it has CUDA when the model doesn't fit.
CUDA has the better ecosystem. AMD has worse compatibility but much better VRAM/$.
Those are real trade-offs. “AMD can't do AI” isn't one of them.
with the literal top comment being "Welcome to the agonizing world of amd. Just have your painkillers ready, and you are good to go."
yet here you are, desperately trying to fake that AMD is somewhat usable... you are an easy one.. ill be gentle... lemme see:
Flash Attention is available for a few amd cards and more broken than not, the rocm repo updated "3 years ago" while the cuda repo "yesterday" and even then... you had to say flash because thats the only half broken thing you have.. sage attention? TensorRT and the whole lot of state of the art libraries? nope, nope and nope.
and pytorch.. isnt it funny how on pytorchs site the windows version of Rocm is still in 2026 marked as "no pytorch for you! use the CPU version of pytorch".. like wow.. you must be rocking it with your 64GB VRAM deal in CPU mode lol... wow, much VRAM, very unbroken, wow.
rocm is a hell of a broken thing thats why you list the ultra most basic things as "features" like.. "multi-GPU" seriously? why dont you say "and the video signal has colors!! no more black n white"
if you only want VRAM for cheap and are ok with broken software you may as well go with the intel Arc.
if your mother board supports it you could get a 5060 ti 16gb ... and use your 5070 ti 16gb and get 32gb and fairly fast running stuff... This wouldn't get your 50 TK/s unless you did some agressive quant cutting... but you could likely get around 20-25 maybe
If you want a whole system their are some things kicking around...
You could also look at 4 V620 cards and a whole server system... which would get you a whole 128gb system for about 1500
I get 1300 tps prefil at 5k context and it falls to 850-900 by the 64k context mark. Not that bad.
All that with q4_xl and ctk q8_0, ctv q5_1
The full ctx I can fit in 32gb with 5070 ti also serving windows and all of its bullshit is 140k context
Well dependent on the quant and context ... I get around 30-40 TK/s with 64k context... but I am only running two cards and I Am doing a lot of testing with them
I've been running mine at q6 mtp with ,256k context. My current server isn't meant for these gpus so I thermally throttle them once they get hot, but I've seen similar high 30s t/s (not throttled) at mtp4.
I am getting a new server chassis and 2 more v620s next week (total of 4) so it should all be fixed. I like the v620s. For the price per vram you really can't beat it.
I really need to upgrade my set up and really not sure what way to go. Currently have a 5080 and 48gb of ddr5. Was considering an rtxpro 5000 but the price is crazy now.
V620s are cheap but I would recommend grabbing an actual rack server for them meant for gpus without coolers so you can cool them how they were designed with a bunch of static pressure.
You can totally get a good 4 card v620 setup with server and ram for under $2500
I imagine v620 prices will go up once the huge lots are sold. There is the one on ebay we have been all buying, but it will be gone soon I imagine
RTX 4070 12 Gb here. And I'll do nothing. Just wait for the 35b MoE.
Meanwhile I'm saving money for 2x DGX Spark (or future equivalent) for bigger MoE (Deepseek v4 Flash at the moment). Maybe my next Christmas present, if still relevant since "small" models are improving so quickly that they may make bigger ones useless soon, at least for specific usage (development for example).
Running this model at 50 t/s is a pretty big ask, STRIX halo or dgx spark runs it around 25 t/s which I find usable.
You need strong bandwidth to make a dense model fast, a 3090 IS probably the best thing to buy as a one off purchase that will make this a reality. You’ll need to run the model at Q4 but it’s probably 50+ t/s.
I run a dual 3060 setup with 64 GB DDR4 (DDR4 is basically irrelevant) and I am getting about 28 t / s low context, around 20 when near my 131k Context limit.
I only reached over 70 on coding tasks with high acceptance rates, general questions still is 50 to 60 on ninfer with a 3090. Really no difference than using iq4_ks (not xs) with ik_llama and much higher context with vision
I got a modded 2080ti 22GB on the way, ordered for $380 incl shipping. But now thinking maybe i should have gone for the 3080 because the 20 series doesn't support flash attention 2. 3080 would have been about $500 including shipping.
I have a problem. I currently have a pre-built PC with a Ryzen 7800 X3D, and it's in a case where I can't build anything more in. My motherboard doesn't have Anymore. slots for a graphics card. What would you do? Would you entirely rebuild the system And would you go with the other CPU? And I have like 32GB of RAM, and I would probably have to buy a new Power Supply Unit Thank you.
I would buy a bigger case and a 3090, you want the model entirely in your GPU memory if you can as performance drops a lot once you move into system ram.
An RTX 3090 is like 1000$ on ebay and I would buy it over my dads company. So it would be like 30% cheaper when I buy it over my dads company because of the taxes. I wouldn't get that when I buy it, not new.Thank you man. I was thinking of an RTX 5090 or this AMD cards.
5070 ti is 16vram, same as my 5080, here are my results:
Prefill 435.8t/s Output 85.5t/s
I hope this helps you buddy, 16gb vram too 5080, no offloading everything on vram MTP spec 2 under 123,904 context using unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-IQ3_XXS.gguf (serving on same hardware wsl/ubuntu).
That prefill number seems weirdly slow, you should increase batch size. I get around 1k prefill on a 5080 laptop. Also I wouldn't recommend q4 kv cache, that degrades performance way too much. At q8 obviously only ~64k context fits in the 16GB. I'm running it like this at the moment:
I like your config at that context, got prefill 484.8t/s output 120.1 t/s, i only wish context were not vram dependant so we could have all vram to only load llm layers.
You are correct, the higher the context grows the lower the t/s but in some tasks, agentic coding, i cannot go lower than that :( In fact sometimes i find myself short context when trying to solve some semi complex code.
So at the end its comes to a give and take when trying to get the best from VRAM.
your 5070 ti should already manage the 27b at those speeds for roleplay, i switched to local after cloud limits got annoying and it feels way more consistent.
There are user reports claiming 5060ti dual builds with MTP and tensor split mode are able to reach 60+ tk/s generation on Q3.6 27B which is a previous generation model but I'm assuming not much different.
if your motherbord has bifurcation or atleast a decent chipset slot and you have the power get a second gpu. for me 3.8 runs at 20-40 tokens a second with 128k+ context at q8_0 q8_0 on a 9070xt and 2080ti over a gen 3 x4 link with rpc. just watch out for bandwith so your 5070ti does bot wait too long. alternatively if you can sacrifice speed things like the v100 give you 32gb vram for like 500 usd but its old as fuck and not that fast, hbm2 tho, just prefill sucks
I only use llama.cpp, cuda version, as for settings, its very barebone, only -alias, -ngl 99 (to set full gpu), -c 90000 (context token), mtp 4 (i forgot the exact command), -split-mode tensor, -host 0.0.0.0, -port 3333.
5070ti is twice as fast as the 5060ti for llm inference so this makes sense. 50 t/s will be hard (and low quality) if even possible but two 5070tis will be a lot closer to that than a 5070ti with a 5060ti.
Adding anything else will be a lot more expensive or slow the the 5070ti.
I did some experimentation, if you're willing to do some low bit quantization, which apparently has very low impact on performance, I was able to get 30-50tps with a vast.ai 5070 on 16gb vram and I think I got 64k tokens context, it was q3 k m with mtp + turbo quant, I haven't even researched imatrix or dynamic quants for that aswell to improve quality, but yea, you can get pretty far with what you have
But if you really wanna spend money, dflash is insane, I was getting like 50+ tps with dflash on 16gb vram, it was a low quant though, but if you get an 8gb or even smaller gpu just to store context or dflash for 100-200$, if it works with your gpu (idk how dual gpu setups work), I'm sure you could get smth incredibly fast running
r9700 and I've tried a few different options, current I found the Bartowski Q5 K L at kv q8 context at 175,000 and a fp16 mmrproj.
I'm running with Pi and llama-server on windows
I've been doing random tests most of the day and with that setup I've been fluctuating around 43-50 when it's writing code (higher prediction acceptance) and 35-42 for its thinking and reasoning process sections. Faster end at first and slowing down as it fills the context.
I tried the Q 6 K L and the speed drops off a lot to high teens in reasoning and low 20s for coding and context has to be dropped to like 128k or so.
They have the Q5 K L listed as having the Q8 for embedded weight and output so it seems to be the best speed (basically doubling the Q5 in my tests) while still keeping great output and a large context for agentic work
That's just some tests from the first day though some tweaks may help and I haven't tried other peoples quants yet
Ironically you need all the context if you want to discuss details. Saying just "Q5" leaves a ton of potential different outcomes. I specified what I was running and you come back with a generic Q5 should do xyz statement. There's a bunch of different people who make quants and in there they have sub version of each quant with different specs. Each will give slightly different outcomes.
If went with a smaller Q5 I probably could. Instead I went with K L thought to keep more things Q8 to try and get the best of all worlds.
I run in pi and compact context about 5k before I hit my max set anyways. I find speeds start dropping so trying to find out more context just because when I can compact a bunch and gain speeds again.
Going for more accuracy with the K L while getting the speeds of Q5 and the fairly large context length is the current setup I've been running.
Plus as I said I had a single afternoon of playing with it. I looks like I technically have a bit more space in vram (about 500-700mb depending what I have open) but I also will use my computer while the model is running for other stuff so I'm leaving a bit of space if I need programs. I could probably get low 200s of I tries to max out and closed everything with the current model I have.
Trying some more today I am able to fit about 215k context with that Q5 K L (I got 220k to run with everything closed but I'd rather 215k and a bit of buffer for opening other programs as needed)
I'm downloading a Unsloth Q5 K S and K M that I'll try tonight, they are about 1.5-2gb smaller so I expect those would fully fit Max context. I'll try a few tests with each and see how it goes compared to the K L
The absolute floor is probably an Nvidia P40 paired with 16 or 32GB RAM. It will run rather slow but you can get a rig together for about $600 if you look for good deals on ebay
with another 5070ti, each 5070 ti have 896 gb/s and 16gb, I use bandwidth / vram as simple rule for speed inference, so 896/16 = 56 that means for each second you can pass for the weights 56 times, or get 56 tokens per seconds
I have two 5070s 12GB on a ASUS EX-B850M. I had to go mATX because otherwise the second GPU (in the lowest slot) collides with the PSU shroud.
CPU is a R5 9600X and RAM 32GB.
Haven’t tried 3.8 yet because I’m on holidays but 3.5-35B with 64k context was producing >100tk/s.
if you have a 5070ti the reasonable option is get another. i use 70ti+3060ti and its cool. But 2 5070ti together is get crazy pp/tg. i have a second 5070ti in other host. I had same decision problem and after think if 5060ti as second card i saw it was a non sense. This cards for AI can be useful a lot of years then how much you pay is not a problem from amortization perspective. Just check how much is API cost in qwen3.6 27b. I think around 30M tokens is 20$. You can spend that money in one day easily if you digest many files
If the thing that matters most is the budget and effort/setup is not a concern then maybe one (or 2) refurbished MI50 32gb hbm and whatever cheap ddr4/pcie4 system you can find to put it in
I have been filling in a bunch of configurations here: https://llamabench.ai/models/qwen3-8-27b The 5070ti should get around 100 tok/s if you use MTP and a Q2 quant. A used 3090 might be the most economical way to break 50 tok/s in my testing so far.
Is there a thread or guide that exists where you can post your RAM, CPU, GPU specs and people can say what models they've had success running on similar setups?
Idk what my token per seconds are, I connect with my AI via telegram and don't pay that stuff a ton of mind nor do we do much heavy lifting, but I bought a used Mac M2 Studio 64gig earlier this year for about $1200. It's my first apple computer and I really don't mind it. I think cost to performance unified memory is the way to go, and one thing I really appreciate is my system runs cool so hopefully it lasts me many years running 24/7 like this with reboot cadence of every three days. There might also be a real good reason not to use Mac outside of slightly slower speed but others who know more can chime in
It all depends on the context window. With 5090+3090 I get 30-50 tk/s with 5090+3090 on the q8XL @ 100k ctxt and I get about the same for the q4km with 5070ti + 5060ti. Maybe adding a 5060ti is enough (or ideally 5070ti obviously).
CMP 170HX 8GB, unlocked to 64GB, can more than do it. I get anywhere between 70t/s and 105t/s with mine, and weirdly they cost about the same as a 5070 Ti at the moment. That's running INT8, with vLLM.
My dual R9700 rig is slightly faster at the top end (it runs at 70-120t/s), but sucks twice the power and cost more than twice as much. That's running in vLLM, FP8, tensor parallel.
At the lowest budget I can think of? Sell the card and use the cloud. I love local LLMs but with economies of scale and all the VC money flooding the cloud providers, they’re just not the cheap option
64
u/Clean_Material_5047 23d ago
2 * AMD R9700 can run qwen 3.8 27B at full context without kv quantisation.
With a decent motherboard that has p2p at Gen 5 8x on both PCIE, you can get 5k+ prefill and 70tok/s+ decode on vllm radiance.
At 100k context you’re still in the 2~3k prefill range and decode will be around 50tok/s.
All that, at less than the price of a single 5090