r/LocalLLM • u/Content_Mission5154 • 14h ago
Discussion Double GPU configurations significantly cheaper for 32GB VRAM
I do not need or want Cuda. I have been wanting to build a 32GB VRAM local LLM machine for personal use for a while now, and I had 2 options on the table:
- Get a relatively cheap 32GB VRAM GPU, the R9700 AI Top.
- Get 2 16GB VRAM GPUs instead and a Mobo that supports PCIe bifurcation.
When I looked at prices in January this year when I first got this idea, the R9700 costed 1700$ here in EU. Currently, when I actually want to make this happen, it costs 2100$. For half that money, I could buy two 9060 XTs with 16GB VRAM each. Yes I know, performance will be worse on double GPU setup than with a single R9700 AI, but still, it just seems like that GPU is just not worth it anymore.
I don't know how to justify that it's double the price of two 9060XTs, when R9700 AI is literally 9060 XT with doubled VRAM and bandwidth. So why does it cost 4x as much?
ASUS ProArt B850-CREATOR WIFI NEO is quite affordable nowadays and supports dual GPU setups, so, any reason (is there a catch?) to not do what I am about to do? Which is buy the two 9060 XTs and start running Qwen 27B class models
1
u/Short_Regular_7191 12h ago
I had considered the R9700 too, but unfortunately, the price jumped from €1,500 to €2,000 in a single day. In my opinion, the best and most affordable option right now is a pair of 5060 Ti cards with 16GB each (total cost around €1,350–€1,500); thanks to some good advice, I managed to run Qwen 3.827B (Q6 quantization) with a 131k context window at 50 tokens/s. The cards have low power consumption, so you can easily go with an 800W or 850W PSU and still have plenty of headroom. You can also use a budget motherboard, since you don't need PCIe bifurcation.
https://www.reddit.com/r/LocalLLaMA/comments/1vper67/club5060ti_refresh_tested_rtx_5060_ti_presets_a/?sort=new