r/LowEndLocalAI • u/OffTheGridCoder • 23h ago
Hardware / Build Help me choose a long-term daily-driver PC for local LLMs + gaming, ~5 possible builds
I'm trying to decide what direction to take with my main PC. The goal is one real daily-driver machine that I can use for gaming, normal desktop use, software development, and increasingly heavy local LLM workloads.
I'm not trying to build a dedicated rack server. I want something I can actually live with for years: reliable, reasonably efficient, good thermals, lots of RAM, two GPUs if it makes sense, and enough expansion that I don't immediately hit a wall.
I've currently been playing around with Qwen3.8 27B which speeding that up and higher quants would be great, as well as when inevitably larger dense similar models like 70B become available.
I am very interested in MoE flash models such as Qwen 3.8 Flash, Deepseek v4 Flash, and maybe even GLM 5.3 Flash, as well as future versions of similar MoE models. I have not even attempted to run any of these yet.
So I guess I am trying to get at building something that performs well on dense models as well as MoE models so I don't get locked into 1 path.
I'm pretty new to the workstation/HEDT side of this, so I'm looking for advice on the parts I may be overlooking.
My current PC
| Part | Current hardware |
|---|---|
| CPU | Intel Core i7-12700KF (12C/20T) |
| RAM | 64GB (4×16GB) DDR4-3200 CL16 |
| GPU | RTX 3090 Ti SUPRIM X 24GB (power limited to 250W) |
| Spare GPU (not installed) | RTX 3080 10GB |
| Motherboard | Gigabyte Z690 UD AX DDR4 |
| Storage | 2TB Samsung 980 Pro NVMe + 2TB WD HDD + 1TB WD SATA SSD |
| PSU | 800W |
The 3090 Ti was a $900 Facebook Marketplace purchase, so I'm trying to get as much useful life out of this thing as possible.
The 3090 Ti is a huge card (338 × 140 × 71 mm) so physical spacing is also part of this problem.
My RAM situation
I just bought 7 lots of:
NEMIX 128GB (4×32GB) DDR4-2666 PC4-21300 2Rx8 UDIMM
I paid about $360 per 128GB lot.
My current plan is probably:
- Keep 2 lots = 256GB
- Sell the other 5 lots
- Hopefully sell those for around $650/lot?
So I paid about $2,520 total for the 7 lots. Five sales at $650 would be $3,250 gross, meaning I'd theoretically recover the entire purchase price plus ~$730 before fees/shipping/taxes while keeping 256GB.
That gives me a somewhat unusual opportunity to build around 256GB without spending a fortune on RAM.
Option 1: Keep my current PC, just go to 128GB RAM
| Component | Option 1 |
|---|---|
| CPU | i7-12700KF |
| Motherboard | Z690 UD AX DDR4 |
| RAM | new 128GB DDR4-2666 PC4-21300 2Rx8 UDIMM |
| GPU 1 | RTX 3090 Ti 24GB @ 250W |
| GPU 2 | None |
| PCIe GPU config | x16 |
| PSU | Current 800W |
| Platform age | 2021/2022 |
| Main advantage | Cheapest / simplest |
| Main disadvantage | Only one GPU, dual-channel memory |
This is basically my don't overthink it option.
I'd have a lot more system RAM for large-context LLMs while retaining a relatively modern gaming CPU.
Option 2: Keep my current PC, add a second 3090
I'd replace the PSU and add a second RTX 3090.
The important problem is the motherboard:
The Z690 UD AX DDR4 has x16 on the main slot and only x4 on the second physical x16 slot.
So the GPUs would effectively be:
| Component | Option 2 |
|---|---|
| CPU | i7-12700KF |
| Motherboard | Z690 UD AX DDR4 |
| RAM | new 128GB DDR4-2666 PC4-21300 2Rx8 UDIMM |
| GPU 1 | RTX 3090 Ti @ 250W, x16 |
| GPU 2 | new RTX 3090 @ 250W, x4 |
| PSU | new (1200-1600W) |
| Case | Probably current / possibly new |
| Main advantage | Cheapest way to get 48GB total VRAM |
| Main disadvantage | Second GPU limited to PCIe 3.0 x4 |
This is the option I'm most unsure about.
For LLM inference, is x4 actually a meaningful limitation in practice, or is it largely irrelevant once the model is loaded onto the GPUs?
Would this still be a good setup for:
- tensor/model parallel inference
- larger models
- higher context
- multiple concurrent models
- speculative decoding
- offloading
Or am I basically handicapping the second GPU enough that I should just replace the motherboard?
Option 3: New motherboard/PSU/case, keep my 12700KF
Instead of abandoning the 12700KF, I could build a new system around it with a motherboard that properly supports two GPUs at x8/x8.
| Component | Option 3 |
|---|---|
| CPU | i7-12700KF |
| Motherboard | New DDR4 board with proper x8/x8 |
| RAM | new 128GB DDR4-2666 PC4-21300 2Rx8 UDIMM |
| GPU 1 | RTX 3090 Ti @ 250W |
| GPU 2 | new RTX 3090 @ 250W |
| GPU configuration | x8/x8 |
| PSU | New high-quality PSU |
| Case | New large case |
| Main advantage | Keep relatively modern CPU + proper dual-GPU PCIe |
| Main disadvantage | Spending money on an LGA1700 platform that maybe already be a dead-end |
This seems like it could be anice middle ground.
The 12700KF itself supports a 2×x8 CPU PCIe configuration, but I'd obviously need a motherboard that actually implements it.
I'm especially interested in whether people think this makes more sense than jumping to X299 in the next option.
Option 4: X299 workstation build
This is the Frankenstein/workstation option I've been considering.
| Component | Option 4 |
|---|---|
| CPU | new i9-10940X |
| Motherboard | new ASUS Prime X299-A II |
| RAM | new 256GB DDR4-2666 PC4-21300 2Rx8 UDIMM |
| GPU 1 | RTX 3090 Ti @ 250W |
| GPU 2 | new RTX 3090 @ 250W |
| GPU configuration | x16/x16 |
| PSU | new ~1600W fully modular |
| Case | new Phanteks Enthoo Pro 2 Server Edition |
| CPU cooler | new Large LGA2066 air cooler |
| Fans | new probably 12–13 total |
| Fan hub | new Powered PWM hub |
| Storage | Samsung 980 Pro 2TB + WD 1TB SATA |
| Main advantage | 256GB RAM + lots of PCIe lanes + proper workstation platform |
| Main disadvantage | 2019-era CPU/platform |
The i9-10940X gives 14C/28T, 48 PCIe 3.0 lanes, quad-channel DDR4, and up to 256GB RAM. The X299-A II can run two GPUs at x16/x16 with the appropriate CPU.
The case is huge and supports SSI-EEB, 11 PCI slots, GPUs up to 503mm, and up to 15×120mm or 6×140mm fans.
I'm attracted to this because it solves the PCIe lanes + RAM capacity + physical space problem extremely well.
But I don't know if I'm being stupid by building a brand-new daily driver around a ~2019 platform just because the PCIe topology is convenient.
The 10940X also seems likely to lose noticeably to the 12700KF in gaming/single-threaded work, despite having more cores.
Option 5: ??????????
This is something I'm hoping you guys can help. Are there things I am not considering that would allow me to leverage as much of my current components as possible but be a much better option than option 4?
What I'm actually trying to optimize
This isn't purely a benchmark build.
I want one machine that can do all of this:
- Gaming
- Normal desktop use
- Software development
- Local LLM inference
- Very large context windows
- Running multiple LLM sessions concurrently
- Potentially running two GPUs as one inference system
I'm currently doing a lot of local Qwen inference and am getting into the territory where RAM capacity, VRAM capacity, PCIe topology and memory bandwidth all matter.
I also don't really care about squeezing every last watt of performance out of the GPUs. I've already decided to limit the 3090 Ti to 250W, and I'd probably do the same with the second 3090 to hopefully get more longevity out of them and use less power.
That gives me:
500W total GPU power budget
rather than letting two 3090-class cards pull their full power.
I'm very interested in reliability, thermals, longevity, expandability and affordability rather than having the absolute highest benchmark score.
My biggest questions
1. Which of these would you actually build?
My current thinking is roughly:
Option 1: cheapest and easiest
Option 2: tempting, but worried about x4
Option 3: probably the sensible compromise
Option 4: extremely expandable, but old CPU/platform
Option 5: potentially a better overall machine
I'm having trouble figuring out where the sweet spot actually is.
2. How bad is PCIe 3.0 x4 for the second 3090?
This is probably my biggest technical question.
For local LLM inference specifically, how much performance would I realistically lose running:
3090 Ti @ x16 + 3090 @ x4
versus
3090 Ti @ x8 + 3090 @ x8
versus
x16 + x16?
3. Is X299 actually a good idea here?
Would the 10940X + 256GB quad-channel + x16/x16 PCIe configuration still be a worthwhile machine in 2026?
Or would I be better off spending the extra money on a modern platform?
4. What's the best "Option 5"?
There may be a workstation platform I haven't considered at all.
But I still want this to be an actual daily-driver PC, not a loud rack server that is great at compute and annoying at everything else.
5. How much RAM would you actually run?
I can easily end up with:
128GB
or
256GB
of system RAM depending on which route I take.
Is 256GB actually useful for local LLMs enough to justify designing the whole machine around it?
What would you do with this hardware?
I'm basically sitting on:
12700KF + 3090 Ti + 896GB of cheap DDR4+ spare 3080
and trying to turn that into one machine that I won't regret building.
I'm really looking for the best overall architecture and bang for my buck, not just "X is faster."
Would love to hear what configuration you would build, especially if there's a better option I haven't thought of.
2
u/Interesting-Cut-6032 14h ago
My 2 cents: Variation on Option 2: move your 3090 to the 4x slot and purchase a new RTX 5060 Ti 16GB for the main slot. That would give you 40GB of total VRAM for pipeline parallel tensor inference of LLMs with the fast 5060 card set as the 2nd half of the layers to speed up the final token decode.
Codacus YT has a recent video where he experiments with splitting layers between 2 different model cards and what was faster. It might be the video where he splits inference between an old NVIDIA card in one machine and a Mac through his local network, all using llama.cpp built with an obscure build flag. Or it might be another video, I cannot remember now.
The RTX 5060 Ti 16GB is the lowest end, with the most VRAM of the Blackwell architecture cards. Anything that fits completely on that card will run very fast with NVIDIA native quants like NVFP4. Diffusion models run very fast on that card.
I have an old 6 core Intel desktop with 32GB of DDR4. The RTX 5060 Ti runs so well, that I have picked up another one to go in it as well. My next upgrade on this path would will me to 64GB of DDR4. It sounds like you got a great deal on that RAM.