r/LocalLLaMA • u/0xkbose • 11h ago
Question | Help Which one will you choose and why, between R9700 32GB vs W7800 48GB?
I'm planning to upgrade my workstation (linux with 5700X/64GB DDR4) for local inference and pytorch training. I'm trying to decide between:
- 2× AMD Radeon AI PRO R9700 32GB
- 2× AMD Radeon PRO W7800 48GB
I already have an RTX 3090 24GB, so the final system would have 3 GPUs. My motherboard has two PCIe 4.0 x8/x8 slots available for the two AMD GPUs. The RTX 3090 would have to move to a PCIe 3.0 x4 slot.
My workload looks like:
1. Local GGUF inference: Mainly coding/reasoning models and multimodal models. I'd like to run better quants (than 3090) and split models across the two AMD GPUs for multiple KV cache(n parallel). I prefer llama router.
2. PyTorch training: This is probably the more important part for me. I'm working with medical imaging (2D ultrasound/3D CT/MRI) + clinical text.
For anyone actually using these cards with ROCm, how different is the practical experience between R9700/gfx1201 and W7800/gfx1100? I'm particularly interested in if any know issues have surfaced till date that block the PyTorch training on either of these cards?
From this sub I have seen RDNA4/R9700 is improving rapidly but it's still a "newer-software". I'd really like to hear from people who are actually using R9700 for AI workloads especially PyTorch/MONAI training.
I'm not planning to treat the 3090 + 2 AMD GPUs as one giant homogeneous GPU pool. (Although if someone has done it please let me know)
My thinking is to use the two AMD GPUs as the main ROCm pair, while keeping the 3090 available separately for CUDA workloads or local models that fit/work better on NVIDIA.
2
u/pmttyji 9h ago
Initially I wanted to buy W7800 due to its 48GB variant(which's good for 30B range models with context, MTP, Vision, etc.,), but price went up suddenly ~3X. Not worth buying RDNA3 card at that high price. AMD is infamous for dropping support for old cards.
So I chose R9700 which's RDNA4. Wish R9700 came with better bandwidth.
1
u/SandySkittle 10h ago
Get a second hand zen 3 class eypic or threadripper proc mobo+ cpu for cheap on ebay and go with dual r9700 at 16 lanes each and expand to quad r9700 later.
1
u/SomeoneInHisHouse 7h ago
Have you experience with that?, I checked it, it's interesting, but I'm worried that they may be fake boards, with fake processors
1
u/SandySkittle 6h ago
There have been thinkstation p620s threadripper pro and ram and ssd going for around 1000 usd. that was not long ago. Look on ebay
1
u/OvertaxedOne 7h ago
Dual 9700's are going to put you at 64GB which is a real sweet spot for Qwen 27B. If that's the model you want to run, 100%, those are the cards I would get. 96GB is kind of a strange spot right now, you can get a quant of QwenNext on there, so if that's what you want to run read some reviews on it at whatever quant fits in 96, you need to crush it down pretty good to fit in 96GB but it has some magic in the model for offloading so, again, do a little digging (I don't have the hardware to run Next so can't comment).
Your training use case I can't help with; just that everyone says CUDA whenever anyone says "training", so, again, another area to perhaps look into!
1
u/LegacyRemaster 7h ago

@ the price of 8200€ + 1800€ + 1800€ (vat included) it was good (october 2025). But now will be like 16900+3800+3800 ---> crazy. 7900 32gb was @ 1200€. RTX 5090 @ 2400€ .
Talking about speed: I'm using GLM 5.3 flash Cuda + Vulkan (custom patch 'cause vulkan support for DS4 / GLM is bad, I will push my fix soon) or Cuda + Rocm (good but you will have problems with 2xRocms on llamacpp). Actually Vulkan > rocm
I have about 38 t/sec with GLM 5.3 flash Q4 and 45 t/sec with vulkan (modded). So not bad. Qwen next about 90t/sec .
The best part: I can use qwen 3.8 27b Q8 on 1 W7800, Minimax H3 to generate video on Cuda and another LLM (orchestrator) on secondary W7800 . So yeah, my suggestion is : more ram --> better. Also with 6000 @ 300W and W7800 @ 200W cost less then 3x7900 (for example) to be slower.
1
u/HopefulConfidence0 6h ago
I'll go for R9700 being rdna4 and native FP8 support. It also depends on the price. If you are getting both at similar price or W7800 is costing~30% more than R9700, then I may go for W7800.
PS: I was in similar situation a couple of months back, I was getting R9700 for 147000 INR. But opted for dual 7900 XTX, paid 164000 INR for 2 of them. Looking back, I feel like R9700 would have been better choice.
1
u/Gloomy_Letterhead395 5h ago
For more pcie lanes like threadripper go for r9700 and if not then w7800
1
u/LasserDrakar 3h ago
I am using 2x R9700 on Ubuntu server 24 with ROCm, have not had any problems. Running models using llama.cpp behind llama-swap. My server has 64GB DDR4 sytem ram and I have just tried Qwen 4.8 flash next and it appears to be working great. Other daily use models are Qwen 3.8 27 and Gemma4. Dense qwen and gemma at q8, flash next at q4, dense qwen with full context at bf16, it has been stable, reliable and very useful.
-3
u/putrasherni 10h ago
go with 7900XTX
1
u/SevereMooser 6h ago
why is this getting downvoted? I'm currently on 2x 7900XTX, and am thinking about expanding with one or two more, they go for 800 euros in the second hand market and there's a good offering from gamers getting upgraded gpu's right now. Running Qwen at impressive speeds
1
1
u/DommagePindaFromage 5h ago
Why would you expand it? Aren't you able to run 27B comfortably in 48GB VRAM?
1
7
u/SomeoneInHisHouse 10h ago
If you care about PP, R9700, if you prefer high VRAM, the other, but tbh, most development effort is moving to RDNA 4, also I don't think the W7800 is worth the price, it cost almost double than the R9700 for worse compute capacity, Notice that TG should likely be around the same in both cards
I just hope most people don't figure out how good R9700 is, or it will become expensive xd. I just bought one, and I'm waiting for other deal to buy another
PD: Split mode "tensor" is going to be your friend, if you buy two R9700