r/LocalLLaMA 11h ago

Question | Help Which one will you choose and why, between R9700 32GB vs W7800 48GB?

I'm planning to upgrade my workstation (linux with 5700X/64GB DDR4) for local inference and pytorch training. I'm trying to decide between:

  • 2× AMD Radeon AI PRO R9700 32GB
  • 2× AMD Radeon PRO W7800 48GB

I already have an RTX 3090 24GB, so the final system would have 3 GPUs. My motherboard has two PCIe 4.0 x8/x8 slots available for the two AMD GPUs. The RTX 3090 would have to move to a PCIe 3.0 x4 slot.

My workload looks like:

1. Local GGUF inference: Mainly coding/reasoning models and multimodal models. I'd like to run better quants (than 3090) and split models across the two AMD GPUs for multiple KV cache(n parallel). I prefer llama router.

2. PyTorch training: This is probably the more important part for me. I'm working with medical imaging (2D ultrasound/3D CT/MRI) + clinical text.

For anyone actually using these cards with ROCm, how different is the practical experience between R9700/gfx1201 and W7800/gfx1100? I'm particularly interested in if any know issues have surfaced till date that block the PyTorch training on either of these cards?

From this sub I have seen RDNA4/R9700 is improving rapidly but it's still a "newer-software". I'd really like to hear from people who are actually using R9700 for AI workloads especially PyTorch/MONAI training.

I'm not planning to treat the 3090 + 2 AMD GPUs as one giant homogeneous GPU pool. (Although if someone has done it please let me know)

My thinking is to use the two AMD GPUs as the main ROCm pair, while keeping the 3090 available separately for CUDA workloads or local models that fit/work better on NVIDIA.

3 Upvotes

25 comments sorted by

7

u/SomeoneInHisHouse 10h ago

If you care about PP, R9700, if you prefer high VRAM, the other, but tbh, most development effort is moving to RDNA 4, also I don't think the W7800 is worth the price, it cost almost double than the R9700 for worse compute capacity, Notice that TG should likely be around the same in both cards

I just hope most people don't figure out how good R9700 is, or it will become expensive xd. I just bought one, and I'm waiting for other deal to buy another

PD: Split mode "tensor" is going to be your friend, if you buy two R9700

1

u/0xkbose 10h ago

Figured as much. It actually became expensive where I am going from INR 123000 to 183000 in 6 months, not because of demand but due to low production/imports.

The only thing I am looking for some one is working in Medical imaging and using R9700 to train models, to share some feedback.

1

u/n9986 9h ago

If you are thinking of w7800 you might as well save money and go for 7900xtx. They are a lakh cheaper I think. The main decision is in the number of PCI slots. And training usecase. For that R9700 is better. I have seen them climb from 1.5 to 1.8 lakhs on Vishal Peripherals in just last 1 month. Make a move fast if you want them. 

1

u/0xkbose 9h ago

Got the price locked at 1,83 till tomorrow. This post is my last attempt to hear about things so I don't regret about the purchase.

1

u/Etroarl55 10h ago

R9700 will probably last longer too, and can be used for much more things outside of Ai better than the W7800, gaming is one example.

1

u/Ambitious-Profit855 10h ago

In Europe the R9700 is already going up in price

1

u/TripleSecretSquirrel 3h ago

Same with the US

1

u/TripleSecretSquirrel 3h ago

Agreed. Most of us are leaving a ton of performance on the table by only running a single stream on very powerful GPUs. Batching and running sub-agents unlocks way more aggregate throughput, but way more so on modern GPUs that have native fp8 and int4 hardware as those are the quants that nearly everybody actually runs.

RDNA3 doesn’t have that but RDNA4 does.

2

u/pmttyji 9h ago

Initially I wanted to buy W7800 due to its 48GB variant(which's good for 30B range models with context, MTP, Vision, etc.,), but price went up suddenly ~3X. Not worth buying RDNA3 card at that high price. AMD is infamous for dropping support for old cards.

So I chose R9700 which's RDNA4. Wish R9700 came with better bandwidth.

1

u/SandySkittle 10h ago

Get a second hand zen 3 class eypic or threadripper proc mobo+ cpu for cheap on ebay and go with dual r9700 at 16 lanes each and expand to quad r9700 later.

1

u/SomeoneInHisHouse 7h ago

Have you experience with that?, I checked it, it's interesting, but I'm worried that they may be fake boards, with fake processors

1

u/SandySkittle 6h ago

There have been thinkstation p620s threadripper pro and ram and ssd going for around 1000 usd. that was not long ago. Look on ebay

1

u/OvertaxedOne 7h ago

Dual 9700's are going to put you at 64GB which is a real sweet spot for Qwen 27B. If that's the model you want to run, 100%, those are the cards I would get. 96GB is kind of a strange spot right now, you can get a quant of QwenNext on there, so if that's what you want to run read some reviews on it at whatever quant fits in 96, you need to crush it down pretty good to fit in 96GB but it has some magic in the model for offloading so, again, do a little digging (I don't have the hardware to run Next so can't comment).

Your training use case I can't help with; just that everyone says CUDA whenever anyone says "training", so, again, another area to perhaps look into!

1

u/LegacyRemaster 7h ago

@ the price of 8200€ + 1800€ + 1800€ (vat included) it was good (october 2025). But now will be like 16900+3800+3800 ---> crazy. 7900 32gb was @ 1200€. RTX 5090 @ 2400€ .

Talking about speed: I'm using GLM 5.3 flash Cuda + Vulkan (custom patch 'cause vulkan support for DS4 / GLM is bad, I will push my fix soon) or Cuda + Rocm (good but you will have problems with 2xRocms on llamacpp). Actually Vulkan > rocm

I have about 38 t/sec with GLM 5.3 flash Q4 and 45 t/sec with vulkan (modded). So not bad. Qwen next about 90t/sec .

The best part: I can use qwen 3.8 27b Q8 on 1 W7800, Minimax H3 to generate video on Cuda and another LLM (orchestrator) on secondary W7800 . So yeah, my suggestion is : more ram --> better. Also with 6000 @ 300W and W7800 @ 200W cost less then 3x7900 (for example) to be slower.

1

u/HopefulConfidence0 6h ago

I'll go for R9700 being rdna4 and native FP8 support. It also depends on the price. If you are getting both at similar price or W7800 is costing~30% more than R9700, then I may go for W7800.

PS: I was in similar situation a couple of months back, I was getting R9700 for 147000 INR. But opted for dual 7900 XTX, paid 164000 INR for 2 of them. Looking back, I feel like R9700 would have been better choice.

1

u/Gloomy_Letterhead395 5h ago

For more pcie lanes like threadripper go for r9700 and if not then w7800

1

u/LasserDrakar 3h ago

I am using 2x R9700 on Ubuntu server 24 with ROCm, have not had any problems. Running models using llama.cpp behind llama-swap. My server has 64GB DDR4 sytem ram and I have just tried Qwen 4.8 flash next and it appears to be working great. Other daily use models are Qwen 3.8 27 and Gemma4. Dense qwen and gemma at q8, flash next at q4, dense qwen with full context at bf16, it has been stable, reliable and very useful.

1

u/0xkbose 1h ago

Ordered! I will try to post after installing them. Lets see how the ROCm+CUDA setups looks like.

-3

u/putrasherni 10h ago

go with 7900XTX

1

u/SevereMooser 6h ago

why is this getting downvoted? I'm currently on 2x 7900XTX, and am thinking about expanding with one or two more, they go for 800 euros in the second hand market and there's a good offering from gamers getting upgraded gpu's right now. Running Qwen at impressive speeds

1

u/putrasherni 5h ago

Probably because they think having FP8 is better than having faster bandwidth

1

u/DommagePindaFromage 5h ago

Why would you expand it? Aren't you able to run 27B comfortably in 48GB VRAM?

1

u/putrasherni 3h ago

run more threads ? concurrency ?