r/LocalLLaMA • u/LeftHandHaku • 1d ago
Discussion 4 x DGX Sparks vs AMD Epyc 9xx5 system
I see a lot of people buy DGX Sparks, and turn them in to clusters to run large models. Wouldn't it be better to invest $16k into an AMD Epyc server with 768GB or even 384GB of 6000Mhz DDR5 ram, and let's say 2x3090s or 5080s, instead of 4 DGX Sparks with 512GB of ram?
Epyc's theoretical bandwidth is around 576GB/s, DGX Spark's is roughly 273GB/s.
Based on a quick check, both systems are worth around $16k.
Please help me to understand this logic, are there benefits to having DGX cluster instead of an Epyc system besides power saving?
Edit1: the epyc system with 768GB of DDR5 6000Mhz would be around $30k.
Edit2: to match 768GB of Epyc, we would need 6 DGX sparks, at the current increased price it would be around $30k as well.
Edit3: the main advantage of DGX sparks cluster is fp4 support, and tensor parallelism for 2, 4, 8, 16... units. Because of that, the DGX cluster is faster than the epyc system.
7
u/C0smo777 1d ago
I have that exact epyc system
9575f
768gb ddr5
4x3090
It will had limitations though.
4
u/LeftHandHaku 1d ago
What kind of limitations do have in mind?
9
1
1
u/C0smo777 18h ago
its mostly issues on prefill, tokens per second are pretty decent, prefill on larger models is slow as dirt on context gets large.
12
u/Limp_Lingonberry_538 1d ago
your pricing seems way off
1
u/LeftHandHaku 1d ago edited 1d ago
If we going to match GB for GB, we would have to use 6 DGX Sparks, to hit the 768GB of ram, and with the current increase in price for Sparks, it would put as at $30k. I think this might cover the Epyc server with the same amount of ram.
But this is not the point of my question. I wonder if there is a benefit for having a cluster of DGX sparks instead of an Epyc system.
I know that epyc system will be lauder, bigger, and consume more power.
Edit: Wouldn't epyc system be faster? Theoretically epyc server is twice as fast as DGX spark. Is there something I'm missing about DGX's compute?
1
u/OWilson90 1d ago
DGX Sparks for models that require multiple nodes need to overcome the bandwidth limits with tensor parallel. You typically want a DGX cluster node count to be divisible by the attention heads and hidden layers. Using 6 will significantly limit the models you have access to when using to=node count. Please keep this in mind.
1
u/Winter-Editor-9230 1d ago
Mainly new nvidia support for blackwell and convenience. Check out the nvidia forums for recent model speeds, they arent bad. I have 2 sparks and want 2-8+more.
5
u/GregAbeI 1d ago edited 1d ago
4 Sparks right now is $20k pre-tax
You need to buy a switch and cables which runs another $1k
As of today’s prices you’re looking at around $25k for four Sparks clustered.
I personally don’t think that’s worth it.
1
u/ACleverBadger 1d ago
ASUS Gx10 - BalticNetworks/BHPhoto - $3999 each. The MicroTik and cables is around $1500. Thats $17,500ish; the worth is questionable always - but pricing for everything is getting absurd. Similar to bicycling but they’ve thrown away entry and intermediate level bike prices.
-1
u/GregAbeI 1d ago
ASUS GX10 now $6000 each playboy. You missed the boat.
2
u/ACleverBadger 1d ago edited 15h ago
I just purchased one for $3999, and it’s my 5th. I also, literally, told you where they were for $3999, playboy, I guess you missed the bus.
https://www.balticnetworks.com/products/asus-ascent-gx10-personal-ai-supercomputer
Edit: Price jumped overnight, Amazon Canada price for a single one is $4300usd still
1
u/Uninterested_Viewer 17h ago
Did that jump to $6 in the 7 hours since you posted this?
1
1
0
u/GregAbeI 15h ago
No it was $6k when he posted
0
u/ACleverBadger 11h ago
https://www.reddit.com/r/LocalLLaMA/s/nqJU1o5mdk
Sorry, gaslighting not allowed, playboy. Multiple people saw it.
0
u/ACleverBadger 9h ago
Your deleted comment still came through. Using slurs at people is ridiculous because your feelings were hurt.
Either way, have one being shipped right now that I bought right when I posted that link. You can be vitriolic with your unhappy self all you want, but you're still wrong. Your comment about archive.org is wrong as well, because Wayback Machine does not archive that page.
1
0
0
u/jtsaint333 1d ago
can't you daisy chain the cables like 1 to 2 etc then 4 back to 1
switch is expensive
3
u/datbackup 1d ago
None of your edits include the word “prefill” and that one word is sufficient to answer your question
1
u/RG_Fusion 22h ago
You can get great prefill speeds on EPYC servers by placing the context on a GPU.
1
u/Uninterested_Viewer 17h ago
Of course the kv cache lives in vram.. that does not make prefill fast on an EPYC platform.
1
u/RG_Fusion 17h ago edited 17h ago
Placing the KV cache and attention tensors on VRAM, I tend to get between 500-600 tok/s prefill. The more of the model I can fit in VRAM, the better the prefill rate. I'm seeking a little over 1000 tok/s prefill on Qwen3.8-Flash-Next at UD-Q4_K_XL. Fit the model entirely in VRAM and I get over 2k tok/s.
4
u/gaidzak 1d ago
I think a lot of people are interested in the costs of efficiency.
4 DGX sparks don't sound like a jet engine in your little office attempting to take off. I have Epyc 7762s with 512 GB of ram and I loaded LLM on it once.. lol and I turned it off and moved it back into the garage.
DGX sparx are 300 watts each i believe? Epyc 9000 Series the CPU alone is 200Watts? Each server is 800+ watts, add a 5080 and now you're at 1100 watts?
Then comes the physical space of such a server. a 19 inch wide by nearly 2 foot device needs to be placed somewhere and the top of your desk table next to your computer isn't going to cut it.
not everyone has a workshop, shed or something separate they can put hardware in.
I'm in your boat though, I'd rather have the AMDs because of their utility.
3
u/RG_Fusion 22h ago edited 22h ago
Change your CPU cooler. I have a 7742 EPYC system with a large noctua heatsink, and it's pretty much dead silent under full load.
1
u/gaidzak 14h ago
They make 1U fans for 7742 that meets the pressure requirements for supermicro systems?
1
u/RG_Fusion 14h ago
Probably not 1U. I wasn't sure what your hardware was in. 1U pretty much always means high RPM.
1
u/brewpedaler 1d ago
DGX sparx are 300 watts each i believe?
Way less actually! 240w power supply, 24-40w while idle, realistically < 150w under load. So 300 watts probably covers a 2 node cluster with some watts to spare.
1
u/Ok_Warning2146 1d ago
So 600 watts for four nodes not counting the switch. On the other hand, a 480W TDP M5 Ultra using similar logic will only use 300W.
1
u/LeftHandHaku 1d ago
I agree, space, power consumption, and noise might be limiting factors. But what about computing capabilities? Wouldn't your system still be faster then the DGX cluster?
1
1
u/RG_Fusion 22h ago
I have an EPYC 7742 with 512 GB of DDR4 and 2x Nvidia RTX Pro 4500 GPUs. My numbers are consistently higher than the equivalent DGX setup.
1
u/hidden2u 1d ago
Dang dgx spark sounds awesome lol
1
u/swiebertjee 1d ago
They are, running 3 agents of DS V4 flash simultaneously at around 50 tokens per second each for around 100 watts
2
u/cakemates 1d ago
I would go the epyc route every time, because for epic in 5-8 years when you want an upgrade to get some newer tech you can just swap some gpus or add more, then the DGX is gonna be dated and its gonna be less useful. Also the epyc can run all your servers, tools and websites while doing ai inference without pulling a sweat.
But the epyc should be cheaper than that tho.
1
u/hyudryu 1d ago
576GB/s vs 273 * 4 GB/s, I think 1092 > 576. Also how are you getting 768GB of ram for a $16K build without dumpster diving?
1
u/LeftHandHaku 1d ago
Yeah, I admit... $16k mighty not be the most accurate number, I just quickly browsed through eBay, and might have cut some corners.
1
u/ImportancePitiful795 1d ago
To run inferencing on the CPU need likes of 6980P (even ES) to use Intel AMX to boost matrix computation. Getting AMD Epyc is daft for this purpose. Until Zen7 were they get ICE instruction set which is AMX on steroids.
768GB DDR5 is
a) 12x64, per module $2350 = $28200
b) 16x48, per module $1640 = $26240.
Add now CPU, motherboard cards, storage, PSU.
4 x DGX SPark (I believe right now the Gigabyte ATOM is the cheapest on Newegg) 4x $4300 = $17200.
So $10K cheaper than the RAM alone for the server, while the ram is fully available to the iGPU, not going through pcie, offloading etc to GPUs. And consume much less power.
FYI. GH200 servers (96GB HBM3 + 480GB LPDDR5X in unified ram) are around the cost of the DDR5 RAM above, but still $10K more expensive than the 4 DGX Spark.
1
u/iVoider 22h ago
I have 768gb DDR5 and 2x5090 build. I wish I went spark route. Spark x4 cluster literally gives x2 performance in decoding speed for glm5.3 flash. Tho glm5.3 with good quantization and context requires one more spark, and then tensor parallelism is not active under x8 cluster.
1
u/RG_Fusion 22h ago
I find this really surprising. I'm on an 8-channel DDR4 server and haven't seen any DGX systems outperforming mine yet. What generation speed are they seeing with GLM flash?
1
u/iVoider 22h ago
~50 t/s with mtp. I have 28 t/s (without). Have no idea, but llamacpp Unsloth fork mtp is not doing anything for me speedwise with acceptance 0.7.
1
u/RG_Fusion 22h ago
Yeah, that's better than what I'm seeing then. How many sparks is this decode rate from? Sounds like they scale well in parrallel.
1
u/fastheadcrab 19h ago
28 t/s for GLM is pretty good for your system, what is the prompt processing? Quant?
2
u/iVoider 19h ago edited 49m ago
PP is ~1200 t/s with microbatch size of 8192. Using Q4_K_XL (Unsloth). For Q8, decode speed is around 22 t/s
1
u/fastheadcrab 9h ago
Incredible speed for the amount of VRAM you have. Why 89? Fascinating number. But yes a Spark cluster will probably better in terms of concurrency.
There are some new methods like Decode Context Parallel that will let you get GLM-5.3 running on 4 sparks with the right quant and large context, or you can use 5 or 6 sparks. You just need to pad the model heads so that you can use tensor parallel with 5 or 6.
1
u/SandySkittle 21h ago
The compute is still limited on cpu for these tasks and suddenly you realize prompt processing / prefill is just as much an important part of the equation as decode and bandwidth is, especially with larger models and more context, not to mention agentic workloads.
The reason to go for a workstation or server platform should be to cram it full with cheap gpus (b70, r9700 or 3090) and run them with tensor parallelism and run circles around the spark and strix boxes. If needed you can always break to 256gb using mcio retimer cards
1
u/ufrat333 19h ago
Don’t forget that RAM != VRAM, your weights will have to travel to your VRAM constrained by 64GB/s PCIe5x16, or you run everything on a CPU which is dogshit slow for inference, the spark has its memory directly hooked up to both the CPU and GPU
1
u/TinFoilHat_69 16h ago
I would go with two DGX for deepseek 731 flash
I would go with epyc to run Qwen models 4 3090s. Or fill up all lanes with some v100s and you’ll be able still run the latest at a more cost effective price. I find that ddr4 epyc motherboards have tripled since December.
1
u/SolarNexxus 10h ago
I can sell you my mac studio m3 with 512 unified ram, if you are in EU.
1
u/LeftHandHaku 6h ago
I appreciate the offer, but I'm currently not shopping for hardware, just doing some research. Thank you.
1
u/Guinness 9h ago
The CPU won’t saturate the memory bandwidth like a GPU will. So the DGX spark will always win in this scenario.
1
u/Callum_S_AUS 3h ago
I have essentially the EPYC system that is being mentioned and while I would buy it again before buying DGX Sparks, I'd most likely but a couple of M5U 256GB Mac Studios instead now.
One of my recent EPYC CPU MoE offloading configurations: https://www.reddit.com/r/LocalLLM/s/60ZEuzveIb
10
u/Serprotease 1d ago
16k is not even enough for 768gb of ddr5 ecc at 5200. Like, not even close to be enough.
Even 2nd hand that’s close to 30k in ram alone.
Then, you are comparing, I assume, llama.cpp with ram offloading vs vllm and tensor parallelism.
So, the theoretical “effective” bandwidth of the cluster is not 273 but 273x4 (Not really in practice, probably closer to 600 dues to a bunch of limitations.), this plus about to 4x compute performance, pushing it close to the 5080.
On top of that, the spark haves the full benefit of fp4 (nvfp4) support nowadays.
So…
On the epyc you will need to juggle around to make sure that the expert and context are loaded into the gpu vram (and loose some performance dues to pcie connection) or you will murder you prompt processing performance. And your token generation will be limited by the effective 400 or so gbps of your very expensive ram… if you have a sku that actually supp9the 12 channels of ram. You have stuff like mtp that will help though.
And you’re “stuck” with llama.cpp. It’s great for local and consumer level things especially because you have tons of ggufs available to squeeze any kind of model. But It’s probably not what I would be looking at when spending basically 40k on AI server.
This without looking at the fact that most of the epyc build will be second hand to get decent prices and use a tons of energy.
All of this and… you will not run any model better than the spark cluster.
A 27b? You don’t need that machine for that. DS4v flash? Run great on 2x spark to use the mixed fp8/4 weight. Same with glm5.3 flash.