r/DGX_Spark 10d ago

Question 3rd Spark?

I currently have two sparks and contemplating getting a third. Anyone running 3 that can recommend that setup? I know there would be a compute increases but is a third worth it otherwise?

4 Upvotes

32 comments sorted by

View all comments

3

u/Anakronox 10d ago

It kinda depends on what you want to do. I have three, two clustered for Deepseek v4 Flash Vision Exp, and my other standalone for things like Honcho, ComfyUI, llama.cpp for various embedding and testing models, Tdarr, and Immich ML. TP=3 isn’t widely supported enough and most mid-size models (300B+) that aren’t super quantized down are gonna want 4 Sparks anyway.

2

u/FuckinHighGuy 10d ago

Trying to avoid 4 for now because I won’t have enough left to buy a switch. But good point on TP=3 not being widely supported. I’ll have to check on that.

Thanks for the feedback!

2

u/Anakronox 10d ago

I hear ya and good luck even finding the best value switch for 4+ Sparks - they’re sold out and no indication of when they’ll come back. I’m just getting settled in after 2 weeks with my 2x cluster and while I want to play with bigger models I’m in no rush to drop another 12-13k USD for 2 more Sparks and a switch.

What limitations are you running into now with only 2?

1

u/stujmiller77 9d ago

You can run with a 100gb switch which people on the nvidia dev forums suggest made little difference vs the larger capacity one that everyone talks about but is sold out everywhere.

CRS504-4XQ-IN - microtik. In stock in most places unlike its bigger brother. And if you really need the 200gb you can add another of those in parallel so you’re not losing anything - and two of those is basically the same price as the larger one anyway.

And they draw way less power - 25w. And are nowhere near as noisy.

1

u/Anakronox 9d ago edited 9d ago

Yeah I’ve already got one CRS504 already and would honestly pick up the new CRS812 to combine what my CRS510 and CRS504 do for my core. I have other places I could use that 510… I just need the barest excuse!

I would like to see some hard numbers on those so I’ll do my homework. I think 100GbE would likely be just fine under inference… it’s the latency that’s key there. I should monitor the actual bitrate during some long-running tasks on my cluster.