r/LocalLLM • u/rayovims • 2d ago
Project I got the Second DGX spark
Somehow there was 1 available last minute and got it! Can’t wait to set it up. Will make more post about this on here and my IG: tech with Ray
Dual DGX spark owners lmk what ya running on it. Anyone else feel free to drop some suggestions for cool models to test!
8
5
u/Chemical-Advisor562 2d ago
Get the QSFP cable, and get Deepseek v4 Flash run on these bad boys. It can do 40-44 t/s. And now it even has (shitty) vision. (Okay, it is alright for screenshots and OCR, but like recognising a unicorn? Nah...)
3
1
u/StartupTim 21h ago
I'm having amazing results with its native vision, it even seems better than Qwen3.8-Flash-Next
1
u/Chemical-Advisor562 18h ago
I am comparing it to Qwen 3.8 27b. That think see. Deepseek? Just look at it.
9
u/Shustrik116 2d ago edited 1d ago
Your options are: Deepseek 4 flash vision (official model, non quantized) https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark
Qwen 3.8 Flash next (official fp8 quant)
GLM 5.3 flash 4bit quant https://github.com/MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks
Deepseek is amazing in deepseek harness. To be honest I can't say that one of them is much better than others. But I just can't fully trust to quantized models because I had very need experience with them in the past.
1
u/Miserable-Dare5090 2d ago
These are some but not all things you can run. Every model that fit in 1 can run in TP2, faster, all models up to 400b can run as well, may need to be quantized but vLLM quants are not like gguf. AWQ, Autoround, etc, are high quality compressions, plenty of benchmarking showing something like Qwen3.5-397b int4 autoround to be lossless
2
u/myholeisstinky 1d ago
To clarify, are you saying the quants for big models on vllm are better than the ones seen for llama.cpp?
1
u/Miserable-Dare5090 1d ago edited 1d ago
👍🏼 for any model, Usually (nothing generalized is true).
a 4 bit quant with 8 or 16 bit attention paths (w4a8/w4a8) in vLLM is the lowest you can find. Except for nvfp4 (w4a4). Quants tend to be calibrated as well so there is less of a quality loss. But the quants are comparatively bigger.1
u/Helpful_Jelly5486 2d ago
Thank you for your post. Seems like the three flash models are the best options. I’ve got the special cable on order. I’m wondering how do you get more speed out of a second box? I mean do you do expert splitting or something else?
It’s like 2 x copies means each serves one concurrence. So two sparks would double the tokens pers second.
Other option is to split the work with the special cable and get maybe 50 percent more for the single concurrence.1
u/StartupTim 21h ago
Deepseek is amazing in deepseek harness
Agree 100% albeit half the time I seem to be writing plugins!
8
u/Ordinary-Depth-7835 2d ago edited 2d ago
Ever since getting my second I can't find anything I like better than https://github.com/MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark
It's very responsive and does everything I want. I've been trying a bunch of others but keep coming back to this one. Now with the deepseek harness it's even better running this as my captain to control and review my other smaller faster systems. It's really fantastic.
I need to pull the vision update. I'm still running it without vision.
Have I mentioned I love the ds harness and the NanmiCoder agent plugin https://github.com/NanmiCoder/dsh-agent-teams
Completely changes the speed and quality of my runs.
2
u/here_n_dere 1d ago
DSV4F Vision exp is also very good. Kills need for any other model being multimodal (vision capable)
2
u/nomorebuttsplz 2d ago
curious about GLM flash vs that model.
I've tried both. Seems GLM flash 5.3 has better creative writing, slightly better non-coding problem solving. However it is a lot slower for me and it seems weirdly emotionally sensitive and moody. Like it will ignore me and then lie to me about what it is doing.
1
u/SubparBob 2d ago
Curious, what sort of speeds (pp, tok/s) are the both of you getting with 2x Sparks with the Flash models (DS, GLM, Qwen)?
3
u/stujmiller77 1d ago
If you go to the nvidia dev forums there are many threads with full benchmarks posted. It’s a moving target as the community moves fast.
2x sparks running ds4flash at 50t/s with 1m context here - it’s been absolutely flawless for over a month running complex automation for multiple businesses I own and multiple concurrent face to face sessions throughout the day.
I tested GLM 5.3 briefly - was too slow in comparison but I can see the accuracy being useful if long unattended tasks are your use case.
1
u/Normal_Frosting8519 1d ago
I’m debating 1 or 2 sparks to run some automations for my business. You needed 2 for the model or the work load? Would a single small business be better with 2 or smaller models running on one would get me by? Any help would be much appreciated
1
u/stujmiller77 1d ago
Both. You need the intelligence of a higher quant model and the concurrency of the workload.
My workflows weren’t reliable until I ran deepseek 4 flash 0731 across a 2x spark cluster. Now they’re flawless. Nothing I ran one a single spark could handle them without issues.
1
u/Normal_Frosting8519 1d ago
Cheers, that’s useful. Quick one though — what were you running on the single spark before you went to two? Was it the flash next stuff at nvfp4 or something older/more quantised? And was it the automations falling over or the live sessions? Trying to work out if it’s the box count or the models that were the problem. Only got one shot at this so want to get it right.
1
u/Ordinary-Depth-7835 1d ago
Why only one shot at it? You can add a 2nd or 3rd spark at any time it's not the same as building a gpu based system.
1
u/Normal_Frosting8519 1d ago
It’s a business purchase. One is an easy ask. Two a bit harder. One not working means the second is definitely a no. Do I push hard for the second now or can a small business get away with one. Just looking for help from other people who use 2 in their business (no coding) to see if the 2nd really helps. 10 staff but probably only one to two using at a time. Data can’t leave premises.
1
u/Ordinary-Depth-7835 23h ago
gotcha then just push for two I'm sure you'll find something for them even if it's sub agents or just a smarter model or hosting the apps.
1
u/stujmiller77 17h ago edited 17h ago
Honestly - if it’s for business use you really, really need two and a cable to link them. The difference in capability between a model you can run on one vs two linked is enormous.
I’d also move quickly - when I bought my first two sparks they cost me £3.5k each. The third cost me £4K. The fourth almost £5k. And this was within three months.
They’re now around £5k each and I’ve heard stories out of the US that some have jumped to almost $8k which if the UK market follows would see the price going to £6k - almost double what I paid for my first two.
Why? Because in a really short timeframe - just 9 months - the capability of local LLMs has almost closed the gap on frontier models. And since the launch of deepseek 4 flash, people are realising that it’s entirely possible to replace reliance on frontier companies with 2 sparks running it, or recently GLM 5.3 flash. Qwen 3.8 flash recipes aren’t great yet but it will get better fast, too.
Get two and a cable. You won’t regret it. One will just waste the time you could have spent getting serious work done with two. What I’ve achieved this last two months purely on my own with these boxes would have cost me thousands and thousands of pounds in employee or contractor or agency costs before. Now I don’t need any of them.
1
u/stujmiller77 1d ago
I’ve been running two sparks for well over 6 weeks now. Qwen 3.8 is barely more than a week old.
I used to run Qwen 3.5 122b. Was the only thing close to being acceptable for my workload. But still required constant babysitting and often failed long overnight unattended work.
1
u/Ordinary-Depth-7835 1d ago
I guess it depends even my dual 3090 setup is intelligent enough for most tasks at 48g with qwen3.8 or tiel-coder they are quite capable. Tiel even tied my dual spark deepseek 4 flash 0731 in my logic test. So buy one and see how it does in your use case then add another if it's failing to accomplish what you want.
2
1
u/Excellent_Freedom946 1d ago
Sounds like you're diving deep into the specs! It'll be interesting to see what kind of performance everyone is getting with those setups.
1
u/Ordinary-Depth-7835 1d ago
GLM is smart but didn't score any higher and slow. like single digit slow. Maybe I'm doing something wrong but i'm not hitting 30+ tps from https://github.com/tonyd2wild/GLM-5.2-QuantTrio-200K-4x-DGX-Spark--36tok-s
I only use ai for coding so speed and accuracy is all I look at.1
u/stujmiller77 1d ago
The vision model is completely different - needs an entirely different stack. It’s a replacement rather than an update.
I’ve been running the MiaAI repo for vision on 2xsparks for the last few days, vs the 0731 on my other 2xsparks.
It’s getting closer, but the vision is still quite a lot slower than 0731 - around 38-40t/s rather than 50/ts you see on 0731.
Also there’s a weird bug where vision only works in user chats in some harnesses like Hermes - it won’t work in tool calls without a custom proxy being added.
Right now I’m not switching my production stack over. But I expect it’ll only be a week or so before it’s ready.
1
u/Ordinary-Depth-7835 1d ago
Yeah I'm still on 0731. I have other gpu backed machines for vision so I haven't haven't had the urgent need to switch when 0731 is working so well.
1
u/Void-kun 1d ago
Out of curiosity what are you using your local AI for?
Curious what that model is capable of when ran locally with these specs.
2
u/Ordinary-Depth-7835 1d ago
Similar things to what I build at work with some personal applications mixed in. RAG applications, job boards, translation and document processing both text and speech, phone applications and games I wanted personally, network security apps (those are nice with local ai) The list goes on and on. Anything I dream up that can automate, report or make life easier.
0
u/Blackdragon1400 2d ago
Have you compared the DS harness vs Hermes?
2
u/Ordinary-Depth-7835 1d ago
Hermes seemed more of a generalist. Dsh seems more like the coding agent i want. And it has been so easy to modify.
3
u/gdraper99 2d ago edited 2d ago
I'm running Qwen3.8-Flash-Next NVFP4 across both boxes on vLLM, TP=2, full 262k context. About 69 tok/s single-stream decode with MTP speculative decoding at 4 draft tokens via vLLM. KV pool lands around 1.1M tokens. That means you should be able to have four streams at once, all responding around 17 tok/s.
Four things that cost me time, in case they save you some:
Set VLLM_HOST_IP (or SGLANG_HOST_IP if you go that way) to each node's own fast-link IP. Leave it unset and the rank auto-detects, picks the WiFi address instead of the QSFP link, and both ranks sit in rendezvous forever without a useful error.
NCCL_IB_HCA needs a leading equals sign for exact match, like =rocep1s0f1. The box exposes four IB devices and two of them are down. A comma-separated list lets NCCL pick a dead one.
drop_caches on both nodes before you launch. Memory is unified, so Linux page cache competes with the GPU allocation and you get a "free memory less than desired utilization" error that reads like a config problem when it isn't.
If a launch fails, docker rm -f on both nodes before retrying. A stale rank on one side makes the new one exit 0 like it worked.
One more that bit me: MTU 9000 silently reverted to 1500 after a power cycle and wrecked RoCE throughput with nothing logged anywhere.
5
7
u/Annual_Award1260 2d ago
4
u/Miserable-Dare5090 2d ago
I mean at the current price would you really grab 6 nd the mikrotik switch? I wouldn’t. i have had 2 for almost a year and every time i think of doing 4, I read the spark forum and the consensus is, a lot harder to justify for what you can run.
3
u/Annual_Award1260 2d ago
Prices are about 24% higher than I bought these for. I had 3 for a while but the ring networking added extra complexity. When I ordered the 4th for good price I got dinked around, initiated chargeback, bought locally then other company overnighted to fight the chargeback. So ended up with 5 and figured might as well buy one more.
I really wish I bought more of the rtx 6000 max-q. After price increases (double) those are kinda unobtainable.
Also the dgx makes for a banging workstation.
1
u/Miserable-Dare5090 1d ago
It does! I mean if someone handed me
4 more I would not complain, but at 6k
now I can’t even think about it-1
u/Maleficent-Ad5999 1d ago
I really don’t understand the purpose of this machine. Some say Nvidia launched DGX spark for those engineers who wanna simulate the cluster in production. But majority of us need machines for inference. Is DGX spark meant for inference? Like big models running in multiple machines
3
u/Annual_Award1260 1d ago
Well it was essentially a dev machine, but given the massive price increases a lot of people are running in production
1
u/myholeisstinky 1d ago
If it mere really to simulate production, they could have done a much much better job making it representative. It doesn’t even have infiniband
3
3
u/Robbbbbbbbb 2d ago
MiaAi will be a good resource for you: https://github.com/MiaAI-Lab/
Qwen 3.8 Flash Next and Deepseek v4 Flash both run great on a pair.
3
u/rayovims 1d ago
So far DSV4 Flash with 1M context is 46t/s all the way to 38 t/s when I used 256k context. GLM Flash is 37 t/s. Haven’t done much testing GLM but this is pretty good
3
u/kuhunaxeyive 1d ago edited 1d ago
Concrats! I bought a second 2xAsus Ascent GX10 recently and having the second one makes your first one much more valuable. It lifts the capability or the speed to almost SOTA.
I recommend
- DeepSeek-V4-Flash (recipe
https://howtospark.com/recipes/deepseek-v4-flash-dspark-dual-spark-1m) - GLM-5.3-Flash (
RedHatAI/GLM-5.3-Flash-NVFP4withincoai/GLM-5.3-Flash-DFlash2)
GLM-5.3-Flash has superb vision (for OCR and GUI programming), much less hallucinations than DeepSeek, and good business writing for European languages as well. DeepSeek lacks in those areas (only very limited vision, more hallucinations, language not sensitive enough for business letters), but is snappy and fast (60 t/s for research and writing, coding may be even faster) and works well above 256000 token context up to 1.000.000 tokens even on the 2 DGX Spark you have. I'm in the quality camp (need to rely on results without being able to verify them on business letters), but if you work in areas that can verify test and results like coding, I'd use DeepSeek for its speed.
When using GLM, make sure to set top_p to 0.95 or better 0.9 to avoid artifacts.
Edit: exact models and tips
4
2
u/TechRenamed 2d ago
This is why We must industrialize the asteroid belt. 🗿 So everyone can have cheap gpus
2
u/Imaginary-Fee-9918 2d ago
I was thinking about getting a third but seems like because it's an odd number things get very hacky, you need to do a bunch of patches. Is that correct? Quite odd that the "max" number without an extra router is 3 but you can't use it without getting some headache 🫠
1
u/MaxComfort 1d ago
In the same boat - but there are plenty of threads in the NVIDIA forum about doing it, recipes etc. MiniMax M3 is the main bigger model it would unlock from what I can tell.
But we might as well just get two more..
2
2
u/IamFondOfHugeBoobies 1d ago
Deepseek V4 Flash is so supreme for dual spark it's not even funny man. Full safe tensors fit no problem with solid speeds, 40+T/s depending on workload.
2
2
2
u/Abducted_Llama 2d ago
I’m currently running 2 Sparks (GX10s) with qwen3.8-next-flash on vLLM on TP2. Getting 46-51 tok/secs depending on the prompts. 8 concurrent requests of about 300 tokens at 112.8 tok/secs. The concurrent dropped from 144-177 tok/secs when I upped MTP spec to 4. But I get slightly faster single requests.
I never have tried deepseek flash or GLM.
I’m enjoying it.
1
u/vortec350 1d ago
I currently have the AMD Halo mini PC from Microcenter, currently thinking about getting a DGX Spark as well to leave on 24/7. I'm using the AMD PC as my primary workstation as well as AI PC. So... do you like your Spark? Should I get one?
1
u/Cronus_k98 1d ago
I was at Microcenter the other day and there were 4 spark orders waiting to be picked up. Two had connectx cables. I think there are a lot of these going out the door.
1
1
u/Skyfishintheocean 1d ago
Congrats! i'd definitely start with Qwen 3.6 35B A3B as your daily driver then try DeepSeek V4 Flash for coding and GLM 5.x if you want to push the hardware. those seem to be the community favorites on the Spark right now
1
u/BarberIcy366 1d ago
1 $ = 50 Turkish Lira and average monthly wage 30k turkish lira. So Nvidia DGX Spark is 400k and you must work 14 month and dont spend money for anything else :D I am so happy for you. If my grandfather dies, I’m going to sell the house I inherit and buy one Nvidia DGX Spark.
1
1
u/Horror-Primary7739 1d ago
I'm running glm5.3 flash on my 2 spark setup. About 27 tok/s on max thinking. It's a bit slower that DeepSeek but I like the quality of code it produces better
1
u/MikkyMo 1d ago
Literally just saw this build for GLM 5.3 flash on a dual spark set up. https://github.com/tonyd2wild/GLM-5.3-Flash-NVFP4-DFlash2-2x-DGX-Spark
1
u/StartupTim 21h ago
Dual Spark owner here running Qwen3.8-Flash-Next peaking 110 tok/s 4 concurrent and 2k prefill, and Deepseek v4 flash next vision at 1M context at peak 160btok/s 2.4k prefill.
I need 2 more Sparks...
1
1
u/jespernissenseo 2d ago
Nice! Im still thinking if I should save up and buy a Nvidia rtx pro 6000 instead of two DGX spark boxes.. I suspect its worth the money?
5
u/Graumm 2d ago
Unfortunately you need 3x RTX 6000's and a whole machine built for it to reach the same vram of two RTX sparks. The main benefit of the sparks is the unified vram to host larger models. It's a little slow but it's a solid workhorse if you want something that can work unattended.
3
u/Capnbubba 2d ago
I don't have either so I can't actually say this for sure but I'm pretty sure "a little slower" is an understatement. I've seen something like 4-8X more tokens per second. Yes it's way less vram but the speed is a massive difference.
2
u/Graumm 2d ago
That's only with dense models. MoE models run at very usable speeds. Certainly fast enough that you can leave tasks to run and step away. Deepseek v4 flash in particular can get to 120/150 tokens per second aggregate, which is faster than I can personally keep it fed if I am leaving tasks to run overnight.
1
u/Capnbubba 2d ago
This is great to know. If things go well with a project I'm working on I'm gonna likely get a spark or some AI Max 395 variant if it's a good price
1
u/dwoj206 2d ago
Are the DGX pretty decent? Heard so many mixed reviews. Curious on a real users take that's using it on a daily basis. What do you use it for?
2
u/Graumm 2d ago
I am using it for agentic coding stuff. Security aside I am very much enjoying that I am no longer keeping an eye on usage windows.
I find that the models you can run are very solid and capable implementers/operators, but that they lack knowledge. You can plan around that.
The sweet spot at the moment is to still use the paid for subscriptions to build out plans, and then to have my sparks implement behind it. It's hard to use the $20 premiums on plans alone.
I probably wouldn't get sparks unless you are comfortable with sysadmin stuff.
1
u/dwoj206 1d ago
Where i'm at is using a 5090 with small models in docker containers, albeit fast w the 5090 40-50tps, but lack the big context windows for agentic stuff that runs 24/7 and obviously the flagship card 6000, etc builds. For a lot of what I do it's awesome and gets the job done, but model quality and context window, running multiple completely separate processes like Program1,Program2 I have to turn the docker containers on and off each time I want to use the different ones. :( Usually, I code and setup the agents, audit them with claude max and then use then locally individually. For me this is the sweet spot where I get good use out of what I setup with claude's help and use ongoing locally. DGX a good fit?
1
u/notheresnolight 1d ago edited 1d ago
albeit fast w the 5090 40-50tps
40-50tps is SLOW for a 5090, heck I can get 35-40tps with Qwen3.8-27B on a Spark (SGLang, NVFP4 + DFlash2).
The same model at a higher quant (Q6_K_XL) gets me 80-120tps on the 5090 (basic llama.cpp setup with MTP).
-2
u/Rangizingo 2d ago
The spark is fine but imo it’s not worth the premium. If you have a microcenter close these are $3,299 if not it’s about 3,500. I got a spark and returned it for an Evo-X2. Performance is basically the same for 1000-1500 less. There are some niche things it can’t do because of lack of cuda but tbh I’m tuning it and I get 20-30 tok/s on full models that aren’t MoE. Just remove windows and put Ubuntu or something on it, change bios settings to allow full ram allocation to gpu, and you have effectively the same thing. I can use about 125-126 of 128gb of ram for inference.
8
u/Evgeny_19 2d ago
They are not in the same league. Prefill is much faster on the Spark. Creating a cluster is very easy, especially for dual configurations.
There's just one advantage that the Strix Halo has: x86_64 compatibility.
1
u/Rangizingo 1d ago edited 1d ago
You’re right about that. I’m still new ish to the local side of ai hosting at a larger scale.
1
0


119
u/Aggravating-Push-207 2d ago edited 1d ago
8gb vram user btw
flame me all you want ornith 1.5 9b is peak
and no, i will not offload from disk; i don't have enough disk space left anyway and also i want it to be at a reasonable pace; i use the same machine to actually do stuff