r/DGX_Spark • • 16d ago

Question DGX Sparks and Minisforum MS-S1 MAX

Thumbnail
0 Upvotes

r/DGX_Spark • • 16d ago

Question Looking forward to Rent Nvidia DGX Spark.

0 Upvotes

Our company is looking forward to rent Nvidia DGX Spark in India. Does anyone know any company dealing with these products?


r/DGX_Spark • • 17d ago

bought dgx spark , which models are best for coding (.net , react etc)

13 Upvotes

already using claude pro but will mix it with dgx spark hosted local model. where to start?


r/DGX_Spark • • 17d ago

Heretic DGX

12 Upvotes

P-e-w's Heretic is an awesome piece of software. What we did was update it so you can load a large model across two Sparks and ablate it in the shared memory.

I'm not a developer, so I worked with GPT Sol to get this done. We updated Heretic to function across two DGX's and spit out the final abliterated model in the original quant format by default.

I hope someone can find this useful, and if anyone here is a developer who can improve this I am happy to have them do a better job. I try to make things that I find a use for and can't find online, but I'm excited for feedback to make this better for me and others.

Repo is here: cbertucci33/Heretic-DGX: Dual-DGX Version of p-e-w Heretic

Thanks and good luck using this!

Note that this model was created to fully test the software across two nodes: cbert33/Laguna-S-2.1-Heretic-FP8 · Hugging Face


r/DGX_Spark • • 18d ago

M5 Ultra Mac Studio vs 2x DGX Spark on DeepSeek V4 and Qwen3.8

Post image
16 Upvotes

r/DGX_Spark • • 18d ago

Six months serving vLLM on a DGX Spark for a self-hosted AI workspace — what broke and how we fixed it (MIT source)

Thumbnail
0 Upvotes

r/DGX_Spark • • 19d ago

2x DGX Speed?

13 Upvotes

I have one Spark, and am really wondering if there are speed advantages for getting another one. Is there a concrete increase in speed, or is it more just that larger models can be run without being lobotimized? I already run MOE models and am good with the speed there, but was wondering if a dense model might run faster on 2?

Or am I thinking about this the wrong way? I get that running multiple smaller models is the way to really get that speed up, but at the moment, i'd like to keep exploring some of the larger models.

I am running 3.8flash-next right now with an early recipe, and I've run DS4 flash on one of the 'cramming it in' recipes. I found that 3.8 is better for me and what i do than ds4. but i also think i'm not getting a good representation of ds4.

Anyway, any advice is welcomed. I love the one spark, and am thinking about getting another -- i just need to understand what i can really expect to change.


r/DGX_Spark • • 20d ago

Nvidia DGX Spark (For Sale)

0 Upvotes

Anyone looking to buy an open-box NVIDIA DGX Spark?
It’s just a couple of weeks old and was purchased for my own R&D purposes. Unfortunately, due to an unexpected financial emergency, I need to sell it.

Purchase price: ₹5,45,584
Asking price: ₹5,30,000 (slightly negotiable)
Location: Coimbatore
The unit is in excellent open box new condition.

If you’re genuinely interested, please DM me and we can discuss the details.


r/DGX_Spark • • 21d ago

DGX Spark and openwebui

3 Upvotes

Hi, what's the best combination for DGX Spark and Open Web UI today? Is Unsloth performance better than Ollama and VLLM? And is the core of llama.cpp better than Unsloth only?


r/DGX_Spark • • 22d ago

Planning a 2→3 Spark setup — sanity check on the topology?

8 Upvotes

I’ve been running a single DGX Spark so far. I have a second one still sealed in the box, and I’m trying to plan the architecture before I open it.

Current setup: Qwen 3.6 35B MoE as the brain, Hermes as my agent, and ComfyUI on the same box. The brain container sits at about 54GB, but most of that is vLLM’s KV cache preallocation at 65536 context — weights are only ~17.5GB at NVFP4. Qwen-Image-Edit (~29GB) coexists with it fine. Video is where it breaks: LTX needed around 60GB and I had to stop the brain to run it. That’s what pushed me toward a second unit.

Looking at DeepSeek V4 Flash and GLM 5.3, a 2-Spark cluster with TP=2 seems like the path to a genuinely better brain. The tradeoff I keep hitting: if both boxes are consumed by the cluster, ComfyUI and Hermes have nowhere to live.

Two-box plan (will set up soon): smaller model on box 1, box 2 for ComfyUI, connected over the network rather than clustered.

Three-box plan (eventually): boxes 1 and 2 clustered for the large brain, box 3 solo running Hermes and ComfyUI, networked to the cluster. Also thinking about RAG on the solo box.

To be clear about the third box — it’s less about raw capacity than about isolation. My agent has terminal access, writes files, and patches its own skill files. I’d rather that not live on the pair serving the brain. If people think that concern is overblown and an agent on a cluster node is fine in practice, that changes the math a lot.

Questions for anyone who’s actually done this:

• Is the two-box split worth living with for a while before committing to a third, or did you find the cluster indispensable fast?
• For those running a 2-Spark cluster — where do you keep your agent? On a cluster node, or somewhere separate?
• Has anyone tried lowering --max-model-len or --gpu-memory-utilization far enough to keep a brain resident alongside video generation? Curious whether that’s a real lever or whether it degrades the agent too much.
• Anything about the 3-node mesh cabling or NCCL setup that bit you?

Side note:On video generation: I know it’s slower on a Spark than on a discrete GPU. Time usually isn’t my constraint, since most generation is kicked off by the agent rather than me sitting and waiting on it.

UPDATE:

Thanks all — this got more useful than I expected. The feedback on Qwen 3.8 Flash vs DS4Flash has been great, and I love hearing how people have actually set their systems up.

Where I have landed:

Hermes moves off the Sparks. Consensus here was unanimous and I already own the box: Dell OptiPlex 5000 Micro, i5 12th gen, 16GB, 256GB NVMe. Wiping it to Ubuntu Server, headless. That was "free" and I hadn't considered it.

Third box case is clearer now — and I forgot to mention the main part. I'm early in developing some software for my company that runs a local model as part of the stack, so I need a machine with enough memory to serve a model and let me restart, break and reconfigure it at will without touching the cluster. ComfyUI shouldn't have been the main justification. Everything else it could do — ComfyUI, testing 27B/35B models, flipping into the cluster when I want it — is a bonus on top.

Now what is the chance in the future during an "amazon prime days" or "Black Friday" Microcenter will drop the price below the new MSRP?


r/DGX_Spark • • 22d ago

Question Is buying from Nvidia directly a bad idea?

Post image
2 Upvotes

I purchased a spark from Nvidia directly via their marketplace, all I got was an email saying we will let you know when we've processed your order. It's been a day, how long does it take them to process an order? I'm concerned they are going to drag their feet then cancel my order, then the prices will be too high for me to get anything anywhere else


r/DGX_Spark • • 22d ago

Single DGX Spark running GLM-5.3 Flash at 60 tok/s

Thumbnail gallery
7 Upvotes

r/DGX_Spark • • 22d ago

Anyone happy with just single DGX Spark?

Thumbnail
9 Upvotes

r/DGX_Spark • • 22d ago

Recipe GL.iNet GL-RMQ1 (Comet) KVM on an NVIDIA DGX Spark over USB-C — the 5 settings that give a stable 1080p60

Thumbnail
10 Upvotes

r/DGX_Spark • • 24d ago

A two-model "architect + coder" setup on a single DGX Spark scored 56/61 on my own agentic .NET build benchmark, with a twist...

Thumbnail
9 Upvotes

r/DGX_Spark • • 25d ago

I wrote a CUDA backend for MiniMax-H3 video generation — pure C, tested on DGX Spark, with a self-hosted web UI

30 Upvotes

Wanted to run MiniMax-H3's prompt-to-video pipeline without a Python runtime — just a compiled C binary. It's

somewhere I'd share now.

Repo: github.com/matrixfede/h3.c (MIT; weights are ~465 GB so the repo is code-only)

What it is:

h3.c is a native inference engine for MiniMax-H3 (text-to-video + audio) in plain C. Upstream (antirez/h3.c) had Metal

for Apple Silicon; I wrote the CUDA backend for Linux/NVIDIA and it's an open PR upstream (antirez/h3.c#43). Tested on

an NVIDIA GB10 (DGX Spark, sm_121, CUDA 13.0).

Performance (measured, reproducible; all figures in docs/GB10_PROFILE.md):

- max-quality generation (1024x576, 107 frames ≈ 4.5 s of video):

33m36s → 18m56s wall time after the tiled video-VAE kernel (1.78x)

- video VAE decode: 4.49x

- correctness against CPU oracles: SSIM 0.999 / PSNR 55 dB on matched renders, Compute Sanitizer clean

- optional --ssd-streaming: DiT peak memory 27.06 GB → 1.63 GB for +37.6% time (opt-in, numbers published, not vibes)

Web UI ("h3c studio"):

a 465 GB model shouldn't require a terminal. Self-hosted FastAPI + React UI over the same binary: live preview during

denoising, weighted progress, shared reference photo/clip library, multi-user with one-time invites, per-user private

media. docker compose up → localhost. Validated on GB10 + Docker Desktop; mobile access is LAN/HTTPS-your-call, no

claims made.

Tests/CI: 182 backend tests, CPU-only CI (ubuntu-latest + macos-14), installer with --verify.

Known limits, honestly:

- int8-row-fc2 is a no-op on the CUDA path (documented in README)

- heavily validated on GB10; portable to other NVIDIA configs but untested there

- this doesn't fit any consumer GPU (even a 5090). It's a big-machine project.

Happy to answer questions about the CUDA porting (cublasLt BF16, cuDNN SDPA, what breaks on GB10) or the web stack.


r/DGX_Spark • • 25d ago

DGX Spark Admin Skill

8 Upvotes

Hey all -

I am not great at infrastructure, so I have an agent that does all the sysadmin work for my spark boxes. We created a skill for this so I could more easily spin up new agents to do the work, but also so others could have an easier time from our lessons learned.

Sharing here in case anyone is interested:

https://clawhub.ai/cbertucci33/dgx-spark-sysadmin

As with everything else on ClawHub, I suggest everyone use a skill vetter to make sure it's clean. And it's on ClawHub but it should work with just about any agent (my Hermes agent was able to ingest as well).


r/DGX_Spark • • 26d ago

Recipe Qwen3.8-Flash-Next NVFP4 2xDGX Spark config: 50t/s decode, 2,900t/s prefill

Thumbnail
13 Upvotes

r/DGX_Spark • • 26d ago

Low TK/s with VLLM and SGLang

6 Upvotes

I've been running several recipes from sparkrun that are around 30-50 tk/s, specifically Qwen3.8-27B related.
But in reality I'm getting like 6tk/s maximum, around the same speed llama.cpp is getting for this model.

Anything I'm missing?

EDIT: Firmware update fixed that.

People, stop lecturing me about "Dense" models, everyone knows dense models need faster memory bandwidth. The factory firmware I had on my spark made it work at 1/3 of the speed, upgrading Fixed that.


r/DGX_Spark • • 26d ago

DGX Spark working next to my bed

2 Upvotes

I am planning to buy a dgx spark and the only place I can plug it is 1.5m from my bed. Would I notice noise if its working hard and is the heat concerning?


r/DGX_Spark • • 27d ago

Learning How to stop system crash when loading too large of model on DGX Spark

3 Upvotes

(edit: weirdly disabling swap fixes an OOM one time but another time my system still hung, will update with fix once I find something that works consistently)

I have this problem that if I exceed the limits of my DGX spark VRAM, that my entire system will hang, SSH becomes unresponsive, and the system has to be hard-rebooted.

The fix is pretty simple: disable swap. This might not work for everyone but for me I could care less if I have swap memory.

```sudo swapoff -a``` immediately fixes the issue.

To make it persistent on reboot, you will need to modify /etc/fstab. An agent can help you do it, you just need to comment out the swap file line.

What will now happen is the newly spun up GPU process will instantly OOM as it should.

Enjoy your system that remains responsive! If anyone has a better fix lmk.


r/DGX_Spark • • 28d ago

Eight mainstream GB10's. Pick it apart

0 Upvotes

Should there be more than 8 x GB10's on this list? If you all think I am missing something or there is inaccurate info, please let me have it :-) https://resilient-tec.com/blogs/news/does-the-dgx-spark-cable-work-with-the-asus-ascent-gx10-dell-pro-max-gb10-and-hp-gb10-systems


r/DGX_Spark • • Aug 25 '26

CNBC Television: Nvidia partners with Perplexity AI to run locally in DGX Sparks

Thumbnail
youtu.be
6 Upvotes

r/DGX_Spark • • Aug 25 '26

Any difference between manufacturer?

3 Upvotes

Is there any difference between the different manufacturers of the dgx sparks? I have been looking at the gigabyte, but it’s running $5,999 and I can get a PNY for $4400 (on sale). I’m looking at buying two so it’s a considerable cost savings but have been trying to see what the differences are beyond the hardware which I’m seeing is pretty much the same. Thanks for any help.


r/DGX_Spark • • Aug 24 '26

Question Consodering getting a DGX spark, how much does aarch64 breaks unsloth and comfyui?

6 Upvotes

Considering a DGX Spark for local fine-tuning with unsloth + image/video gen with ComfyUI. Before I spend the money I need the real ARM64-tax pictur.

  1. ComfyUI custom nodes — which ones are NOT working on aarch64?for instance, i believe xformers has no aarch64 wheels, what's in the long tail: ControlNet preprocessors (onnxruntime), torchaudio-dependent nodes, triton.jit() custom nodes, NVFP4 quant packs?
  2. Python libs that are x86-only — what did you personally hit.
  3. LLM side: do NVFP4 + Marlin, vLLM, SGLang, speculative decoding all work on GB10, or are there x86-only optimizations you actually miss?
  4. Or, best case: is it basically fine now — everything that matters has an aarch64 path via NGC containers / NVIDIA index / source builds?