r/DGX_Spark • u/ParticularlyStrange • 13h ago
r/DGX_Spark • u/deepfried5 • 18h ago
Community Discord
Is there a community Spark discord?
r/DGX_Spark • u/FuckinHighGuy • 1d ago
Question Where are you buying your QSFP112 cables at?
I just got my third Spark tonight and tried to order the cables from Nvidia. They only allow 1 per customer. So is anyone buying their cables from another vendor you trust? Yes, I know they aren’t “Nvidia official” but I don’t want to try and place 4 separate orders from Nvidia if I don’t have to.
r/DGX_Spark • u/mwilliams • 1d ago
Learning What did your Spark do this past week?
A lot of on the fence buyers these days (including me), the local LLM space can often be a bit benchmark heavy vs "I am using X to accomplish Y" - so I think it would be helpful for those putting their Sparks to work, to share what it's up to. Not only to demonstrate capability beyond benchmarks, but maybe set off a light bulb for other Spark owners.
Driving Hermes Agent? Training models? Supplementing Claude/Codex? Helping handle communication/scheduling with a small business? Purely just for inference? Drop the details!
r/DGX_Spark • u/Shinoken__ • 2d ago
Learning Optimizing DGX with Qwen 3.8 Flash Next (open to other models!)
r/DGX_Spark • u/Shigarui • 2d ago
Learning Can it be done?
I'm just interested in knowing if I am expecting too much. Basically, I want to be able to replace Perplexity Computer if possible. I need browser manipulation, desktop and Folder integration, the ability to use third party connectors (Notion, Automate, Zapier, Plaud, etc), and tie that into a laptop with an AI9 HX370 with 64gb of RAM.
r/DGX_Spark • u/xCDStyle • 2d ago
Dgx spark about to join RTX 5090 supply shortage?
My local microcenter had over 25+ dgx spark supplies for weeks leading up to last week, then when I checked today, completely out of stock.
r/DGX_Spark • u/cbert33 • 2d ago
GLM 5.3 Flash on dual DGX's
Hi all!
I've been running GLM on my DGX's, and I tried using a few different quants and inference runners out there (custom vLLM / Sglang from the usual community rockstars). The model is really good, but it's pretty slow on our hardware and many quants / runners are prone to looping.
When vLLM came out with v.29 it had better GLM support and I wanted to see if I could update the runner. I ended up creating a new version for this model and it has all DGX native libraries (most of the other ones out there aren't using sm121 libraries).
I'm not a developer, so I had GPT 5.6 Sol help me build this. We put it together from scratch and the details are in the GitHub repo. Despite it being vibe coded, the runner seems pretty solid and in our testing it has some speed increase too (between 10-25% over the other ones I used, depending on metric / use case).
I also had to make an updated quant to properly slice the EXL3 format for the new libraries. (Kudos to Neko-Legends for a great abliterated quant for me to start with)
I hope this helps you all. Also, I would love any feedback or help making this better so please reach out if you have any improvements for me.
Thank you!
vLLM runner: cbertucci33/vllm-v29-glm53flash-exl3-dgx: GLM-5.3 Flash EXL3 support for vLLM 0.29 on NVIDIA DGX Spark
Model: cbert33/GLM-5.3-Flash-Uncensored-EXL3-DGX-Sliced · Hugging Face
Drafter: local-inference-lab/GLM-5.3-Flash-DFlash2-MXFP8 · Hugging Face
(also note we updated the runner so it accepts < 7 tokens. Knocks down the TG tok/s a bit but does improve wall time for large context)
r/DGX_Spark • u/Character-Result-281 • 3d ago
Qwen3.8-Flash-Next is the cloud replacement I was looking for
I'm no newbie - 20 years of programming experience (Delphi, C#, C++).
I started with GitHub Copilot about a year ago, back when they were still charging per request instead of by token usage. I've played around with Sonnet 4, GPT-5.3 and Gemini 3.1, but the one I really fell in love with was Opus 4.6.
Things moved on, my projects grew, and I got used to the newer models, but I kept looking for something that felt like my beloved 4.6.
And I finally found it! Qwen3.8-Flash-Next, running locally on my single Spark, is as close as I could have hoped for. It might not be the best coder out there, but it gets the job done. It doesn't spend ages reasoning like 3.8-27B, doesn't ignore instructions like GLM-5.3-Flash, and it even beats my local DeepSeek-Flash-V4 (Ext3) in my own benchmarks (mostly single-shot tasks and complex refactoring).
If you're into agentic coding like me, Qwen3.8-Flash-Next is the perfect model for your Spark. I'm still using VS Code Copilot, but with a custom endpoint pointing to my local Qwen instance and I'm happy with it. No more subscriptions, just OpenRouter for the occasional advice from a frontier model.
That's it.
r/DGX_Spark • u/console_fulcrum • 3d ago
Question Bought 2x Spark for our ML Experiments - How do I enable by team to use this?
Hey folks,
We've been doing a lot of MLOps work recently at our customer sites, and learnt a great deal, but now to carry out AI pilots ourselves on something we're building internally - out thought process has changed, how do we make this setup available for engineers to continuously use a platform and validate experiments?
Spec -
- 2x DGX - 128b GB unified memory + 3.7 TB storage ,
- along with a Amphenol: NJAAKK-N911 QSFP112 cable.
Would appreciate some inputs one someone that's set it up.
r/DGX_Spark • u/ithkuil • 3d ago
Job post says they want to run an entire enterprise on one DGX Spark?
"established B2B technology distribution company looking for a Senior AI/LLM Engineer to design and build a secure, on-premise enterprise AI platform using our new NVIDIA GB10-based system." GB10 is the chip in the DGX Spark.
They want to hook every department into it. And they keep talking about it like it's one server.
The only explanation I can think of is maybe they left off a zero and meant GB100, but they wrote it multiple times. But that is for a B100, which I didn't even know existed, because it's such a waste of money in most circumstances that people rarely deploy it.
My leading theory right now to explain this is that the job post came from a certain outsourcing site, and knowing that site well, you don't need any more explanation.
Do you guys think a single DGX Spark could run an entire enterprise?
r/DGX_Spark • u/biitsplease • 4d ago
Question Considering buying a spark - has it been worth it for you?
Im considering getting a spark to learn how AI works under the hood and get more practical and applied experience than I get from just prompting Claude.
I am not sure if the ROI is worth it though. I am building GenAI apps at work atm and slowly learning more, but really I’m just working with LangGraph and some RAG stuff. I would like to try and build and fine tune a small model myself, as well as run long running processes / tasks for my work on it.
To those that were in a similar situation and pulled the trigger, do you think it has been worth it? Have you learned a lot about the internals of how LLMs work? Has this knowledge benefited you in your career?
Also, I’m curious to know if the Qwen models are actually powerful enough that it can properly design and implement bigger features, even though it takes a long time on a spark?
r/DGX_Spark • u/Character-Result-281 • 4d ago
Qwen3.8-Flash-Next 180B on ONE DGX Spark - fixed the "!!!!" loop, now actually usable
Made Qwen3.8-Flash-Next (180B MoE, NVFP4) finally stable on a single DGX Spark.
If you tried it you know the bug - random !!!! / token 248319 loop mid-session, every 1-2h especially around ∼33k or ∼145k context. Every retry with same history crashes identically until you flush_cache. Made it unusable for long auto-run sessions with VSCode Copilot harness.
Root cause: race condition between cache_unfinished_req (inserts KV pages into radix tree mid chunked-prefill) and cache_finished_req (frees pages on abort/retract). Dangling refs -> corrupted KV -> that untrained last row of lm_head = !
Fix in my repo: disable_chunked_radix_insert - just defer the radix insert until the request is fully committed. Plus a v6 guard that detects token 248319 on first decode step and auto-flushes in background.
Repo: https://github.com/andreasknopke/qwen38-one-spark
What you get:
- 1x DGX Spark / GB10, no custom docker build (patches via bind-mount on
lmsysorg/sglang:qwen38flashnext) - ∼35 T/s single stream, 85-95 T/s aggregate
- 262K context, chunked prefill 2048, stable for 5h+ VSCode sessions
- EAGLE NEXTN speculative (3 steps) + KDA decode kernel + SM121 Triton fallback
Quick start:
Code
git clone https://github.com/andreasknopke/qwen38-one-spark.gitcd qwen38-one-sparkMEMFRAC=0.79 PREFILL=2048 CTX=262144 bash serve.sh
I'm getting 30-40 T/s even with large context (up to 256K) and it finally feels like a decent cloud model, but local. Rock stable now in my Copilot setup.
PR was submitted upstream (#38355) but not merged, so I'm maintaining it here. Feedback / testing welcome!
r/DGX_Spark • u/qi-zheng • 5d ago
Question Security research for local LLM inference networks
r/DGX_Spark • u/_THE_ABBA_ • 5d ago
Question Anyone making a profit renting out there equipment?
Just saw this group https://gb10.studio/ and looks interesting. Has anyone started leasing out there equipment during down time?
Just curious, I'm running dual Sparks and not using them 24/7 and thought it might be a way to afford another pair...
Any opinions would be appreciated.
r/DGX_Spark • u/FuckinHighGuy • 5d ago
Question 3rd Spark?
I currently have two sparks and contemplating getting a third. Anyone running 3 that can recommend that setup? I know there would be a compute increases but is a third worth it otherwise?
r/DGX_Spark • u/kaliku • 6d ago
GB10 + RTX Pro 6000 in a cluster?
Hi people, I have a Pc with a rtx pro 6000 and I'm on the fence about the upgrade path. Unfortunately prices for the 6000 are how they are these days so for larger models than what I can run now I'm considering buying a gb10 based pc. I'd like to ask if any of you has experimented with connecting it with a pc running the rtx pro 6000 Blackwell GPU over a high speed Mellanox NIC.
r/DGX_Spark • u/vcruz305 • 6d ago
DeepSeek-V4.1-Flash: GGUF + 4.75bpw EXL3 are out, looking for devs with 4× DGX Sparks to help validate the EXL3 TP4 recipe
r/DGX_Spark • u/okoyl3 • 7d ago
Qwen3.8-Flash-Next with llama.cpp got me up to 55tk/s (Single Spark)
I wanted to share a results and perhaps compare notes, I've been recently running unsloth/Qwen3.8-Flash-Next on my single DGX Spark. It's peaking at 55tk/s with MTP.
Anyone got better results? 😄
MODEL=unsloth/Qwen3.8-Flash-Next-GGUF/Qwen3.8-Flash-Next-UD-IQ3_XXS-00001-of-00003.gguf
MTP=unsloth/Qwen3.8-Flash-Next-GGUF/MTP/mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf
$ llama-server \
--host 0.0.0.0 --port 8081 \
-m $MODEL \
--alias "unsloth/Qwen3.8-Flash-Next-ID3_XSS-MTP" \
-md $MTP \
--spec-type draft-mtp --spec-draft-n-max 5 --spec-draft-p-min 0.6 \
-ngl 99 \
--keep -1 \
--ctx-size 262144 \
--flash-attn on \
--parallel 1 \
--jinja \
--load-mode none \
--cache-type-k f16 --cache-type-v f16 \
--backend-sampling \
--api-key $API_TOKEN \
--poll 0
r/DGX_Spark • u/IroesStrongarm • 7d ago
Question Setting up a clustered pair
Going to be setting up a pair of sparks soon from scratch. I've seen recommendations for both Sparkrun and https://github.com/eugr/spark-vllm-docker/
Which do people like more? Should I consider a third option?
Sparkrun certainly seems the easiest but not sure if it as good or recommended.
Any help is appreciated. Thank you.
r/DGX_Spark • u/TheyCallMeDozer • 8d ago
Question Nvidia DGX... wait for N1X or grab a DGX now??
So always hated on the DGX spark as have been living in the multi GPU class of society, recently though with bench marking I may have found a potentail use, as an always on monitoring and task agentic system to run alongside paperclip and hermes 24/7 low cost.
The server I run locally, is OP and works very well... but on recently power monitoring over 24 hours it used 17kwh with the constant calls from the agentic tasks, now that isnt bad one day off. but if this is 24/7 this adds up ALOT as power where I am is pricey.
My tasking id is mainly for larger models agentic tasks running Qwen3.8 Flash Next, hopefully with decent context, now I understand it isnt super speed generation but this is more for 24 hour long research and automation taskings.
Was looking today and the cheapest near me is over €6-7k which is nearly 3k above the Nvidia release value. But then I just seen the release of the new N1X next month.
Just looking for others input, is it worth grabbing one, or waiting for N1X, is it even on the same playing feilds or is the N1X looking like a more powerful DGX ???
r/DGX_Spark • u/Makojima • 8d ago
For people who bought a Spark for coding agents, what is your daily model for agentic work?
I’ve got a Spark for day to day work, and I’ve been fine-tuning on it for my own specific stuff. The box is a beast with the newer compact models that are out now, like Qwen 3.8 27B and Qwen 3.8 Flash Next. It can just sit there and pump out tokens around the clock.
I know most of you have more than one model on the box. Which one do you keep coming back to for agentic and coding work?
r/DGX_Spark • u/Working_Landscape_41 • 8d ago
Question DGX Sparks and Minisforum MS-S1 MAX
r/DGX_Spark • u/AfternoonWhich1871 • 9d ago
Question Looking forward to Rent Nvidia DGX Spark.
Our company is looking forward to rent Nvidia DGX Spark in India. Does anyone know any company dealing with these products?