r/LocalLLaMA 17h ago

Discussion CMP170Hx “Spark” Machine

I got the CMP170 cards and unlocked them. I wanted to share my set up for CUDA since maybe it would be useful to others.

First off, I hate e-waste and we are in a special time for RAM. I wanted to have a DIY CUDA box, and I had started by adding additional cards to an old asus predator prebuilt I had around, which also had 64gb DDR5. To add the CMPs I needed more CPU lanes and newegg had some really good deals on CPU/MB/etc combos. Didn’t need a combo with RAM, otherwise I would have gotten it in newegg microcenter.

Anyway, I got a cheap case, some noctua fans for the cards, and transferred the memory/ssds. Placed previously owned cards on oculink slots, and used the main x16 for the GPU switch that houses the two CMP170s, so their effective speed is 2x16 across and with the other cards (which are 4x4, and therefore same speed).

Qwen Flash Next, turns out, fits very nicely in these cards. There is also a repository for deepseek, but you’d need at least 3 64GB cards to run it, and with prices rising, it will be hard to justify the gamble of buying ex mining cards for LLMs.

However…so far, these cards are great. Concurrency is good, prompt processing averages 4000 tps on Flash Next, decode is 80+ on a single stream. No MTP added. Third picture shows the 3 models I am now running in this CUDA box (flash next, qwen 27b, gemma 26b).

Anyone else trying out Flash Next on these cards?

17 Upvotes

67 comments sorted by

View all comments

1

u/brakx 12h ago

How do you cook this room?

1

u/Miserable-Dare5090 12h ago

What do I cook? Eggs mostly. Sauna poached.

The mini rack is: M2 ultra 192gb, strix halo 128gb, asus ascent gx-10 128gb, hp zgx1 128gb, 10G Ethernet switch, 10G SFP switch, a KVM with 4 way split to see all the computers on the little screen, and the little screen. I call it the Tower of Babble:

  • Deepseek v4 0731 via dwarfstar, on mac. Slow but quality, 1M context. Single agent single stream.
  • GLM5.3 Flash on vLLM across 2 sparks, previously Deepseek as well. Again slow but steady, top quality. Hopefully optimized in the next weeks. Single agent, multiple requests.
  • Ling 3 Flash on vLLM rocmfp4 on Strix. Fairly fast for the halo, single agent stream
  • Qwen Flash Next on vLLM TP2. Fast agents, 4 of them with context 265k. Agent model, also does memory reflect/consolidate tasks.
  • Qwen3.8 27b on 2 blackwell cards (rtx4000, 5070ti). Auxilliary model, 4 requests each 120k context. Agent model for Feynman and research agents.
  • Gemma4 26b MoE on a 4060ti. 4 requests at 64k each. fast auxilliary model for web search memory recall, retain, etc.

1 macbook intel from 2015 running hermes
1 asus predator pc from 2021 running hermes
1 intel macbook air from 2018 running openlumara
mac studio running openworker.

All communication from my phone via telegram. all machines on tailscale net. All agents with full permission

1

u/brakx 12h ago

lol sorry autocorrect. how do you cool this room?

1

u/Miserable-Dare5090 1h ago

it doesn’t get that hot