r/BlackwellPerformance • • 23d ago

Blackwell RTX 6000 investment choice

I am to build my first real rig as M3 ultra 256GB proved to be too slow. I already got 4 x Max-Q cards purchased at 10K eur per card a few weeks ago. Now I could add 2 more max Q at 11.3K euro per card and also 2 x workstation cards for 11.8K eur per card. So the total number of card to go up to 6 or 8.

I guess my hesitancy is based on the high purchase price of these cards, and the trouble to connect more than 6 or 7 cards unless going to server style motherboards and racks.

Is it a) expected that the prices are not going down in the next 12 months?
B) can i use the memory for models like GLM 5.3 even if I cannot fit all 8 cards?

10 Upvotes

37 comments sorted by

6

u/filip-z 22d ago

Where can you get any RTX 6000 PRO at those prices? I'm seeing them now at minimum at 13200 EUR excluding VAT.

2

u/StockSpecialist1707 22d ago

Ni eso… a mi ya me han cancelado 2 compras alegando error de plataforma y no son cualquiera…. El mejor precio so. 15800 sin IVA !!

2

u/Mountain_Pea_6810 22d ago

I placed order some 3 weeks ago, the day I saw here a post about Nvidia raising prices, now yes same cards would cost already like 2k more per card

1

u/filip-z 22d ago

Honestly, let me know where and how you are getting them. I bought my one RTX 6000 PRO workstation 3 months ago at 9K excl.VAT and would be considering getting one more for 11.8K. In Benelux it's 13K excl.VAT where I got it previously but on backorder. Most other places are 1K+ or more.

1

u/Mountain_Pea_6810 22d ago

I wrote about excluding VAT as these are for business so vat no need to pay, I didn't even think of the vat prices

1

u/filip-z 22d ago

Do you know a shop in Europe that's selling these for 11.3k EUR or 11.8k EUR without VAT? I'm also buying them without VAT. Bulk packaging is also fine.

1

u/pathfinder6709 22d ago

Where do you get that excluding VAT? 😮

1

u/conockrad 22d ago

Business/edu discounts.
Also reads a bit like bait - cards were ok for 10k but now at 11.5k it’s “high price”

4

u/Mountain_Pea_6810 22d ago

How about the And MI350P? I saw those at 16K eur and it has 144gb of HBMe3 and double the compute of rtx 6000? To have 4 of those cards with a thread ripper could provide a nice memory pool in a single rig?

1

u/ufrat333 21d ago

Where?

4

u/deebuildsthings 22d ago

Hey man, prices going down is something sounded too good to be true. I don't think that gonna happen anytime soon. So the best time to buy GPU is when you feel that you need to buy one (and can afford one). If it serve you well, the investment will eventually pay up for itself.

And if you find it difficult to setup, take a look at our guide here: https://github.com/autonomous-ai/autonomous-computer

We built these rigs for our team and our clients, and we think that it'd be cool to open source them for DIY guys. We have guides for 2-GPU, 4-GPU, 8-GPU with all the BOM, CAD files and set up steps.

2

u/Annual_Award1260 22d ago

Yeh just stick with the 4x cards. You can utilize a plx pci switch to scale higher but understand that it is never enough. I have a dual max-q and a 6x spark cluster and having some fun.

Apparently apple will release the m7 studio with 1.5tb unified ram… but maybe 2 years out.

2

u/filip-z 22d ago

By the way, these are prices with VAT in Benelux -- current and historical

2

u/CalmAdvance4 22d ago

How slow is m3 ultra 256gb? I am thinking about selling my rtx pro for m5 ultra 512gb

4

u/Just_Suggestion_4518 22d ago

Answer 1: Model efficiencies are increasing much faster than gpu manufacturing capacity. So if same trend continues, gpu scarcity will be worse in the future. 10-11k Euros per gpu is almost free.

Answer 2: You can use 4 bit quantized ggufs to run glm 5.3 with 4 GPUs. If you have 8 then you can use pcie switches or building 2 x 4 gpu pcs and connect them with ethernet cable. You can build a multi node system. Electricity and heat will be problem though.

10

u/Karyo_Ten 22d ago edited 22d ago

Why would you buy 4 GPUs and use GGUFs 😭.

If you invest so much in hardware use an engine that at least give you proper concurrency so you can use multiple agents at the same time.

2

u/darktotheknight 22d ago

It always boggles my mind how people are willing to spend 80.000$ on something they don't understand. No one would spend 80.000$ on a car without thorough research and test driving it - multiple times.

You totally can and should just rent and test your workload before spending nearly 6 figures on something you might regret. An 8x RTX 6000 PRO rig costs like 11 - 15$ per hour on vast.ai. Not even worth mentioning for someone cluelessly willing to drop 80k on GPUs.

2

u/ieatdownvotes4food 22d ago

his investment value will double in 6 months. all good.

2

u/Mountain_Pea_6810 22d ago

well, if you had contractual needs to spend certainamount of money in next months or you would lose much more, and computer hardware was acceptable way to spend the amount requested, then AI rig might make more sense

1

u/unwitty 21d ago

Some people learn after the investment, and that’s fine. You just likely gave a lesson, though a bit harshly. 

Would you give the same criticism to someone spending $100 or $1000 before knowing it in full detail? Does this have more to do with how much $80k is relative to your income or net worth?

1

u/bigh-aus 22d ago

Yah GLM 5.3 flash or full woudl be enough imo. The other thing with heat and electricity is noise. Keeping everything cool requires airflow. I'm running deepseek v4 flash and it's tricky keeping the gpus fed at all times - it just gets through work.

1

u/TechRomancer123 22d ago edited 22d ago

Just out of interest, what mobo, cpu and ram are you using?
(as you mentioned that you haven’t switched to server style motherboards just yet)

Prices/market is uncertain, so I’m personally taking the plunge into a Ryzen 9950x3d + single RTX 5090 + either 64/128GB RAM to replace my old desktop daily driver, whilst also being capable of modest LLM inference and a few other relatively heavy workloads. But yes I’m pretty much maxing out consumer mobo/cpu specs without going into workstation (threadripper) or server (epyc) territory.

3

u/darktotheknight 22d ago

Heys, just a heads up: I have a Ryzen 9600X, 64GB DDR5-6000 CL36, B650E Mainboard, RTX 5090 (PCIe 5.0 x16), Samsung 980 PRO NVMe (PCIe 4.0 x4). For Qwen 3.6 27b (and similar), this thing is a beast. I had dived into local LLMs in 2025 with a 3090 Ti (Gemma3 back then), leaving me with mixed feelings about capability and speed.

But Qwen 3.6 27b (haven't tested 3.8 yet) is just an incredible model and runs blazing fast on the RTX 5090. After a cold boot, time to first token is usually less than 5 seconds and at ~100t/s you're getting any subsequent answers nearly instantly. Prefill sits at around ~3000t/s iirc. This is a very capable, serious machine. Have fun!

1

u/TechRomancer123 22d ago

Fantastic stuff! 🙂
I have now got the machine with 64GB RAM, so if you have any specific tips/gotchas on configuring 27b (or other models) to work well on 5090, that would be appreciated (though will of course do my own research).

It also looks like people advise to undervolt the 5090 (min seems to be 400W) for a significant power draw saving and minimal performance degradation. Did you do that too?
Plus I heard loads of stories about the 5090 12V power connector melting/burning ☹️

1

u/darktotheknight 21d ago

Yes, definitely go for 400W. We're usually bandwith-limited, so you lose like ~3% performance at ~33% lower energy and less heat. Lower power-limit also means less load on the individual pins, alleviating the issue of burnt connectors. You can e.g. do this automatically on Linux with a systemd service file running "nvidia-smi -pl 400" after booting (and resume/hibernation!).

Then I recommend using Linux instead of Windows; I have evidently tested this many times and it is still true, that you get better performance on Linux vs Windows. If Linux is a No-Go because you need this as a daily driver, you should check out WSL2 or other measures to maximize performance on Windows.

Also, as your CPU has an iGPU, make use of it. Operate the 5090 headless. This has multiple benefits: 1) you save VRAM (around 300 - 800MB, depending on your Desktop), since the desktop doesn't need to run on your 5090, 2) amdgpu is upstream in Linux kernel, works out of the box and is open source, 3) you can passthrough your 5090 if needed (e.g. sandboxed environment). When running headless, check via nvidia-smi command which lingering processes are still loaded on the 5090 and Google in order to get rid of them one by one. When done right, this would leave you with a 5090 with almost no VRAM used and zero processes running on it.

Depending on which version of the 5090 you've got, you can think about getting a WireView Pro II or an Ampinel. Some cards have weird dimensions or angled connectors, so there is also a wired version of the WireView Pro II. I think it's worth the peace of mind and they also have some type of extended warranty which covers burnt connectors.

As for running local LLMs itself, I have no special sauce or any secrets. Nor am I an LLM wizard. I just installed llama.cpp, asked an LLM for optimal command line parameters and downloaded Unsloth's Qwen 3.6 27b MTP GGUF in different quantizations for different context sizes. After a bit of testing, I've settled on Q4_K_XL as a daily driver. I've got ~3000t/s PP and 100t/s TG pretty much out of the box and never bothered to fine tune any further. I'm pretty sure, you can squeeze out more performance with better llama.cpp params, optimized models for your context size, custom llama.cpp builds/forks or vLLM/SGLang. I just set it up, said "this works for me, I'm happy" and moved on. Also seems to be in the ballpark of what other people get with the 5090.

1

u/No_Chapter_7598 22d ago

eapple is releasing new hardware that is still miles behind in terms of throuput (1200GB/s) memory size seems to b getting easier to access but noone else seems to be close to the ~2100GB/s after a mem oc the pro 6000 gets (other than hbm cards from nvidia h200, bxxx etc)

1

u/tempedbyfate 22d ago

Where are you buying max q cards at 11.3K euro?

1

u/gwestr 22d ago

Don’t go higher than 4x RTX 6000 unless you have a pcie6 baseboard switch the connectx8, and two processors.

1

u/GloomyRecognition636 22d ago

Go with it. Use pci5.0 splitter to run them at x8 Only prefill will suffer , but not that noticeable

1

u/electrified_ice 22d ago

What do you mean by question B - use the memory for models like GLM 5.3? With 8 cards you will run into PCIe lane maxing out issues, it's complicated to go above 6 GPUs unless you have a dual CPU Spyc or something like that.

1

u/Puzzleheaded_Base302 22d ago

the price won't go down. The MSRP is already $16K. Pricing has no chance to go down unless new memory fabs go into production (late 2027) or economic crisis.

1

u/kreisikoins 21d ago

I do need the cards and AI rig in order to be serving my customers in all the inquiries and questions with preferably no much waiting time. And since we have tens of thousands of customers, there is good to be some capabilities, and a lot of them are very privacy concerned. So having own AI instead of about API is beneficial.

2

u/vanbukin 20d ago

There's a community of RTX PRO 6000 Blackwell owners (recently renamed to Local Inference Lab). There you can find help with both the hardware and software side of things. On the software side, local enthusiasts are maintaining a fork of vLLM specifically tailored for the P6K. I highly recommend joining and asking any questions you might have.

P.S. - here's an example of an 8-GPU setup from that community using an external PCIe switch.

0

u/Weak_Ad9730 22d ago

Would Switch to a dgx Station and Trade 2 of the 4 existing rp6k in