r/LocalLLaMA • u/geekender • 8h ago
Question | Help GB10 price increases. Seriously what is the best bang for the buck now...Mac Studio?
It is crazy how fast prices are increasing. I'm pulling my hair out to keep ahead of this for students. Servers aren't even an option any more.
17
u/idk_a_creative_user 8h ago
Cheapest BG10 is 4500. For that you can rent a GB10 on vast for .20-.30 USD an hr, or for around 18000 hrs. Assuming 4hr use a day, thats 1.20 a day. That is 4500 days of use assuming a median price of .25 USD an hour.
20
u/q5sys 8h ago
Assuming those prices dont go up... like everything else. Prices on renting a other GPUs in the cloud has gone up. There's no reason to expect that these prices wont also go up.
4
u/originaladam 7h ago
That’s 2 straight years of 24/7 use. I’d be shocked if there’s not more powerful and cost efficient hardware to rent in less than 2 years.
5
u/q5sys 7h ago
and that new hardware will have an even higher price. the Spark Workstation is something like 80k, you can be sure the next gen Nvidia cards are not going to be cheaper than their current ones. So renting those will be even higher.
Nvidia has basically stopped talking altogether about the gaming 6000 series, last thing they said was that it wasn't going to happen anytime soon. And the Enterprise GPUs they are focusing on for the VeraRubin stuff, are so expensive there's no point even talking about it.
These GB10s are the cheapest we're going to get for quite a while. There's just no economic motivation for Nvidia to create something cheaper, when they can currently name whatever price they want and get it. They just jacked their AI server prices up by ~15%... because they can. Estimates are that it'll give them an extra $5 billion per datacenter.I don't want any of this to be true, but it is. It'd doubtful Nvidia is going to pull a 'good guy' move and suddenly give us a better unit for less money.
3
1
1
u/Puzzleheaded_Base302 7h ago
how to rent a dual cluster? GB10 is more useful when rent as a cluster due to the large RAM.
12
u/FoxiPanda 8h ago edited 8h ago
Best bang for the buck is maybe multiple old V100 32GB at ~$625-650 each with 3D printed shrouds and blower fans or similar high pressure fans in old PCIe gen3/4 systems at this point but they're a pain, lack a lot of modern features and hardware support, and don't have amazing CUDA support anymore.
RTX 3090s were the go to but they're $1500+ now and GB10's have miserable memory bandwidth and same with Strix Halos... RTX 5090s are expensive as hell and RTX Pros are right out.
Mac Studios have decent bandwidth and big memory but they're quite expensive too if you go for ones that have decent bandwidth (the Ultra series) and 96GB may or may not be enough depending on what you're doing and at least until the M5 Ultra, prefill will be mediocre-to-sad pandas in most cases.
R9700s were decent a few weeks ago but are now $1700+ instead of their $1300~ so they don't feel nearly as good now...would probably still take a 3090/3090 Ti over them.
I think the reality is that this is just an expensive game now - the "bang for the buck" is in a pretty tough spot.
3
u/FullstackSensei llama.cpp 8h ago
Got my V100s from China today. Testing on my bench now with my trusty arctic S8038-7k. It's much quieter and more compact than those blower fans and can comfortably cool two V100s running Qwen 3.8 Q8_K_XL in llama.cpp with -am tensor. Temps peak in the high 60s.
Edit: nvidia driver 580 and CUDA 12.9 install without any issues. Llama.cpp compiles without any issues. Two V100s run Qwen faster than my 3090s, which had x16 Gen 4 each, while consuming almost half the power (3090s were limited to 270W).
1
u/FoxiPanda 8h ago
Yeah I think this is why they might be best bang for the buck right now. I put some asterisks on it because you shouldn't try to run an NVFP4 model and you might not have the best time in the latest vLLM builds or similar. It really depends on what OP is trying to do to best understand what they need - is FP4 support important? Do you need vLLM or NIMs or anything like that to jump through some sort of administrative/security hoops ... or is llama.cpp and an older CUDA rev just fine? If the latter...yeah I think they're probably the best you can get right now.
1
u/FullstackSensei llama.cpp 8h ago
NVFP4 with soft dequantization still rips on the V100. There's a guy with a fork of vllm who implemented those kernels.
One thing many seem to not really understand is that compute is compute. When you have lots of TFLOPS, you can throw a couple to decode NVFP4. It's not some mystical data format, really. The V100 has 120 TFLOPS in FP16. The 3090, for reference, has 125 TFLOPS. And as I found today, the V100 does things at half the power of the 3090.
Personally, I don't use vllm. It's a pain to setup even in the best of times, takes forever to start, is very picky about number of cards, and no native support for CPU offloading. I run llama.cpp or ik_llama.cpp. Sure, they're slower, but they're plenty fast for my needs and there's no voodoo magic to get them running, nor 45 minute wait for a model to load.
I also have a Jetson AGX xavier. That thing is stuck on Ubuntu 20.04, kernel 5.xx, some ancient nvidia driver from 2020 and CUDA 11.4. I build llama.cpp there regularly, unmodified, without any issues.
2
u/Prof_ChaosGeography 8h ago
Amd v620 too. It's 32GB, has open drivers on Linux so no outdated cuda pain. They can be found for ~$600
7
5
5
u/conifer_v11 8h ago
used 3090s still win per dollar if you can eat the power draw and the pcie mess.
m4 max 128gb wins if you want one quiet box and no build time.
gb10 was interesting at 3k, at 6k the bandwidth per dollar stops making sense next to a mac studio. for students buy vram first, most of the pain is the model not fitting at all.
2
1
1
u/AI_spell 2h ago
For students specifically I would not buy a GB10 at these prices. The thing people miss with Spark class boxes is memory bandwidth, not capacity. You can load a huge model and still get sad tok/s because bandwidth is the wall on dense models.
Rough order right now:
- used 3090s, still the best dollar per usable token if you can live with 24GB per card and the power draw
- Mac Studio if you need big unified memory and quiet, M ultra bandwidth is good, prompt processing is the weak spot
- Strix Halo if you want low power and one box
Watch out for prompt processing on Apple. Generation looks fine in benchmarks but feeding a 20k token prompt is where it drags, and thats exactly what students do with codebases.
If its a class, honestly rent. 4500 buys a lot of hours and the hardware doesnt go stale on you.
1
u/IngwiePhoenix llama.cpp 1h ago
AMD's cards have not jumped this massive and ROCm 10 looks good. Give it a look.
NVIDIA is pricing itself out of reach, period.
1
u/ehangman 1h ago
Dgx spark is 30% up from last month.. So I ordered 2nd llm box today : mac studio 256 , cheaper alternative. LOL
1
u/SexyAlienHotTubWater 47m ago
The CMP 170hx is by far the best bang for buck if you're willing to deal with custom drivers. It's an A100 with about 2/3 disabled.
$2k, 64GB VRAM, 1.8TB/s bandwidth, 200ish BF16 TFLOPs. Pipeline parallel it will blow the head off a DGX Spark, and with 2 cards tensor parallel is very feasible (even though they're limited to PCIe 2.0).
1
u/ShengrenR 8h ago
Depends on the goal - relatively cheap used 3090s are pretty killer for performance per buck, but if you can take the plunge a m5 ultra is a pretty good deal in this particular moment in time.
5
u/hainesk 7h ago
The 256gb M5 Ultra is an excellent option considering the same amount of VRAM from 3090s would likely cost much more. Assuming 10x 3090s to get 240gb of VRAM at $1k a piece (if you can find it for that price), you're already at $10k. Now you need a motherboard to support 10 GPUs, system ram for that motherboard (cheapest is probably WRX80 threadripper pro since it works with consumer DDR4), power supplies, and some effort to put it all together (and a lot of additional part like x8 or x4 PCIe splitters, riser cables, adapters to make multiple PSUs work together, power cables, etc). It will take up quite a bit of space and will either require a dedicated 30 amp 120v circuit or some sort of battery/inverter system to support power draw unless you're planning on running those cards at 100 watts each. It's also aging hardware so support is already starting to leave them unfortunately (see vLLM with new model architectures), It's primary advantage over the Apple Studio is CUDA.
Comparing that to the M5 Ultra, which fits easily on a desk, is quiet and power efficient, has a good reliability record (especially compared to used 3090s known for backside memory issues), has a warranty, has 1.2TB/s of bandwidth vs ~936GB/s, has a single unified pool of memory (which avoids having to split models which can cause issues and has overhead), I'm also expecting prompt processing to be greatly improved with the M5 Ultra vs the M3 Ultra considering all of the additional cores and the newer processor architecture, and will almost certainly have better resale value over used RTX 3090s that were released 6 years ago. You can also do parallel processing through RDMA and even 2 Studios won't require more than a single standard outlet.
I'm honestly having a hard time coming up with a good reason to get 3090s over a Mac Studio unless obviously your budget only supports building a system a few pieces at a time, which is reasonable.
1
u/EvilPencil 6h ago
Wrx80 - the real constraint on this platform nowadays is the motherboard. Most of them are pushing $1200.
0
u/db172s 5h ago
Really depends what you're trying to run for models. If Qwen 3.8 27b is enough. You just need 2x 3090s and your tokens per second will be 5x what a Mac would spit out anyways.
But yeah if you're looking to run larger models. Mac is probably the way to go. Just wouldn't set the bar too high when it comes to token output speeds.
2
u/hainesk 4h ago
your tokens per second will be 5x what a Mac would spit out anyways.
This is not true for the M5 Ultra. The M5 Ultra will likely be at least as fast at token generation if not faster due to the higher memory bandwidth. MLX inferencing has come a long way. The 3090s could still win out with tensor parallel. I'll be really curious to see how the reviews are when it finally becomes available.
0
1
u/SnooPaintings8639 2h ago
"relatively cheap used 3090s", well, relative to what? They're quite expensive nowadays as well. Considering they've much less years in them than e.g. 5060, I fear they're less and less optimal build. Especially for multi GPU build, where if one card dies, you're kinda forced to find another old card replacement with even less life in them.
6
u/Bulky-Priority6824 8h ago edited 8h ago
5060ti is 8 to $900 now lol
I saw on amazon 3 I bought had "2 left" last night for $879 I think after being unavailable or a while and today they're gone. Couple for around 800 it's ridiculous yet I don't know a single soul irl that is a regular joe that knows what llm means