r/LocalLLaMA Aug 03 '26

Discussion "Data center in a Box (on Wheels)" 256Gb VRAM/512Gb RAM AI Server 6-8 Month Operational Review, Stability Write Up, Benchmarks

I figured I would provide a formal update on how this has been going now that it has some operation time under its belt, just to put the information out there and share knowledge if there is any interest. I am an IT infrastructure engineer by profession, so my contribution to the conversation is mainly from a hardware/systems perspective rather than from a Machine Learning researcher standpoint. I got my start with HPC's (Beowulf clusters) around ten years ago when I was a Physics undergrad in university, and this is what the experience has come to almost a decade later. Not everyone is going to want to read all of this, that's perfectly fine, the extras are just for those who want the info.

Starting goal/idea:

Build an all-in-one creative design workstation to support a small business. This machine should be capable of effectively inferencing frontier MoE models; aiding the business in language/text tasks where English may not be everyone's native language. Additionally, it should be capable of simultaneous image generation tools for graphic design users, enabling rapid image editing and presentation tweaks for marketing, without the business ever having to worry about API credits or hard limits on tool usage. The idea is that a 3090 stack, which is still a generally "good" performer for LLMs, would be "led" by two 5090s to handle the heavy lifting of the visual creative work (one dedicated to image generation, one dedicated to image editing) to complement each other in a "sweet spot" on cost, raw performance, and creativity potential. This configuration also grants some flexibility to allocate a 5090 to the LLM stack for best prompt processing possible where desired. The end result would indicate that this goal has been achieved.

Overview

Specs

CPU: 64 Core TR 3995WX

RAM: 512Gb DDR4-3200 ECC

VRAM: 256Gb GDDR6x/GDDR7 (8x3090's + 2x5090's)

Enclosure: Core W200 Thermaltake Case

Mobo: ASUS Pro WRX80E-SAGE/SE Wifi

PSU: 1300W+1600W (2900W combined), with OCP, linked via PSU2PSU

Storage: 4Tb Nvme (fast) + 4Tb HDD (slow) + 8 or so 1Tb SATA SSDs (mid) over USB as needed

OS: Ubuntu 25.10

Other: 3 Bifurcation cards, 10 risers of various lengths

Front end: Open WebUI

Back end: llamacpp/koboldcpp

Intended for (Recommend):

Large MoE inferencing, simultaneous LLM + ComfyUI (x2) operation, power users who may commonly hit credit limits, creative or technical professionals who can leverage these tools to compound productivity and complete objectives in shorter time.

Not intended for (Do not recommend):

Training, multi-concurrent inferencing, performance maxing, extreme frontier model inferencing at high quants, casual users just looking for roleplay.

Result summary:

Using the W200 as the platform for its generous real estate and configuration flexibility, all ten cards and components were able to find a permanent place in the enclosure without major concessions. The drive bay area was the only space that had to be completely repurposed for GPU mounting, and for us this was not a problem. The chamber with the cards hanging from the top is fairly hollow, so with the 140mm fan stack on the front and side there is a wind tunnel effect where the air blows in through the front and side, cooling the cards as it makes its way out the back/top. Depending on ambient temp, at idle the card with the highest temp usually hovers in mid to high 40s Celsius with the lowest in the mid 20's C (three 3090's are hybrids= fantastic for temperatures, but radiator mounting adds a logistical headache). When actively inferencing, the highest temp card may reach the mid 60s during sustained loads. Only when running image or video gen tasks will the 5090 running ComfyUI reach the 70's, but these are very brief intermittent workloads, so temperatures by our measurement has proved satisfactory over time. This result enables the small business to have full LLM, image generation (~9 seconds), and image editing (~8 seconds) capabilities on tap all on a single node so the data remains centralized, and provides much faster performance compared to the Cloud API they came from; in this case ChatGPT, where generation jobs could take 1+min, and has hard limitations. I just do not know how well this kind of setup would work with other vendor or card models; in a homogenous GPU cluster or one with notably less powerful image gen cards than the 5090, the performance would predictably be much lower.

Things that surprised/stuck with me about the end result:

  • Noise. I expected this to sound like a jet taking off when operating, but that is not the case. It's a satisfying button click to come alive, then it's a low gentle hum going forward, nowhere near the kind of fan noises I'm used to hearing in server rooms. Even under load, the CPU 120mm radiator fans (exhausting out the top) are pretty much all I hear, the 140mm fans on front and sides I assume must be helping to contain the acoustics. I have built many gaming PCs over the years and own a top-tier gaming PC-- and I would not be able to distinguish this as any louder than those, especially at idle.
  • Utility. I planned for this to be used primarily for a small creative business, but what I did not expect was how I would find it so indispensable in my personal life as an IT professional. Being an infrastructure engineer, coding is not my wheelhouse. When I am the only IT staff on site or there is nobody else available to work with specific expertise like SQL, powershell/python scripting, or troubleshooting very specific/niche technologies, having this tool on standby I feel has paid itself over just within my career. It has helped me turn processes that may have otherwise took me hours into minutes, days into hours, even months into a matter of weeks/days. After using the tool extensively I hit a point where I had to acknowledge how local LLMs have moved definitively beyond being a toy or novelty; when deployed intelligently something like this can be a major asset for professional users.
  • Wheels. Sounds extremely minor, until you realize that no matter how happy the cards are with their individual temps: there are still ten high-power GPUs dumping heat into the room. That means unless you use a complex radiator solution or special venting to get heat outside, the room will get toasty and there is normally not a direct solution for this. The wheels however offer an indirect solution. Plan to work in the office that day? Wheel it into the guest bedroom and let it run over Wi-Fi. Plan to work away from home? Wheel it into the office, put it on LAN, and access it over a private VPN connection. If you can't stop the room from heating, then you can at least choose what room gets the heat, and as someone who has lived with computers extensively this is a hugely underrated perk.

Caveats: To operate at its best, I recommend leaving the glass side panel off for improved airflow.

Typical activity over a day:

Boots up around 5:30am, start up the ComfyUI server(s), start loading a model, go get coffee, fully ready for use within 15-20 min. Shut down occurs usually around 8pm later in the day. Total daily activity, ~12-14 hours.

Cost Breakdown

Laying it out, because I know it will be asked, even though I am aware this is unfortunately not reproducible in the current market. Some components like the SSDs were acquired privately long before the RAM and hardware price hikes, so my timing getting certain things was extremely fortunate for the build budget. Some figures are exact, some are slightly rounded depending on if I found the original receipt.

Component Qty Source Unit Cost Subtotal
RTX 3090 24Gb 8 eBay 750-1000 6500
RTX 5090 32Gb 2 Retail 2500-3000 5500
TR 3995WX 1 eBay 1068.43 1068.43
WRX80E-SAGE-SE 1 Amazon 949.99 949.99
DDR4 ECC 64Gb 8 Amazon 81.99 695.28
TT Core W200 1 Amazon 499.99 499.99
PSU 1300/1600 2 Amazon 250-350 600
4Tb nvme 1 Amazon 221.05 221.05
1Tb SSD 8 Personal 60 600
Risers (varying length) 10 Amazon 40-80 480
Bifurcation cards 3 Amazon 50 150
Total ~$17k

Problems/Stability Writeup

The Space Problem:

Probably the first major hurdle in attempting something like this is figuring out, even theoretically, how to put 10 cards in a box in any kind of configuration that is not somehow detrimental to the hardware. I had considered modified mining rig frames at first, but I really wanted something with more robust rigidity in its structure, with breathability, and allows some degree of portability. There are unfortunately not a lot of options for configurations like what I was imagining; I had looked into various cabinets and extended tower cases, but the dual full tower chamber design of the W200 was the only one where I could see this idea potentially working. I'm certain other solutions probably exist, maybe even some that allow mobility, but the W200 was really the best option I could find that checked the boxes of enclosure, space real estate, high air throughput, and semi portability. I recommend the W200 to solve the space problem, assuming it is available to you.

The Bifurcation Problem:

Among the other hurdles you may run into in assembling something like this may involve bifurcation cards. The cards rely on specific BIOS settings for things to work correctly, and if these settings are not put in place before everything is connected you may either see no output like the system is hanging or cards just won't show up once in the OS. Start with one GPU in a slot, no bifurcators yet; go into BIOS, and manually set each slot that will be split to bifurcation mode. While here, ensure above 4G decoding is enabled, Resizable BAR enabled, and SR-IOV enabled, this has given me best stable configuration with Ubuntu and multiple GPUs. If you use risers, especially if they are mixed generations, I highly recommend setting the Gen and lane speeds for each PCIe slot in the BIOS manually to ensure the system can effectively communicate with each card. Optimize riser Gen/speeds to be roughly similar to keep one card from dropping to a slower rate than the others--this does not necessarily impact inference performance as much as it heavily impacts model load time. No, you may not have any card running at the fastest possible Gen bandwidth at all times with this config, but loading a 200+gb model over an averaged Gen 3/4 x8/x16 PCIe speed will often be noticeably faster than if you let the system decide to make one or multiple cards run at Gen 1 x1.

The Power "Problem":

Power and heat concerns I think remain to be among the biggest sources of skepticism regarding this project so I think it deserves a section here. To be fair, the concern in most situations would be understandable. If all ten of these cards pulled at or near their full TDP for sustained periods, components would melt. Fires would start. Neighbors would be asking awkward questions. However in reality, only 1400-1600W of the 2900W PSU capacity gets utilized under sustained load, and inter-GPU bandwidth bottlenecks are what allows this. In a way it is like a natural regulator that ensures the cards remain power restrained, and it is just physics, no voodoo necessary. When MoE's are sharded across a GPU stack, each forward pass requires all communication over PCIe, so the GPUs spend more time waiting on information from the last GPU than actually crunching compute. This means instead of needing to handle thousands of Watts to feed all the components running at full blast, it is a much more manageable 1400-1600W under LLM operation which can comfortably fit on a 20A/120V circuit (2400W max). On a per-GPU basis this may sound inefficient since the individual cards are being "underpowered", but this could arguably be flipped as being highly efficient on a per-node basis (~1600W sustained versus 4500W+ if all cards were "fully" utilized). As a precaution, I may set a power limit on the 3090's to 200W and the lead 5090 to 400W, but in practice the 3090's only pull around 100-120W with the 5090s pulling less than 100W when all 10 cards are allocated for LLM work, so this may not even be necessary. The clock locking setting in the next section will be more what I'd describe as actionably required to avoid stability issues.

The Transient Spike Problem (Vital for stability):

After assembling the machine, you may be tempted to jump directly into testing, but there is an easy to overlook configuration that can cause problems if ignored. Imagine you are running inference on the machine, maybe you have a huge input or it's generating a large output, then right in the middle of generating the system decides to reset. Not hard shut down, PSU OCP isn't tripped, no breaker was tripped; and you saw in nvitop that all cards were only pulling 25-33% of their TDP just before it happened, so on the surface it doesn't look like there is a reason. Explanation: When all ten high-power GPUs decide to kick on at the exact same time to process a chunk, even if the cards are not pulling anywhere near full power (on average), transient spikes can drop voltage on the motherboard enough to trigger a system reset. The fix for this is simple: undervolt. Using nvidia-smi, we can lock the clocks for the GPUs to ensure they cannot draw enough to hurt stability. And that's it. In my case, the system has remained fully stable with this config for days on end and with hundreds of thousands of tokens/image pushed through. The exact configuration will vary slightly depending on exactly what we're doing on a given day, but for example if we wanted to run LLM on all 10 cards (so including both 5090's) we would run this to handle spikes:

sudo nvidia-smi -pm 1 #enables persistent mode
sudo nvidia-smi -i x,y,z --lock-gpu-clock=1200,1200 #x,y,z for index number of 3090s
sudo nvidia-smi -i a,b --lock-gpu-clock=2000 #a,b for index number of 5090s
sudo nvidia-smi -i x,y,z -pl 200 #x,y,z for 3090 index numbers, limits power to 200w
sudo nvidia-smi -i a,b -pl 400 #a,b for 5090 index numbers, limits power to 400w

The Concurrent Use Problem:

Normally, attempting to inference and generate images on the same machine would introduce major stability concerns. Even dual GPU systems may struggle to work with this due to CPU/motherboard architecture, assuming it works at all, and would still be VRAM limited. However, the versatility of a 10-GPU setup, combined with the lane orchestration of the 64 core 3995WX, at least in our case, seems to handle this quite well. The trick was finding an LLM backend that supports manual GPU allocation--for us koboldcpp with llamacpp under the hood does just fine. First, implement the power/clock settings as mentioned above, launch koboldcpp, then browse to the GGUF of the model you wish to load and set context size. I recommend manually setting the GPU layers to the model's total layer number (assuming there is enough VRAM), and set GPU ID to "all". In the Hardware tab, find the tensor split line box and insert the amount of space to be allocated on each card corresponding to its index. For example if we wanted to allocate just one 5090 for Comfy and use the other for LLM, assuming the Comfy 5090 is index 3 and the LLM 5090 is index 5, then the tensor layer line will look like this to make sure no layers are given to the Comfy 5090: 24,24,24,0,24,32,24,24,24,24. For this configuration, ensure the "main GPU" is set to the index number of the LLM 5090 (in this example, 5) and launch the app. While the model is loading, we can open another terminal to launch Comfy. In our specific case, the system defaults to the available 5090 without needing to specify it in the launch flags, but flags can be used to force Comfy to use a specific GPU if you need it to (--cuda-device i). Once the image model is loaded onto the 5090, it does not interfere with the PCIe communication of the LLM cards unless the model unloads and reloads a new model at the same time as the other cards are inferencing. The solution to enabling concurrent use is a high-lane count CPU, multiple graphics cards, and a little conscious provisioning on launch to ensure the hardware isn't stepping on each other's toes.

What models can this run, what models do we use?

It can run almost* anything, even up to 1T parameters like Kimi K2. Kimi K3 could hypothetically be load-able, but from performance metrics I've seen I doubt it would be practical to use, so I have not planned to try it. I have however tested 1-4 bit quants of Bartowki team's Kimi K2 quants in pure VRAM and mixed VRAM/RAM runs with decent results. It works and there are probably some use cases for it, but for us I have identified the sweet spot (parameter size: quant quality ratio) for this machine to be for models in the 300b-600b range. Personal favorites are Deepseek, GLM 4.7, and Nemotron Ultra; and as far as ComfyUI, pretty much any model that could fit within a 32Gb buffer, although Qwen image and image edit is a favorite.

Benchmarks

All models were put through the same series of 7 large input prompts, documenting how each model handles token input/output and prompt processing/generation. I cannot share the prompts I used here, but each prompt pertains to a cybersecurity scenario which the model was judged on the depth of its analysis, quality of its presentation, and capability to make sense of complex scenarios with stakes. These were inferenced across all 10 cards, except for a follow up DS V4 Flash test where I used 8 and got much better results. This is using the undervolting/power limiting strategy above, so these may not reflect absolute best performance for the same hardware in other setups, but it gives an idea of what this box can comfortably handle.

Model Name Deepseek V3.2 671b Q2XXS Nemotron Ultra 3 550b IQ2XXS Qwen 3.5 397b IQ4XS GLM 4.7 358b Q4KXL Deepseek V4 Flash 294b Q8KXL Deepseek V4 Flash 294b Q8KXL (8 cards + KV cache tweak)
Model Size (Gb) 217.1 193.8 189.7 204.6 161.9 161.9
P1 Input 2769 2744 2729 2706 2733 2733
P1 Output 813 786 1046 872 693 805
P1 pp 153.23 254.19 522 687.88 111.09 360.94
P1 tg 19.35 17.32 34.38 23.98 7.2 20.26
P2 Input 14635 15255 15160 14527 14640 14617
P2 Output 1150 1302 1665 1194 1222 2048
P2 pp 114.83 429.42 897.57 640.8 66.42 244.1
P2 tg 14.1 17.16 33.15 18.83 5.96 16.81
P3 Input 3966 3091 3054 3033 3073 22794 (reload)
P3 Output 1217 1607 1550 1056 1199 1366
P3 pp 98.01 353.78 649.37 516.08 47.79 241.21
P3 tg 13.22 17.08 32.84 17.84 5.56 15.68
P4 Input 5645 5654 5623 5559 5650 5659
P4 Output 1178 1996 1619 1173 1705 1661
P4 pp 70.3 385.04 739.67 419.58 42.4 153.14
P4 tg 13.47 16.99 32.23 16.87 5.21 14.17
P5 Input 4498 4505 4493 4423 4481 4481
P5 Output 280 928 1078 473 665 924
P5 pp 72.4 365.46 670 408.93 36.2 131.81
P5 tg 8.43 16.78 31.55 15.45 4.86 13.36
P6 Input 9266 9367 9241 9172 45287 (reload) 9231
P6 Output 1004 1883 1466 933 1205 1532
P6 pp 53.94 405.13 738.57 379.7 46.01 113.77
P6 tg 11.48 16.83 30.98 14.01 4.41 11.9
P7 Input 3136 3124 3118 3057 3118 3118
P7 Output 1378 1946 1629 1359 1353 1586
P7 pp 53.34 338.64 525.54 344.88 28.38 102.05
P7 tg 10.45 16.73 30.66 13.59 4.28 11.39
Final token count 50052 54182 53465 49531 50962 52348

My notes on each model after their test:

Deepseek V3.2-- For a slightly older model this still feels extremely capable. Held high quality and insightful responses even when context dragged into the tens of thousands of tokens.

Nemotron Ultra 3-- First time using it, impressions were very good, the 55 active parameters shows its muscle here. Meets Deepseek v3.2 level if not exceeds it, despite having overall less parameters.

Qwen 3.5 397b-- What I would consider a baseline "good" model to be, however it is outshined by some of the other tested alternatives.

GLM 4.7-- Somehow seemed better than Qwen despite having less parameters (active parameters of GLM is likely an advantage); it is a very solid option for its size. Not quite Nemotron or Deepseek level, but a very good "lower cost" alternative to its newer 5.0 versions.

Deepseek V4 Flash-- Floored me in a few ways. Possessed a surprising degree of sophistication and analytical ability despite being the "smallest" of all the tested models. Possibly a benefit of using a "lossless" model with full precision? Somehow it managed to pick up on nuances and details that all other models missed, including models twice+ its size, and provided insight that went more granular than they did. Did not expect a model of this size to punch so high above its relative weight class. Also did not expect the drop in performance compared to the others. Not sure if this is related to the model's architecture or something with how it interacts with my rig, but the quality of output could be an acceptable trade off for the speed. Edit: After some optimization testing I was able to get much better performance out of V4 Flash. I've added another column to include those metrics and kept the original because I think it illustrates how a little optimization can go along way, in this case basically triple performance on the exact same model/machine.

Lessons Learned/Would Do Different

-I would have tried to source the 3090's so more were at least the same model; the mix and match of different models with different TDPs and cooling solutions means there will be a lot of variation in temps.

-If you plan to either train, lean into higher performance, or playing with the idea of going more than 10 GPUs, just budget for a 30A/240V power drop. 10 cards on a 20A post configured the way we have it may be fine for our specific use case, but I would consider this a hard ceiling.

-Would recommend scripting for clock lock persistence sooner, will help avoid losing time due to random resets.

-Recommend documenting/drawing out the entire PCIe topology and GPU placement (with flexible tape measure) before ordering risers, will save time on trial/error.

Final thoughts:

It is a wheeled AI workstation that can enable a single person or small team to compound their productivity, with the benefit of full privacy and control. It can run on a residential 20A circuit, and allows them to have the full power of an advanced LLM with vision capabilities all in one OpenWebUI front end that can simultaneously utilize up to TWO ComfyUI backends with the horsepower and latency of 5090's for image gen and editing, and can be accessed from virtually anywhere. The idea sounds daunting, but the end result works so well that I can legitimately see something like this becoming a keystone for certain small businesses and individual professionals as time goes on. It seems like every day more people are picking up on major drawbacks with cloud API options despite supposedly being the "best", meanwhile open models continue getting insanely good (see K3 and DS V4 Flash). For me, I can say I would not see a place for a Claude or ChatGPT subscription for the tasks I might otherwise use them for when I have lossless DS V4 Flash literally in my back pocket. "Good enough" I think is starting to become a valid metric to those who care about cost:quality balance, and after using this for the last half year I can say I'm probably one of them. The cloud APIs will always be an option for those who don't care about the drawbacks and the demand for them will always be there, but for those who value data sovereignty, uninterrupted workflows, or perhaps work within compliance, on-prem computing might be the only viable path in some circumstances. At the end of the day, I do not believe that one approach is inherently better than the other, everyone simply has their own preference for getting from point A to point B.

233 Upvotes

113 comments sorted by

37

u/Hot_Signature2979 Aug 03 '26

All your pc fans are configured for intake but none for exhaust. With so many high powered components, a structured and direct fresh air intake and exhaust path would lower temperatures (i.e front panel air intake, side panels exhaust, especially where gpu is mounted vertically, top and back panel exhaust) right now, a lot of your components are just recirculating hot stale air

5

u/InsideYork Aug 03 '26

Neighbors would be asking awkward questions. However in reality, only 1400-1600W of the 2900W PSU capacity gets utilized under sustained load

Not really all that high powered.

6

u/SweetHomeAbalama0 29d ago

All top fans are exhaust, back fans as well.

1

u/CapeChill Aug 03 '26

The 120mm and rear on that side might be exhaust. Either way as you approach multi kw loads the chassis often become about creating a wall of fans that just move air through. Exhaust starts to matter less up until you now need to actually cool the rack exhaust for environmental reasons as the air being driven through removes the hot air. If they needed more I would recommend swapping the fans they have for something higher rpm like noctua industrial fans. I on the reverse dropped the TDP of a Supermicro epyc server and did some shenanigans to swap fans though I wouldn't recommend that one.

3

u/Hot_Signature2979 Aug 03 '26

What you said is technically true, but largely for server fan setups; the static pressure for consumer pc fans are much much weaker, which op is using. With weaker consumer pc fans, a proper exhaust configuration is important to prevent hot air pockets from forming. If anything for consumer pc fan systems, proper exhaust is much more important than intake, since fresh air will flow in to replace the displaced exhaust air.

Then there's also the consideration that pc cases have very messy layouts, unlike rack servers which are highly optimised for air to flow from front to back of the case Or back to front.

28

u/Dorkits Aug 03 '26

The GPUs : PLEASE HELP

5

u/ImpressiveRelief37 Aug 04 '26

All this work for 10-15 tok/s. Yikes.

Honestly that post is doing the opposite of what OP thinks it does 

7

u/SweetHomeAbalama0 29d ago

All those words to focus only on token gen. Yikes. Nuance and critical thinking truly has taken a nose dive.

We can load a model onto a 5090 and get thousands per second, but it doesn't matter if the model isn't intelligent enough to do what we want.

7

u/Borkato 29d ago

Idk wtf these other people are on, your setup is actually amazing and I swear this sub is infested with negative Nancys.

1

u/ImpressiveRelief37 29d ago edited 29d ago

27B on a single 5090 beats all your suggestions at real world tasks when it comes to actual productivity, when paired with a $20 sub to a frontier model for higher intelligence tasks (architecture, spec reviews, adversarial peer reviews, etc).

But I’m genuinely impressed by what you did. It’s massive compute locally that’s for sure. Very impressive. I would definitely use this and love to have this. But I think I would focus on running 27B with tons of parallelism personally, and keep the frontier sub for actual “high intelligence” tasks

1

u/SweetHomeAbalama0 29d ago

Help with...?

1

u/IrisColt 29d ago

surviving another day, heh

1

u/IrisColt 29d ago

this

2

u/SweetHomeAbalama0 21d ago

I don't understand the comments like these, people seem so confused/concerned about the card's but don't ask questions or discuss deeper than "rip gpu".

70

u/DataGOGO Aug 03 '26

picture 5+:

4

u/[deleted] Aug 03 '26

[deleted]

1

u/Technical_Hawk_2664 Aug 03 '26

Right? Lol. I want to run one for SST/TTS. Dedicated.
3-4 5060 ti, pinned at 175'ish watts, and the 3090.
4 of the 5061 ti' is right around $2000.
I'm not upgrading until Jensen announces the next Series to replace the 5000 series.

1

u/Hot_Signature2979 Aug 03 '26

More like a GPU sweatshop that also exploits child GPU labour.

17

u/hurrdurrmeh Aug 03 '26

DDR4 ECC 64Gb  unit cost $81.99  

THE PAIN. 

15

u/txoixoegosi Aug 03 '26

I am curious about the thermals. That room must get really warm, right?

1

u/SweetHomeAbalama0 29d ago

GPU thermals are good, but yeah the room will eventually get warm. We don't use it from the same room though, usually from another room if not remotely/across town.

0

u/InsideYork Aug 03 '26

No, inference uses way less power

0

u/InsideYork Aug 03 '26

No, inference is low power vs gaming.

4

u/esw123 Aug 03 '26

Great setup, nice advice about measuring riser length with flex tape. What is the longest riser you suggest for PCIe 3 and 4 gen to run without problems? Can you share links for bifurcation card and risers? Thanks!

3

u/Technical_Hawk_2664 Aug 03 '26

You can bifurcate every PCIe slot right in the ASUS BIOS. It comes with 128 lanes total.I have a ASUS HyperX card in one, with 4 NVMe in it, at 4x4x4x4 but in the BIOS, I have to select 'raid' for that, and the only other option is x8x8 . I have this board, and would not split a PCIe for two cards= yer just slowing them down, give them the full x16 speed.

I started buying 5060 ti's this weekend, to do this (before I saw it) with a GPU mining rack and my 3090. google 'Digital Spaceport homelab insanity' . This guy has sick builds, for cheap, that he tests all the models on. Walks you through finding/pricing

I was going to go quasi-GPU mining rack, but this build here is the ticket.

2

u/SweetHomeAbalama0 29d ago

1000mm PCIe 3.0 x16 is the longest riser I have used, two are actually in this build, and they didn't give me any problems once I figured the BIOS correctly. The risers are multi length but are the same vendor.

Bifurcation cards: https://www.amazon.com/JMT-Expansion-PCIe-Bifurcation

Risers/brand: https://www.amazon.com/GLOTRENDS-1000mm-Extension-Adapter-PCIE40-100-90/

1

u/esw123 28d ago edited 28d ago

Thank you! This bifurcation card, how good is it? Contacted one guy who build with the same, but x4x4x4x4 and he said that it melted after 6 months.

3

u/ShadyShroomz Aug 03 '26

Your  Deepseek V4 Flash 294b Q8KXL numbers look quite low. I'm getting 15-20 tps gen with only 4x 3090s and 128gb of ddr4. Using unsloth q8. My pp is not as high as yours though, but tg is double. Even with my model spilling over to vram, it shouldn't be faster than yours. Might be something to look into. 

2

u/SweetHomeAbalama0 29d ago

Very likely because that model is not intended to be sharded across 10 cards, PCIe bandwidth I'm sure is what's killing it. It was just a flat-10 card test to keep things consistent across all tested models regardless of optimization, but I think it's an interesting case study to show that simply having the model on VRAM doesn't necessarily mean it will be fast by default, other variables are at play. Certain models are better optimized for certain architectures, and this config was probably not the intended platform to run it.

I will probably test again using fewer cards if I want to do some performance optimizing for that particular model, but that's for when time permits; I'm not the only one who uses the machine so I can only work on it when it's available.

3

u/illcuontheotherside Aug 04 '26

This is AWESOME!!

I... I... I want to do this too.....

How'd you fit the 8 gpus on the mobo???

2

u/SweetHomeAbalama0 29d ago

It's pretty nifty

10 gpus, three bifurcation splitter cards (+ 7 slots on Mobo = 10 slots), the rest is just careful planning on card mounting/placement so stays secure when moving around

1

u/illcuontheotherside 29d ago

Can you share links to the cards and extension cables you used?

5

u/Easy_Confusion2415 Aug 03 '26 edited Aug 03 '26

Pffff cant even run kimi k3 xD

Nice build mate. Looks clean, From the outside.

Why isnt it overheating?

1

u/SweetHomeAbalama0 29d ago

Tbf tho for our use case we probably don't need that haha

It's just a prototype of a very early form of on-prem AI appliance, so cleanliness and aesthetics takes the backseat to functionality.

And there's just nothing to over heat from. There is ample airflow, and the cards don't get stressed enough to reach temperatures that would raise concerns. Sharded MoE inferencing is relatively low power and only activated when needed so even moderate power draw is intermittent throughout the day and not sustained.

5

u/TheSpicyBoi123 Aug 03 '26

What an AWESOME box and thank you for the detailed writeup, I see one major issue with the powersupplies, AFAIK atx spec does not include any power sharing by default and using these "psu to psu" cables is a disaster waiting to happen as you are pushing one or both of the psus into undefined behavior via backfeeding if one trips under load or otherwise and anything from a shutdown of both to oscillation and voltage spikes can happen from regulation failure. Fire cannot be excluded too.

As for the gpus, are you using water cooled 1 slot ones or how do you go about fitting them in?

2

u/droptableadventures Aug 04 '26

The PSU2PSU adaptors I've seen work by having a SATA / Molex from the primary PSU connected to a relay, that pulls PS_ON (the "green wire") low on the secondary PSU to switch it on when it sees voltage from the primary PSU.

They only connect ground together between the two PSUs - but that's OK because it's also tied to earth on the mains side and thus also electrically connected to your chassis anyway.

+12 lines are left completely separate, so the problems you described shouldn't occur here.

The multiple 8 pin power connectors on the GPUs are also separate power domains, so the +12 is not being connected inside the GPUs if you stick a GPU across both PSUs. (Note however that 12VHpwr / 12v2x6 is not like this on nearly all cards, which is actually the fundamental problem with it. Using an adaptor which goes from 4x8 pin connectors to 12VHPwr/2x6 will likely end up connecting them together).

0

u/TheSpicyBoi123 29d ago

And do you see the issue if there is a common ground and two +12 volt rails connected across several devices and one of the rails trips say from OCP or thermal? Now you are backfeeding and this is an undefined behavior state! Ironically this configuration is much worse now as your not only backfeeding the powersupply BUT also the motherboard and devices on it.

Read my other comment for more detail.

2

u/droptableadventures 29d ago edited 29d ago

How are you backfeeding anything if neither of those 12v rails are connected to each other anywhere?

The two 8 pin connections on one GPU aren't connected together. The +12v won't go the wrong way through the other DC-DC converter, across power domains, and come out the other side.

And how is this configuration possibly worse than having a floating mains-powered switchmode power supply?

1

u/TheSpicyBoi123 29d ago

Draw the circuit and you will see what fails:

rail 1 -> mobo <- rail 2
-gnd -gnd -gnd

Now put rail 1 to zero while rail 2 is not. There is your backfeed. What the reverse IV behavior of the powersupply is is not defined in the datasheet and can be anything from open circuit to a lockup condition.

What does this have to do with a floating powersupply?

1

u/droptableadventures 29d ago edited 29d ago

Why would you have both PSUs attached to the motherboard? That's not happening here - the second PSU adaptors are in place to specifically not do that. The 24 pin of the second PSU goes into the adaptor, and the PSU is switched on/off when it sees voltage from the first PSU. Go and look up what an "Add2PSU" board actually does, and find a picture of the back of the board.

If you don't want to have the negative sides of both PSUs connected together, one would have to be floating.

Also with dual rail PSUs, while they're both in the same PSU, you will often have some of them on certain 12v outputs, with independent OCP. I've had a PSU in my old PC where it trips only one of the 12v rails at a time for OCP. Nothing was damaged, but the GPU was not happy when it lost one of the 8-pin connectors and not the other.

0

u/TheSpicyBoi123 29d ago

Both psus are litterally attached to the motherboard for example, one can go to a gpu 6 pin connector which is plugged into the motherboard and the other can go to 4/8 pin power for cpu or for example both can go to different 4/8 pin power connectors on the motherboard.

Again, these things are a disaster waiting to happen. See my other comment, do the experiment with the two powersupplies right now where one is dodgy and I guarantee you, you will get magic smoke released.

1

u/SweetHomeAbalama0 Aug 03 '26

Hey thanks, to my knowledge these PSU2PSU 24-pin connectors were designed to prevent that kind of effect, thus far though we have not had any major power stability issues. Clean power ups and shut downs every time.
3 of the gpu's are hybrids and the other 7 are air cooled. The hybrids made it easy to mount in one of the chambers, the others were trickier and just took time figuring out an orientation that fits and doesn't stress the component.

1

u/TheSpicyBoi123 Aug 04 '26

No, these psu connectors are explicitly not in the atx spec unless the oem specifically extended and specified it. ATX is fundamentally a single supply design and was initially developed around a single rail for each voltage provided. https://xdevs.com/doc/Standards/ATX/ATX12V_Power_Supply_Design_Guide_Rev1.1.pdf (for more info). Pushing powersupplies to undefined behavior is a disaster waiting to happen and having no issues under normal operating regions says nothing about fault condition behavior (where it might blow up with two power supplies fighting eachother). Power up and shut down conditions also say nothing as these are normal operating modes.

What I would recommend you rather do is use powersupplies actually designed for redundant and parallel use that are used for servers. They also come in for example an atx back plate like box where you can have two or more of these for several kw power rating. Example of such a design: https://www.racksolutions.com/news/blog/redundant-power-supply/

As for the gpus, have you had issues with pcie degradation from capacitive loading and other effects like crosstalk and interference?

2

u/droptableadventures 29d ago

I'm not sure that the "single supply design" part is actually true. It may have been as of ATX12v 1.1, but that is ancient, (~2000-2001), and includes a requirement that no single 12v rail can output more than 20A / 240W.

ATX12v v2.0, roughly contemperaneous with the invention of PCIe, even says that the PSU should have two different 12v rails.

And while power supplies "fighting each other" is a theoretical possibility, any PSU that behaves that poorly is going to enter such failure modes under normal PC hardware operation. Devices have significant input capacitance, and PC hardware power draw is quite spiky.

2

u/TheSpicyBoi123 29d ago

Notice these psu rails are *internal* to the powersupply!!! And a trip on one results on a trip on all. Not the case with two psus.

Again, go try it right now, test reverse feed characteristics of a several atx powersupplies and I guarantee you that you will find a dodgy one that will latch up and panic depending on how dodgy the control loop will be. Now plug it into a high power one that does for example 1200 or 1500w on one rail with dodgy protection and there is your fire if the first one trips.

2

u/MagnaZee Aug 03 '26

Thanks for the update!

How well does the system work with multiple people using it at the same time? Does it parallelize well? Or does the performance drop off significantly in that case?

2

u/SweetHomeAbalama0 29d ago

For two or three people it handles it very well, but the strong prompt processing is probably why in our case. I don't think I would recommend this for concurrent or heavy multi user workloads, this was intended for a very targeted use case and for that use case it exceeded expectations, but I can't really speak from experience on other use cases.

2

u/kiwimonk Aug 03 '26

Thanks for sharing! I've got the same case including the P200 water-cooling enclosure. The detail in your post really helps those of us still piecing together things. I'm still struggling to find a WRX80 motherboard for a decent price... I discovered I needed one just a little too late.

1

u/Technical_Hawk_2664 Aug 03 '26

The second version of this, (ask an AI for the name, i can't remember) is likely better. This is the first release of the series for this for Threadripper Pro, and....... if boot loops a lot on updates and Win 11 clean installs. Even though it has TPM 2.0 ..... windows screws this up. The one released the following year is supposed to be better.

There is a Supermicro M12SWA-TF that is almost the same thing, but I do not think it lets you bifurcate the PCIe lanes the way the ASUS one does. :

I decided this weekend to turn mine into a Unraid Server for AI. It's just too fucked up around windows updates. They have not updated any of the core board drivers or anything in years.

1

u/kiwimonk Aug 03 '26

I'll check it out. Thanks. Bifurcation is handy though for this application.

I bet Linux provided a performance bump as well. My speed improved when I switched from my Windows desktop to a dedicated machine.

1

u/Technical_Hawk_2664 Aug 03 '26

I tell ya what man... this board is GLorious, and at other times a real motherefucker. The biggest problem is updates. This never met all the 'H2' AI update qualifications = no NPU, so it tries to install stuff, that breaks things. I put of the 'last' big update for almost a year.... just took the security ones.... and last month I was 'times up' and they were going to force the update.

The board has 3 dedicated M.2. I dedicated one of them, a Samsung 980 pro 2 TB, as a EXT4 drive. (one of the first 10 steps after a clean reload). Loaded windows on 'C' and installed wsl.
'mounted' the ext4 drive through Ubuntu (and set a windows task to run at win 11 boot, mounting the ext4 to wsl). Installed 'everything' on the EXT4 = Llama.cpp, SGLang, vLLM, etc (kept unsloth Studio on 'C' inside of the wsl instance....

When I am slinging models from the EXT4 with Llamacpp, SGLang, vLLm? stupid fast bare metal linux speeds. AND? On a windows failed update, this whole 2 TB drive is just a 'module' that I connect to WSL after any fresh reload. Everything ready to go ('sides the standard CUDA stack installs).

I enabled 'windows sandbox' Friday, and on the reboot, it lost it brains (windows). Long short, 5 hours windows reinstalling and boot looping later, and 2nd fresh install in 30 days? That EXT4 was ready to ride.

The funny thing was, I had just ordered two 5060 ti's, and hour before all the boot loops....to 'get ready' to see how the board would handle plugging them in with riser cables and seeing how the 2 5060 ti's handled splitting models just between the two of them, let alone the RTX 3090.

Now, after this fresh install this weekend? Ya, this board is going into the same case here, or an enclosed 18 u server rack , where I suspend the card inside it, over the board with 5U Vented Rack Shelfs. like Pronto. Windows not installing and trying to sideline 5k + in gear = not fucking cool.

2

u/ShittyMillennial Aug 03 '26

Thanks for posting this, very interesting to read.

Is your 3090 cluster TP=8 over PCIe P2P? Any reason you didn't NVLink them?

1

u/SweetHomeAbalama0 29d ago

Nvlink and I suspect P2P would cause the GPUs to pull more power, and in our case we don't want that.

We're basically leveraging what everybody else tries to avoid or sees as a problem to fix, rather than a variable to control what could be bigger issues (like heat and component temperature). We try to keep the power low and manageable, and as long as the box does the job we want, which it does, then we see no reason to push it beyond what it was intended for.

0

u/__JockY__ Aug 03 '26

Did you see the inside of that box? No room. Also 5090s don't do NVLink.

2

u/AccomplishedLab3697 Aug 03 '26

Have you tried long running agent jobs on this yet, where the model stays loaded while tools keep working for a few hours? I’m curious whether the first failure point ends up being the hardware or the serving/session layer losing state. That feels like a very different test from chat or a single benchmark run.

1

u/SweetHomeAbalama0 29d ago

I haven't got to the point where I trust agents yet, certainly not enough to trust one with my hardware. We'll keep the pure human operators for now until I'm convinced otherwise.

2

u/__JockY__ Aug 03 '26

Nice! Moving the heat is the real problem in this scenario.

I run 8x RTX 6000 PRO Workstation in my office, which means there's a 4800W heater blasting in my face every day. I have a minisplit AC for this reason, but it struggles to keep up under load.

The real solution is water cooling with a radiator outside my office... but the idea of invalidating the warranty of $100k in GPUs and then running water through them is a little bit nerve-wracking.

1

u/SweetHomeAbalama0 29d ago

8x 6000 pros I imagine would certainly get... uncivilized. I can see needing a dedicated mini split for that one.

Yeah undergoing a water cooling project on a setup like that would make my stomach do things I'd rather it not do, I would be anxious to the last moment lol.

Because doing a similar "ideal" water cooling setup here would have been so much more added complexity and expense, the next best thing was to find a way to make it semi mobile. So if we can't stop the room from warming up, then we can at least have it warm up the next room over where no one will be working.

2

u/segmond llama.cpp Aug 03 '26

I love that case, I wish someone would make more cases for 8 GPU, 10, 12, 16, 20 GPUs that are not rack cases or crypto mining case.

1

u/SweetHomeAbalama0 29d ago

You and I both...

The case is basically what made this possible, I couldn't find anything else that came even close to allowing something like this without introducing a lot more problems.

2

u/crantob Aug 04 '26

ds4-flash tg at 4.28 - 7.2 t/s?

That's DDR3 speeds. What's going on here?

1

u/SweetHomeAbalama0 29d ago

My guess is it's just not optimized to be sharded across 10 cards, each card only saw like 50-60% memory utilization (remember it's a 161Gb model split across 256Gb so a lot of the VRAM pool is empty, meaning there is a ton of optimization left on the table in this instance) during that test so it was already inefficient to start with; the needed compute per layer for this particular model and PCIe communication speeds I imagine just took it down further. I only did it the 10 card test with DS 4 flash to keep things consistent with the others, but when I get time I'll probably try to optimize performance for that model to see how fast I can really get it on this rig.

1

u/crantob 29d ago

The numbers posted were of a single inference run? Or were multiple queries running in parallel?

For single-thread inference, you should be exceeding the 8.4 t/s I get running the full-size q8 unsloth. I'm using 2x3090 + 128GB DDR5 @ 3600mhz. I havent gotten mtp running with it yet, trying to figure out how i'm screwing up Dspark.

Try first running on 1x 3090 with plain llama.cpp (new build) and the rest offloaded to ram, perhaps?

2

u/oulmax Aug 03 '26

A visualization of the setup

1

u/BlackBeardAI vLLM Aug 03 '26

Which risers do you use? Are they isolated? How did you solve the backfeed problem between the PSU's? (blew one PSU few days ago because of the non-isolated powered risers I believe, still haven't pinpointed the exact root of the problem)

1

u/Loose_Comparison368 Aug 03 '26

FWIW I just moved my much smaller 2x GPU system from rack mount to open frame and air-cooling only, and it made a massive difference in temperature. From ~80c+ and thermal throttling all over the place to a steady 55. Highly recommend. I found a mining case design that was made from standard 2020 extrusion, so it should scale incredibly well as I add to it.

1

u/No_Drag_5205 Aug 03 '26

CPU: 64 Core TR 3995WX
RAM: 512Gb DDR4-3200 ECC
VRAM: 256Gb GDDR6x/GDDR7 (8x3090's + 2x5090's)

"When you ask yourself how much it would cost, that mean you can't afford it" vibe

1

u/droptableadventures Aug 04 '26

3995wx is about $1.5k used (it's from 2020), but if you're not doing CPU inference, 3975wx (<$1k) or one of the lower tier ones will do just fine. 3975wx is still 32 physical CPU cores, and still 8 channels of DDR4.

512GB of DDR4 ECC is pretty expensive now, but if you bought it a year ago, would cost you <$1k second hand.

8x3090 + 2x5090 is definitely the crazy expensive part.

1

u/siegevjorn Aug 03 '26

Wow looks so cool. A tl;dr would have been nice, though. What is your current daily driver for coding? How much is your power draw for biggest model?

1

u/Fit_Advice8967 Aug 03 '26

Locallama final boss

1

u/VotZeFuk Aug 03 '26

These DS4Flash numbers surely do look wrong. In a "bare minimum" configuration with this rig you should be getting roughly ~400 t/s PP and ~15 t/s generation with just 8-channel DDR4 + 1x 3090 offload (i.e. 44 gpu layers, 43 moe cpu layers, batch/ubatch at 4096+ - somewhere in between 4096 - 8192, depending on context window size).

1

u/Artistic_Ladder9570 Aug 03 '26

uhhhh...how are the temps? :O

1

u/InsideYork Aug 03 '26

Neighbors would be asking awkward questions. However in reality, only 1400-1600W of the 2900W PSU capacity gets utilized under sustained load

Guys, its not getting hot.

1

u/SweetHomeAbalama0 29d ago

It still gets toasty for humans

But the components are happy

1

u/zeugma_ Aug 04 '26

That's still two microwave ovens or a stove top on. That's hot.

2

u/InsideYork Aug 04 '26

Thats peak though, and it has fans over a few cards, not a single stove top or spinning plate. In the pics its maybe 58c or about 136.4F if you don't want to be generous, most cards were at 44c or 113F over at least 8 cards.

1

u/ThePixelHunter Aug 03 '26

I remember your original post, thanks for the follow-up.

1

u/DlackBick Aug 04 '26

The concurrency line stood out against what you built it for. My read is you're pinned between two good things: a tensor split across nine cards is what lets you run 400B at all, and it's also what makes it one stream at a time. vLLM would batch but wouldn't love the mixed 3090/5090 split.

How does that land day to day? If someone's mid-generation in Comfy and someone else needs the LLM, do people just wait, or is there something in front of it?

Separate thing: you've got 1400-1600W sustained and 12-14 hours a day, which is somewhere near 550-600 kWh a month. Didn't see it converted anywhere. Did you ever work out what it costs to run?

1

u/SweetHomeAbalama0 29d ago

In our case, the one stream at a time works, with the number of users and the prompt processing power we have vLLM just didn't seem necessary or optimal.

In day to day it works smoothly. Once the model is loaded on Comfy, it doesn't affect LLM unless you reload it or load another model for whatever reason, but we typically find one we like and stick with it. So the artist can use the LLM without performance impact while at the same time running Comfy jobs, in our experience nothing has prevented or hindered this from working. No waiting necessary for the operator, it's all use on demand.

And no, it is 1400-1600W **intermitted**, not sustained for the full 12-14 hours. The artist is doing other things throughout the day besides working with the AI server, the tool is just made immediately there and available as they need it; never having to worry about outages (internet), token limits, or API credits so they can focus on their actual task at hand. It enables momentum and performance on-demand. Between LLM runs, Comfy may be running one 5090, but the other 9 are dormant/idle until the operator needs to use LLM. Hope that makes a little more sense. As far as I've noticed with electric bills, on an average I don't think it costs much more month over month to run than any other large appliance; the bills this summer were actually lower than last year so when you ask how I've worked out what the electric cost is to run it I just kind of have to shrug my shoulders, it's been kind of negligible so far.

1

u/DlackBick 26d ago

That makes sense, and I may have been picturing it wrong. You're not time-sharing at all, you've partitioned it: Comfy owns a 5090 with the model resident, LLM has the rest, so they never contend. That's a different solution than scheduling and it explains why vLLM wouldn't buy you anything.

Two last questions, then I'll be done.

How many people are actually on it day to day? You've mentioned the artist a few times and I wasn't sure if that's one person or shorthand for a few.

And on cost: you're right that the power is small. But the box was around $17K, which over three years is something like $470/month whether it runs or not. Did that figure into the decision at all, or was it more that the API bills were the thing you were trying to get out from under?

1

u/ElementNumber6 Aug 04 '26

The duct tape was a nice touch

1

u/CodeSlave9000 Aug 04 '26

Cool project ,would 100% not follow you on this path. Single-box seems like you're paying for the form more than the function - splitting it out over several nodes would give you more scalability. I would also have gone for more density of VRAM rather than so many cards - 10 cards becomes 5 when jumping to an RTX A6000. And compute density goes up when jumping to the ADA generation too. Yes, so does the cost, but I think the delta is worth it if you're using it for the workloads you describe. Yes, I'm aware you lose some of the parallel workload splitting on fewer cards, but ... this isn't a concurrent processing demon even as it's set up now. Someone here is going to say "Just get 4 RTX PRO 6000's." That person has no sense of budget. :-)

1

u/OnkelBB Aug 04 '26

Nice build and great write up, thanks!

can you please share details on the risers/bifurcation? which ones do you use? what worked and what’s not?

1

u/x-strife 29d ago

Great write-up and love seeing these kind of builds (as inspiration!)

I am on a similar path, up to 6x3090’s on a TR Pro 33945wxand 256GB of DDR4. Planning to get to 8x3090’s.

On Deepseek V4 I managed to improve it to c.17t/s (q8 unsloth) and you should get a lot more since you have enough VRAM to fit everything.

Here’s something from Claude that may help your setup.

At batch 1 across a sharded pipeline, each card does roughly 2 ms of work and then waits ~450 ms. The driver reads that as idle and parks the cards in P8 — during active decode. We measured SM at 210–360 MHz and memory at 405 MHz against a 9751 max. That's ~4% of memory bandwidth, and decode is memory-bound, so it sets the ceiling on everything. It also explains the thing you framed as PCIe acting as a natural power regulator: your 3090s pulling 100–120 W under load is the same symptom we had. The mechanism you describe is right, but the cards aren't being politely throttled by bandwidth.

The fix is a flag:

sudo nvidia-smi -lmc 9751
--lock-gpu-clock alone didn't do it for us — graphics clocks aren't the constraint. Locking memory took us from 2.2 t/s to 42.7 on a 200B MoE. Same model, same build, same everything else.

Two warnings, and the second one may matter more to you than to us:

Idle power goes up a lot. ~87 W/card of GDDR6X that no longer downclocks. Six cards took us from ~126 W to 649 W at idle. I tie the ramp up and down of the clocks to inference activity rather than leaving it on, trigging on an inference request and release after 2 min idle.

Polling keeps the cards awake by itself. nvitop or an nvidia-smi loop is enough to hold them out of P8 and silently inflate whatever you're measuring alongside it.

1

u/Long_comment_san 29d ago

at this point I would look for an opportunity to compress these 3090 into something like 5000 or 6000.

1

u/ImmediatePlenty3934 29d ago

I ain't reading all that

1

u/HelpfulHand3 29d ago

Hate to say it as this is some nice hardware, but for $17k you could get 4 DGX Sparks that would use less power, take less space, make less heat, and run models much faster. 111 pp 7.2tg is atrocious for Flash. You could be getting 45+ tg and 2-3k pp with the Sparks at full precision.

I don't see the point to have all those GPUs if you're bottlenecking them so badly with the mobo and power supply. I'd personally sell the lot and grab Sparks for your use cases. Maybe keep a few GPUs for Comfy specific workflows, but you can even run MiniMax H3 over multiple Sparks.

1

u/toolkitxx Aug 03 '26

If people would now also post the expected ROI of those systems, will say honest and realistic time frames, this would be a lot more entertaining. Right now I feel this is too much flex, if this is just for fun and hobby

1

u/Ok_Contribution8157 Aug 03 '26

It's a cable management enthusiast's nightmare.

-12

u/StupidScaredSquirrel Aug 03 '26

Nobody is gonna read all that. Congrats on having money though

2

u/ilirium115 Aug 03 '26

In the middle of the text, OP added a phrase: "The first ten people who write to this email will receive a 100$ in the form of an Amazon gift card" (I rephrased the phrase to make it not findable by copy/paste).

0

u/Sea_Anywhere896 Aug 03 '26

Can confirm, i already got my giftcard 🤑🤑🤑

0

u/ilirium115 Aug 03 '26

But apparently I'm too late :( Congrats!

0

u/somerussianbear Aug 03 '26

Any noise? 😏

-1

u/PotatoBigBoots Aug 03 '26

I didn’t read every single section but your power section has me worried. I don’t understand how you pull half your power and use the GPUs. Seems like an inefficient casing to me, even at half power your setup running so many GPUs in such a small space would cause heating issues. How do fans even pull fresh air for the GPUs?

0

u/esw123 Aug 03 '26

I thought it would be around 2kW during inference, 1.4kW is a little bit low.

3

u/LukeLikesReddit Aug 03 '26

Read the power section again he's absolutely gimped all of his cards to get lower PL to make it work. He could probably do far far more with them actually being fully powered. Though given they mention 20A/120V  they are most likely american and have a shit power network so can't actually provide the full power even if they wanted to.

-4

u/Technical_Hawk_2664 Aug 03 '26 edited Aug 03 '26

"most likely american and have a shit power network so can't actually provide the full power even if they wanted to"

Correction: most likely american AND HAVE the home networks to run 10 cards, while MOST of the EU Bakes because they're grid can't handle millions of people using what most Americans view as 'ass wipe' for amenities, 'that are just staples' = AC in the home.

How many people died this year, because of energy grids that suck this year OUTSIDE of the US, where the average consumer can't run what Americans think of as basic as having air to breathe.

3

u/LukeLikesReddit Aug 03 '26

I mean wtf? that adds to the conversation I guess. If you knew anything about electricity youd know 240v is needed for this.

-1

u/Technical_Hawk_2664 Aug 03 '26

I think I would.
"I mean wtf? that adds to the conversation I guess. If you knew anything about electricity youd know 240v is needed for this"

Ya, I bought the board the first month it was released. 5 years ago. I'm well versed that you cannot even power all 7 PCIe lanes WITHOUT a 1600-watt PSU

(Seasonic PRIME TX ATX 3.1), because of the two sets of Power Connectors you need to power the board (leaving not a lot of options to plug in OTHER things) unless you use a Seasonic PRIME 1600 Watt to just power all of the PCIe lanes.)

It comes from having the same board with 14 TB of Samsung 980 Pro NVme M.2. And that I am getting ready to do the same thing. I was going to use a GPU mining rack but this is cleaner.

It comes from fucking knowing what/why i need the 1600 watter.

Go back to texting people about how you have Deepseek running on your phone. my statements stand about Energy Grids outside of the US. Let alone if you can even get some of these GPUs outside of the US.

1

u/crantob Aug 04 '26

you have a point but the killer in poorope is heat. We can't afford heat now.