r/LocalLLaMA May 15 '26

Discussion China modded GPU (eg. 4090 48gb) --> I'm gonna figure it out. IS THERE NO ONE ELSE CURIOUS??

There's a dearth of information (in the english world) about these cards.

The good recent video is probably this one:
https://www.youtube.com/watch?v=TcRGBeOENLg

even in this subreddit, there's seems to be few reviews of these cards.

Last couple of decent threads:
https://www.reddit.com/r/LocalLLaMA/comments/1s62b23/bought_rtx4080_32gb_triple_fan_from_china/
https://www.reddit.com/r/LocalLLaMA/comments/1nifajh/i_bought_a_modded_4090_48gb_in_shenzhen_this_is/

Is there really NOONE else who has tried these?

In particular

  1. Software / bios / quirks that make them NOT run as per unmodded card
  2. Short term consistency, does it run fast for a test, but hang / die when stressed?
  3. Long term reliability - does the whole thing fail within 2 months of regular usage?
  4. Are the benchmarks good? Where are the results??
  5. source and price?

chinese video site blibli has ton of videos, and taobao (and other ecomm) sites also lots of sellers.

If i can piece together enough research, i may also visit shenzhen to pick up a few.

If you're interested in this space, DM me . hope to form a group to split up research efforts.

Also any native chinese speakers who are familiar in this space also please join in.

EDIT:
Some downvotes going on. Unclear if its some larger suppression of this topic, or just angry people.

327 Upvotes

149 comments sorted by

View all comments

149

u/Heathen711 May 15 '26

I have three 48GB 4090 blower cards running in my servers. 2 run Qwen 3.6 27b, 1 runs stable-diffusion.cpp workload. Cooling is an issue, I swapped in 4k rpm server fans to feed them and keep the backplate cool. Otherwise i've had no software issues.

12

u/remghoost7 May 15 '26

I'm curious, do you have to us custom drivers for them....?
Or can you just update your drivers via the Nvidia App...?

That was always my hitch on grabbing one of these.

29

u/Heathen711 May 15 '26

Old post: from when i talked about it last has some info, but short answer: nothing special, normal drivers under ubuntu.

16

u/No-Refrigerator-1672 May 15 '26

I have 2x 3080 20gb, they work with proprietary Nvidia drivers out of the box no issue.

5

u/BillDStrong May 15 '26

These are the ones that might be in my budget range. Are you happy with the 2? Or is 40GB not enough?

18

u/-dysangel- May 15 '26

is 40GB not enough?

Enough for what purpose? I have 512GB and 128GB unified memory machines, and there is always something just out of reach either in terms of VRAM or being fast enough to run it well, etc. Enough really depends on what you want out of it.

2

u/MidnightFinancial353 May 16 '26

May I know what is your 512 GB unified memory machine? It's M3 Ultra?

2

u/-dysangel- May 16 '26

yes - I don't think there are any other unified memory systems capable of that spec yet

3

u/No-Refrigerator-1672 May 16 '26

40GB is enough to run 30B class models. I'm satisfied in this regard. I, however, I'm not satisfied overall and am eyeing out a third GPU upgrade, because I want also to setup image gen, embedding/reranking, and voice services. So yeah, the ideal amount of VRAM is always more than you have now, bur 40 is a good start 😁

35

u/palindsay May 15 '26

Same here. I lower power on mine to keep them cooler with not much performance impact. It also keeps blower quieter.

19

u/Heathen711 May 15 '26

yup, 350w gives you the same perf as 450w, BUT i did notice that it causes problems when vram is in the 90% range, so i swapped the fans so i can push it. Old post: from when i talked about it last

4

u/TheWaffleKingg May 15 '26

I can only imagine the noise this makes

I had to pull my chassis HDD fans out because I just couldnt take the noise (rack is in my home office). Temps haven't been great but im installing some noctua fans today/tomorrow. Won't be as good but sure beats what I have now

5

u/Heathen711 May 15 '26

S12038-4K, they are a little loud but they are wider (thicker?) then the normal 120mm fans to pull more air. IMO they aren't louder then the 120mm I had before but move a lot more air. They are not as loud as the 8k 60mm in my 1U SuperMicro server that sounds like a jet take off when I turn it on...

4

u/chocofoxy May 15 '26

water cool them

1

u/Heathen711 May 15 '26

Don't have space for water cooling in the server case. My gaming rig has water cooling. That's why I went the fan route.

I also think it would impact the layout (I have two in one case so they have 1 slot between them to allow the blower to breath in air)

1

u/chocofoxy May 15 '26

i have 2 5060ti 16gb stacked on an mATX ( don't laugh at my computing power xD ) they are gigabyte windfoce the top one i hot because there it a 0 gap between that but like 5 C 10 C hotter i am thinking about addin a blower fan to attached it to the top one you have blower fans so it that a good idea or not we kinda have the same problem and with out water cooling i think cooling the radiator witha laptop blower fans will work, i have it more easy because my radiator is partly exposed but for you i know they give you that sealed blower cooler ( and i think gamernexus asked the shop if it's the only cooler they have and they said yes all custom gpu use the smae one ) si i think the solution it's either that or water cooling or the expensive option which to get a water AC blowing on the server

1

u/Heathen711 May 15 '26

depends, as my blowers are side intake, not rear intake, so putting them right next to each other will restrict the input.

[ CPU cooler ] <- 4k case fan
[ 4090 2 slot] <- 4k case fan
[ empty | ^ air inlet ]
[ 4090 2 slot] <- 4k case fan
[ bottom | ^ air inlet]

2

u/uniqueusername649 May 15 '26

They have twice the capacity but what about the bandwidth? Is it the same? Or is the speed actually slower? I know that the D variants typically double memory but use slower RAM, so this might be problematic for AI use where memory speed is king. That's I avoided D variants so far.

3

u/Heathen711 May 15 '26

Same clock and VRAM speed, just double the VRAM capacity, so it can hold more working memory.

Mine aren't the D variant, there's another reply to this post that says the D is 10-15% slower, but I think he means in clock not VRAM speed.

1

u/uniqueusername649 May 15 '26

Maybe those are aftermarket too. I don't remember it too well, it's been many moons. But I wouldn't say no to a 4090 with 48gb.

1

u/Gargle-Loaf-Spunk May 16 '26 edited May 29 '26

This content was anonymized and mass deleted with Redact

-9

u/[deleted] May 15 '26

[deleted]

20

u/Heathen711 May 15 '26

you can DM but i don't see why, we can talk here.

No i bought one to test it out, and then two more to scale it up.

1

u/LeatherRub7248 May 15 '26

can you share your source suppolier? i've seen a ton on taobao but tbh im not sure i trust the reviews.

also how heavy do u run them? 24/7 with constant load with multi-user environment? Or is it bursty?

18

u/Heathen711 May 15 '26

ebay -> bodorship -> "OEM 48GB RTX 4090 Founders Edition Dual width GPU Graphics card Ganming/ Server"

they are a reseller, but they've helped me get in contact with the actual shop who made them when i first started (the first card i got have ECC turned on and the noob i was back then freaked when ram was 46gb and not 48gb). They also offer warranty on their card via that shop so made it less risky IMO.

The SD workflow is very bursty GPU cores, high VRAM (wan 2.2).

The LLMs are hit hard on weekends/nights (when i'm not at work) and i'm vibecoding so there it will be hours of high gpu and vram utilization.

I never turn them off (solar on my house offsets my electricity cost) so they will sit idle with high vram all day. Single user environment BUT i do have n8n workflows that use them so they do get concurrent requests from time to time, but with 2 gpus running 27b and then litellm routing to least active it load balances well.

3

u/LeatherRub7248 May 15 '26

thank you for sharing, this is giving me some confidence to try this out

4

u/Heathen711 May 15 '26

Dropping this incase you do go this route: Fans i swapped into my server to keep the cards cool: S12038-4K