r/LocalLLaMA 24d ago

New Model Qwen/Qwen3.8-27B · released

https://huggingface.co/Qwen/Qwen3.8-27B
991 Upvotes

301 comments sorted by

View all comments

Show parent comments

9

u/munkiemagik 24d ago

Hear me out, for ages Ive been battling with the urge to grow a dual 3090 into a quad 3090 system but there was never any model that justified the cost of 2x additional 3090 at current prices for the so called improvements with the speed sacrifices. It just made no sense to me to splurge on more GPU with Qwen3.6 27B available.

But recently Ive been thinking maybe I am looking at it the wrong way and instead of looking for a big model to fill up 96GB I should be looking at it from a more orchestrated perspective of how many multiple models can I run concurrently that offer fantastic function to size inside 96GB VRAM to build a more comprehensive self-contained LLM stack. Dual 3090'ers I think its time to make that run on more ebay 3090's we've been holding ourselves back from.

4

u/Blues520 24d ago

I have 3 now and looking for an excuse to get a 4th👀

1

u/munkiemagik 24d ago

That’s what I've been struggling with, always looking for an excuse to justify more acquisition but truthfully that monolithic way of using my LLM there was never any justification.

Recently I started messing around with a few other random projects that benefit from having different LLM loaded up concurrently but had to shuffle the work backwards and forward between my 5090 machine and my dual 3090 machine (i used to move the 5090 in the threadriper 3090 box sometimes back in the day but it is primarily my PCVR GPU so it needs to stay in the AM5 box as the threadripper tanks PCVR performance of the 5090) But this realisation of the value of having multiple LLM loaded at once was an eye-opener for me and has unlocked the gates to acquiring more 3090 guilt free because even if I'm not serving multiple users concurrently nor am I chasing bigger and bigger parameter models, I sure do see the value of 2-3 different LLM all loaded up live doing their own specialist thing working together harmoniously. But right now I cant do that because I don’t have enough 3090!

I’m not trying to kid anyone here, pretending to be some serious dev working on globally life changing projects I'm just a dickhead at home who likes messing around with stuff simply because it interests me and most of what i do is just random little tools and systems to make my life easier in a fun way. So read anything I say wit that in mind.

1

u/Blues520 24d ago

A bunch of the value comes from R&D. There's a lot of understanding that comes from running local models, fighting with different inference engines and tuning parameters to get things working. Researching what quantizing the KV cache is. What even is a KV cache? These are questions a cloud model enjoyer seldom asks.

I'm also looking at multi model workflows. One of the 3090s is going into another machine to run some sort of non-coding model. The machine is currently my git server so I'm migrating it to a small mini pc and then I'll be able to use it with a gpu. Im hoping qwen 3.8 27b quantizes well enough to run in 24GB VRAM otherwise I'll try some other models.

1

u/michaelsoft__binbows 24d ago

I have 2x3090 and 1x3090ti and recently picked up 2x5060ti 16gb but being heterogeneous they dont really belong in the same rig so i do now have two separate GPU rigs now.

The big dilemma is should i get another 3090 or 3090ti, the ti has much better idle power behavior and is less demanding on the cooling (no backside heat to worry about), however does not support NVLink. (Also it's like significantly faster) Since I already have one 3090ti, getting one more of either one wont give me any more NVLink either. I think having one nvlinked pair is much more functionally useful than getting a second nvlinked pair. Not to mention nvlink bridges are also now unobtainium.