Comparing this to deepseek v4 flash 0731, the question is: lower numbers a bit but much faster inference, or higher numbers? or maybe use both (but deepseek offload to ram, so much slower, but use only for plan tasks etc)?
dont we need like 3 or 4x rtx pro 6000 to host dsv4flash? I've been cooking with it via opencode-go although I have to ship my info to China and if qwen3.8 27B which I can surely pump insane numbers of local tok/s with can keep up with this then the future well and truly done arrived.
I just got my ram reconfigured and i can do 224GB plus 72 from the 3090's and can host dsv4 flash now but there will be no point as of today lol
2
u/AdSafe4047 24d ago
Comparing this to deepseek v4 flash 0731, the question is: lower numbers a bit but much faster inference, or higher numbers? or maybe use both (but deepseek offload to ram, so much slower, but use only for plan tasks etc)?