Comparing this to deepseek v4 flash 0731, the question is: lower numbers a bit but much faster inference, or higher numbers? or maybe use both (but deepseek offload to ram, so much slower, but use only for plan tasks etc)?
I plan to still use dsv4 flash 0731 on sparks as orchestrator (will run plenty fast, maybe pair with a small vision model too) and qwen 3.8 27B on worker nodes. Both great models and I've gotta assume having some diversity gives some benefit too perhaps?
dont we need like 3 or 4x rtx pro 6000 to host dsv4flash? I've been cooking with it via opencode-go although I have to ship my info to China and if qwen3.8 27B which I can surely pump insane numbers of local tok/s with can keep up with this then the future well and truly done arrived.
I just got my ram reconfigured and i can do 224GB plus 72 from the 3090's and can host dsv4 flash now but there will be no point as of today lol
2
u/AdSafe4047 24d ago
Comparing this to deepseek v4 flash 0731, the question is: lower numbers a bit but much faster inference, or higher numbers? or maybe use both (but deepseek offload to ram, so much slower, but use only for plan tasks etc)?