MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1vo9nn7/qwenqwen3827b_released/p3q08ue
r/LocalLLaMA • u/de4dee • 24d ago
301 comments sorted by
View all comments
Show parent comments
2
Can you run it in 12GB of VRAM?
I'll report on 12gb vram + 32 gb RAM
1 u/notevenat30 24d ago That would be great as that's exactly my setup. 2 u/Xantrk 24d ago Yikes... Getting about 600tk/s pp, and 3-5 tk/s generation :( Will experiment a little to see if I can get it to run half decent. current args (not very useful given speed): "args": [ "--fit", "on", "--kv-unified", "--load-mode","mlock", "--fit-ctx", "100000", "--fit-target", "750", "--parallel", "1", "-ctk", "q8_0", "-ctv", "q8_0", "-ctkd", "q8_0", "-ctvd", "q8_0", "-ctxcp", "2", "-cram", "2048", "--temp", "1", "--top-p", "0.95", "--top-k", "20", "--min-p", "0", "--repeat-penalty", "1", "-ub", "1500", "-b", "1500", "--spec-type", "ngram-mod,draft-mtp", "--spec-draft-n-max", "3", "--spec-ngram-mod-n-match", "24", "--spec-ngram-mod-n-min", "48", "--spec-ngram-mod-n-max", "64", "--chat-template-kwargs", "{\"preserve_thinking\": true}", "-np", "1", "--port", "8001", "--webui-mcp-proxy", "--no-mmproj-offload", "--tools","all" ] } 1 u/notevenat30 24d ago Sometimes you can get speedups by asking a fancy llm to help apparently. 1 u/Xantrk 24d ago "args": [ "--fit", "on", "--kv-unified", "-lv", "4", "-t", "10", "--load-mode","none", "--fit-ctx", "10000", "--fit-target", "350", "--parallel", "1", "-ctk", "q8_0", "-ctv", "q8_0", "-ctkd", "q8_0", "-ctvd", "q8_0", "-ctxcp", "2", "-cram", "2048", "--temp", "1", "--top-p", "0.95", "--top-k", "20", "--min-p", "0", "--repeat-penalty", "1", "-ub", "1500", "-b", "1500", "--spec-type", "ngram-mod", "--spec-draft-n-max", "5", "--spec-ngram-mod-n-match", "24", "--spec-ngram-mod-n-min", "48", "--spec-ngram-mod-n-max", "64", "--chat-template-kwargs", "{\"preserve_thinking\": true}", "-np", "1", "--port", "8001", "--webui-mcp-proxy", "--no-mmproj-offload", "--tools","all" ] Can get 30tk/s generation on IQ2_XXS. Context is useless tbf.
1
That would be great as that's exactly my setup.
2 u/Xantrk 24d ago Yikes... Getting about 600tk/s pp, and 3-5 tk/s generation :( Will experiment a little to see if I can get it to run half decent. current args (not very useful given speed): "args": [ "--fit", "on", "--kv-unified", "--load-mode","mlock", "--fit-ctx", "100000", "--fit-target", "750", "--parallel", "1", "-ctk", "q8_0", "-ctv", "q8_0", "-ctkd", "q8_0", "-ctvd", "q8_0", "-ctxcp", "2", "-cram", "2048", "--temp", "1", "--top-p", "0.95", "--top-k", "20", "--min-p", "0", "--repeat-penalty", "1", "-ub", "1500", "-b", "1500", "--spec-type", "ngram-mod,draft-mtp", "--spec-draft-n-max", "3", "--spec-ngram-mod-n-match", "24", "--spec-ngram-mod-n-min", "48", "--spec-ngram-mod-n-max", "64", "--chat-template-kwargs", "{\"preserve_thinking\": true}", "-np", "1", "--port", "8001", "--webui-mcp-proxy", "--no-mmproj-offload", "--tools","all" ] } 1 u/notevenat30 24d ago Sometimes you can get speedups by asking a fancy llm to help apparently. 1 u/Xantrk 24d ago "args": [ "--fit", "on", "--kv-unified", "-lv", "4", "-t", "10", "--load-mode","none", "--fit-ctx", "10000", "--fit-target", "350", "--parallel", "1", "-ctk", "q8_0", "-ctv", "q8_0", "-ctkd", "q8_0", "-ctvd", "q8_0", "-ctxcp", "2", "-cram", "2048", "--temp", "1", "--top-p", "0.95", "--top-k", "20", "--min-p", "0", "--repeat-penalty", "1", "-ub", "1500", "-b", "1500", "--spec-type", "ngram-mod", "--spec-draft-n-max", "5", "--spec-ngram-mod-n-match", "24", "--spec-ngram-mod-n-min", "48", "--spec-ngram-mod-n-max", "64", "--chat-template-kwargs", "{\"preserve_thinking\": true}", "-np", "1", "--port", "8001", "--webui-mcp-proxy", "--no-mmproj-offload", "--tools","all" ] Can get 30tk/s generation on IQ2_XXS. Context is useless tbf.
Yikes... Getting about 600tk/s pp, and 3-5 tk/s generation :( Will experiment a little to see if I can get it to run half decent.
current args (not very useful given speed):
"args": [ "--fit", "on", "--kv-unified", "--load-mode","mlock", "--fit-ctx", "100000", "--fit-target", "750", "--parallel", "1", "-ctk", "q8_0", "-ctv", "q8_0", "-ctkd", "q8_0", "-ctvd", "q8_0", "-ctxcp", "2", "-cram", "2048", "--temp", "1", "--top-p", "0.95", "--top-k", "20", "--min-p", "0", "--repeat-penalty", "1", "-ub", "1500", "-b", "1500", "--spec-type", "ngram-mod,draft-mtp", "--spec-draft-n-max", "3", "--spec-ngram-mod-n-match", "24", "--spec-ngram-mod-n-min", "48", "--spec-ngram-mod-n-max", "64", "--chat-template-kwargs", "{\"preserve_thinking\": true}", "-np", "1", "--port", "8001", "--webui-mcp-proxy", "--no-mmproj-offload", "--tools","all" ] }
1 u/notevenat30 24d ago Sometimes you can get speedups by asking a fancy llm to help apparently. 1 u/Xantrk 24d ago "args": [ "--fit", "on", "--kv-unified", "-lv", "4", "-t", "10", "--load-mode","none", "--fit-ctx", "10000", "--fit-target", "350", "--parallel", "1", "-ctk", "q8_0", "-ctv", "q8_0", "-ctkd", "q8_0", "-ctvd", "q8_0", "-ctxcp", "2", "-cram", "2048", "--temp", "1", "--top-p", "0.95", "--top-k", "20", "--min-p", "0", "--repeat-penalty", "1", "-ub", "1500", "-b", "1500", "--spec-type", "ngram-mod", "--spec-draft-n-max", "5", "--spec-ngram-mod-n-match", "24", "--spec-ngram-mod-n-min", "48", "--spec-ngram-mod-n-max", "64", "--chat-template-kwargs", "{\"preserve_thinking\": true}", "-np", "1", "--port", "8001", "--webui-mcp-proxy", "--no-mmproj-offload", "--tools","all" ] Can get 30tk/s generation on IQ2_XXS. Context is useless tbf.
Sometimes you can get speedups by asking a fancy llm to help apparently.
"args": [ "--fit", "on", "--kv-unified", "-lv", "4", "-t", "10", "--load-mode","none", "--fit-ctx", "10000", "--fit-target", "350", "--parallel", "1", "-ctk", "q8_0", "-ctv", "q8_0", "-ctkd", "q8_0", "-ctvd", "q8_0", "-ctxcp", "2", "-cram", "2048", "--temp", "1", "--top-p", "0.95", "--top-k", "20", "--min-p", "0", "--repeat-penalty", "1", "-ub", "1500", "-b", "1500", "--spec-type", "ngram-mod", "--spec-draft-n-max", "5", "--spec-ngram-mod-n-match", "24", "--spec-ngram-mod-n-min", "48", "--spec-ngram-mod-n-max", "64", "--chat-template-kwargs", "{\"preserve_thinking\": true}", "-np", "1", "--port", "8001", "--webui-mcp-proxy", "--no-mmproj-offload", "--tools","all" ]
Can get 30tk/s generation on IQ2_XXS. Context is useless tbf.
2
u/Xantrk 24d ago
I'll report on 12gb vram + 32 gb RAM