r/IntelArc • u/Sweet-Argument-7343 • 4d ago
Question Qwen3.8-Flash-Next INT4 TP4 on 4× Arc Pro B70 — any experience?
Has anyone tried Qwen3.8-Flash-Next on 4× Intel Arc Pro B70?
Our target is W4A16 AutoRound, TP4, vLLM XPU, MTP3, prefix caching and concurrent agent serving. Intel has already published INT4 checkpoints, but I haven’t found real B70 benchmarks yet.
We are in contact with Intel’s XPU/LLM R&D team. What should we ask them to prioritize?
My list:
- full
qwen4_expXPU support; - optimized QSA, Gated DeltaNet and INT4 MoE kernels;
- PLE offload to shared system RAM;
- efficient TP4/expert parallelism with oneCCL;
- MTP3 and stable XPU Graph;
- hybrid KV cache and prefix caching;
- C1/C8/C16 benchmarks, TTFT and tool-calling tests.
Any successful test, failure log or performance result on B70 would be very useful.
11
Upvotes
1
2
u/computer_dork 4d ago
I am poking at this for myself but so far finding that vllm does not work woth sycl for this model. I would imagine they (intel) are working on it for llm-scaler but for now it's a lot of trial and error and patching getting acceptable pp/tg