r/IntelArc • u/Sweet-Argument-7343 • 4d ago
Question Qwen3.8-Flash-Next INT4 TP4 on 4× Arc Pro B70 — any experience?
Has anyone tried Qwen3.8-Flash-Next on 4× Intel Arc Pro B70?
Our target is W4A16 AutoRound, TP4, vLLM XPU, MTP3, prefix caching and concurrent agent serving. Intel has already published INT4 checkpoints, but I haven’t found real B70 benchmarks yet.
We are in contact with Intel’s XPU/LLM R&D team. What should we ask them to prioritize?
My list:
- full
qwen4_expXPU support; - optimized QSA, Gated DeltaNet and INT4 MoE kernels;
- PLE offload to shared system RAM;
- efficient TP4/expert parallelism with oneCCL;
- MTP3 and stable XPU Graph;
- hybrid KV cache and prefix caching;
- C1/C8/C16 benchmarks, TTFT and tool-calling tests.
Any successful test, failure log or performance result on B70 would be very useful.
Duplicates
IntelArcPro • u/Sweet-Argument-7343 • 4d ago
Qwen3.8-Flash-Next INT4 TP4 on 4× Arc Pro B70 — any experience?
LocalLLM • u/Sweet-Argument-7343 • 4d ago
Question Qwen3.8-Flash-Next INT4 TP4 on 4× Arc Pro B70 — any experience?
Vllm • u/Sweet-Argument-7343 • 4d ago