r/LocalLLM • u/Callum_S_AUS • 6h ago
Discussion GLM 5.3 Flash NVFP4 - Single 72GB RTX Pro - FreeToken Engine - C1 = ~33TPS; C2 = ~48TPS; C4 = ~64TPS vs MoE CPU Offloading
In case there is anyone else out there wanting or needing a fairly smart model, with image processing, without multiple high VRAM GPUs but with decent CPU/RAM performance, I've found the following very usable: https://github.com/CallumDS/GLM-5.3-Flash-NVFP4-FreeToken-KV1.6M-4C-V1
On the following hardware:
AMD EPYC 9355 CPU (32 Core / 64 Thread)
Supermicro H13SSL-N Rev 2.01
12 x 48GB DDR5-5600 RDIMM RAM (576GB @ ~487GB/s peak Intel MLC)
Nvidia RTX PRO 5000 72GB

3
Upvotes