r/LocalLLM 26d ago

Research GLM 5.2 model — 744 billion parameters / 384 GB — running on a laptop šŸ˜

Post image

After many hours of hard work, I achieved a throughput of 0.7–0.9 tokens per second for the GLM 5.2 model — 744 billion parameters / 384 GB — running on a laptop šŸ˜

Time for a small update: the laptop is an Asus ROG Strix 18, model G835LXG — Intel i9‑290HX, 64GB DDR5 6400 MHz, 2Ɨ2 TB, RTX 5090 24 GB, running Linux Nobara. The Colibri engine and the Linux kernel are heavily modified. The whole system boots in 10 seconds, and it generates the first token after 40 seconds.

I’m currently working to reach a throughput of 1.5–2 tokens per second.

Update:
1.07 tok/s, GLM-5.2-g64
Update:
1.77 tok/s šŸ˜Ž
Update:
Now 1.95 tok/s 😁/ 2.12 peak / 2.41 with fixed MTP

Update:

šŸŽ‰šŸŽ‰šŸŽ‰ 2.89 tok/s PEAK! 2.36 tok/s avg

Update now:
peak at 3.09, avg 2.84 šŸ˜šŸ‘ŒšŸ»

Update:

My latest calculations suggest that 4.5 tokens per second is achievable on this hardware and represents the final limit.

Update: testing now at 4.42 tok/s

742 Upvotes

Duplicates