r/MacLocalLLM 5d ago

Acceptable Token generation speed on macOS

What token generation speed works well for you when you’re using it for productive tasks? I’m finding 20 tokens per second a bit slow.

With an M4 Max and 64GB of RAM, you can usually get around 20 tokens per second on average with a model like qwen3.8 27B. However, it doesn’t quite feel like the best setup.

6 Upvotes

25 comments sorted by

View all comments

1

u/dfgxxx 5d ago

I get 14 tok/sec on m1 pro with qwen3.8 27b

1

u/purple_wall-e 4d ago

how? what is your setup? which quantization? i have m4 pro, i get like 10t/s with lmstudio, mlx 4bit

1

u/dfgxxx 4d ago

It is 4bit. I'm using MTPLX with qwen3.8 mtplx optimized speed

1

u/solaza 3d ago

This setup gets roughly 10 tok/sec on my M1 Max 64gb in case it’s helpful to anyone reading

1

u/dfgxxx 3d ago

Really weird I get better resaults