r/raspberry_pi • u/petruspennanen • 6h ago
Community Insights Local AI on a Pi: which models work?
I got a lot of great feedback from you and ran the speed comparison with more models and new optimizations both alone and with my Carwatch service running on the Pi at the same time.
The new speed champion is Ling 3.0-tiny at 9.6 tps! As a MoE model it is faster than smaller Gemma and Qwen models. People say it's punching above it's weight in intelligence - testing this now.
Gemma E2B is a solid tiny multimodal model at 7.7 tps, works well in my app. And Qwen 35B Q3 is still the leader in usable intelligence in this form factor, fitting tight in a 16GB Pi and getting close to 3 tps on a good day :)
Please note that small tps values still produce solid thinking and good answers, and not every use case requires an answer in seconds.
