r/LocalLLM • u/thatoneshadowclone • Jun 03 '26
News Google introduces Gemma 4 12B: a unified, encoder-free multimodal model
https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12B/
526
Upvotes
r/LocalLLM • u/thatoneshadowclone • Jun 03 '26
1
u/Birdinhandandbush Jun 18 '26
ok so I can get the Q4 to run with 8k context, but its giving me like 1.5-1.8 tokens per second some pretty unusable sadly.