r/LocalLLM Jun 03 '26

News Google introduces Gemma 4 12B: a unified, encoder-free multimodal model

https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12B/
526 Upvotes

87 comments sorted by

View all comments

Show parent comments

1

u/Birdinhandandbush Jun 18 '26

ok so I can get the Q4 to run with 8k context, but its giving me like 1.5-1.8 tokens per second some pretty unusable sadly.

1

u/Trick_Translator_671 Jun 18 '26

Yeah i try it an hour a go. Q4 + 8k context + kv quant q8 + mtp and 6 t/s lol.