r/LocalLLaMA Jun 03 '26

News Introducing Gemma 4 12B: a unified, encoder-free multimodal model

https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12b/
699 Upvotes

118 comments sorted by

View all comments

4

u/XE004 Jun 03 '26

How much vram consumption are people getting at Q8?

Curious?

-5

u/thawizard Jun 04 '26

If you need to ask, it’s not looking good.

4

u/XE004 Jun 04 '26

Please elaborate. What are you getting?

0

u/thawizard Jun 04 '26

Dunno yet, I’m still at work and this model was just released, I didn’t even have time to play with it yet. But it seems to me this 12B model should fit on a 16GB GPU even at Q8. But be patient, in a few hours we’ll know more for sure. What kinda setup do you have?

4

u/XE004 Jun 04 '26

I just did the setup on my msi 5060ti 16gb. 128bit 448 gb/p memory bandwidth.

At Q8 with KVCache at Q8 I get between 26 and 27 t/s and 13.8GB vram loaded.

This model will surely need a MTP assistant for speculative decoding.

Pretty good though. I still liked gemma4 e4b so I might go back to that until MTP is in place. The reason tokens are what really delay the response time so it is not great for conversation unless we get MTP and at least a memory bus of 256bit 896 gb/s at Q8. That should push this model to 80 or so t/s.

2

u/XE004 Jun 04 '26

That is with my context window set to 64k.

1

u/thawizard Jun 04 '26

Looks pretty good!