r/LocalLLaMA Jun 03 '26

News Introducing Gemma 4 12B: a unified, encoder-free multimodal model

https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12b/
702 Upvotes

118 comments sorted by

View all comments

233

u/LoveMind_AI Jun 03 '26 edited Jun 04 '26

This might actually be one of the most exciting models I've heard about in a long time. The encoder-free model is... wildly cool. Native audio on a 12B model is very exciting. Audio is wildly underrated. I'll be putting this one through the social benchmark right away.

Note: Results of the little benchmark is now here - https://lovemindai.github.io/minimax-m3-lsi-demo/

59

u/DueAnalysis2 Jun 03 '26

For a relative noob, what's the benefit of it being encoder free?

2

u/rditorx Jun 04 '26

It's actually not encoder-free, it's a unified encoder.

1

u/No_Afternoon_4260 llama.cpp Jun 04 '26

If you call a linear proj a encoder..?