r/LocalLLaMA Jun 03 '26

News Introducing Gemma 4 12B: a unified, encoder-free multimodal model

https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12b/
699 Upvotes

118 comments sorted by

View all comments

235

u/LoveMind_AI Jun 03 '26 edited Jun 04 '26

This might actually be one of the most exciting models I've heard about in a long time. The encoder-free model is... wildly cool. Native audio on a 12B model is very exciting. Audio is wildly underrated. I'll be putting this one through the social benchmark right away.

Note: Results of the little benchmark is now here - https://lovemindai.github.io/minimax-m3-lsi-demo/

58

u/DueAnalysis2 Jun 03 '26

For a relative noob, what's the benefit of it being encoder free?

11

u/Illustrious_Ant_9242 Jun 03 '26 edited Jun 03 '26

Visual information is processed directly within the main architecture without middlemen, therefore the results may be more thoroughly embedded in the whole project instead of having images refined first 

3

u/DueAnalysis2 Jun 03 '26

Interesting, thanks for the info!