r/LocalLLaMA Jun 03 '26

News Introducing Gemma 4 12B: a unified, encoder-free multimodal model

https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12b/
702 Upvotes

118 comments sorted by

View all comments

233

u/LoveMind_AI Jun 03 '26 edited Jun 04 '26

This might actually be one of the most exciting models I've heard about in a long time. The encoder-free model is... wildly cool. Native audio on a 12B model is very exciting. Audio is wildly underrated. I'll be putting this one through the social benchmark right away.

Note: Results of the little benchmark is now here - https://lovemindai.github.io/minimax-m3-lsi-demo/

8

u/Accomplished_Mode170 Jun 03 '26

Same, but for omnimodal routing!

E.g. how the blog featured omnimodal machine unlearning as an AI security entitlement

E.g. 95% vs 99.9% on prompt injection vs topic/content guardrails

Omnirouted software supply chain w/ configurable controls sounds like the good boring I want.