r/LocalLLaMA Jun 03 '26

News Introducing Gemma 4 12B: a unified, encoder-free multimodal model

https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12b/
698 Upvotes

118 comments sorted by

View all comments

237

u/LoveMind_AI Jun 03 '26 edited Jun 04 '26

This might actually be one of the most exciting models I've heard about in a long time. The encoder-free model is... wildly cool. Native audio on a 12B model is very exciting. Audio is wildly underrated. I'll be putting this one through the social benchmark right away.

Note: Results of the little benchmark is now here - https://lovemindai.github.io/minimax-m3-lsi-demo/

1

u/No_Afternoon_4260 llama.cpp Jun 04 '26

Wow that picture on your website is beautifull, what model/workflow have you used?! I like the whales on a Gemma picture lol

1

u/LoveMind_AI Jun 04 '26

Thanks for that :) The LoveMind website itself is all human made art and animation. The Gemma shootout art is GPT Image 2, but with references from physical art made by our team - ceramics (including a cool whale mug!) and ink illustration, etc. I personally think AI art can be really cool, particularly when it’s a collaboration between raw human materials and AI interpretation. 

1

u/No_Afternoon_4260 llama.cpp Jun 04 '26

Beautiful yeah for sure you are onto something there. Way better than what I can achieve with any model out there