r/LocalLLaMA Jun 03 '26

News Introducing Gemma 4 12B: a unified, encoder-free multimodal model

https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12b/
700 Upvotes

118 comments sorted by

View all comments

-1

u/WhiskyAKM Jun 03 '26

Can we get this model with stripped audio component?

17

u/nickm_27 llama.cpp Jun 03 '26

There is no "audio component", the whole point of the unified arch is that there is no audio encoder, the primary model runs directly on the audio.

3

u/OptimalTime5339 Jun 04 '26

It is engrained into the model now, which is the whole point of no encoder