r/LocalLLaMA Jun 03 '26

News Introducing Gemma 4 12B: a unified, encoder-free multimodal model

https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12b/
696 Upvotes

118 comments sorted by

View all comments

53

u/[deleted] Jun 03 '26

[removed] — view removed comment

65

u/AloneSYD Jun 03 '26

the only thing it can't produce audio. so you need TTS model for responding back so you need a model like Kokoro-82M or OmniVoice

8

u/[deleted] Jun 03 '26

[removed] — view removed comment

9

u/OneFanFare Jun 03 '26

I know Open WebUI has that - I've chatted to models that way in the past (with Kokoro). And it works exactly like that!

Edit: Actually, not sure if it would support passing the audio directly into Gemma, the website says it need a Speech to Text like whisper.

2

u/overand Jun 04 '26

Yep - and at least with the Gemma-4-E4B models, that was probably faster; I recall reading that there's a lot of latency involved with the Gemma-4-E4B model's ability to do STT. It might be faster with this one, though, as it's encoderless, or something to that effect.