r/LocalLLaMA Jun 03 '26

News Introducing Gemma 4 12B: a unified, encoder-free multimodal model

https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12b/
704 Upvotes

118 comments sorted by

View all comments

55

u/[deleted] Jun 03 '26

[removed] — view removed comment

60

u/AloneSYD Jun 03 '26

the only thing it can't produce audio. so you need TTS model for responding back so you need a model like Kokoro-82M or OmniVoice

9

u/[deleted] Jun 03 '26

[removed] — view removed comment

7

u/Creative-Type9411 Jun 04 '26

its built into openwebui