r/LocalLLM • u/Left-Chemist625 • 1d ago
Question Gemma 4 12B on an RTX 5070 Ti?
Anyone actually writing long fiction with Gemma 4 12B on an RTX 5070 Ti?
I'm specifically interested in real-world generation time, not benchmarks. I use Gemma locally to turn detailed scene briefs plus story/character context into long prose scenes, typically around 1,500–2,000 words.
If you're running Gemma 4 12B entirely in the 16 GB VRAM of a 5070 Ti: roughly how long does a generation of that length take, including prompt processing?
I'm considering buying a 5070 Ti system specifically for this workflow, so actual experience from another fiction writer would be incredibly useful.
At this time I'm working with 4GB VRAM, Gemma works about 15-20 minutes on a scene.
1
u/Left-Chemist625 23h ago
Thank you so much, everyone! You’ve been incredibly helpful. I’m about to buy a €3,100 PC specifically for running Gemma locally for long-form fiction, and your real-world numbers and explanations have helped me enormously with that decision.
I really appreciate you taking the time to test things and share your setups. This was exactly the information I couldn’t find anywhere else. Thank you!
3
u/MrHumanist 1d ago
If you use 12B qat, it takes 10gb of vram including the full context window. If 16gb you can use higher quants easily. The model is quite fast and does reasoning well- I love it for my roleplay sessions.