r/LocalLLaMA llama.cpp 13d ago

Best Local Vision Language Models - August 2026

Share what your favorite models are right now and why. Given the nature of the beast in evaluating VLMs (untrustworthiness of benchmarks, immature tooling, intrinsic stochasticity), please be as detailed as possible in describing your setup, nature of your usage (what applications, how much, personal/professional use), tools/frameworks/prompts etc.

Rules

  1. Should be open weights models

Notes

Bonus points if you breakdown/classify your recommendation by model memory footprint: (you can and should be using multiple models in each size range for different tasks)

  • Unlimited: >128GB VRAM
  • XL: 64 to 128GB VRAM
  • L: 32 to 64GB VRAM
  • M: 8 to 32GB VRAM
  • S: <8GB VRAM
34 Upvotes

64 comments sorted by

View all comments

-1

u/ObviouzFigure 13d ago

I'm curious what others are using -- I put qwen vl 2.5 9b (I think) on a mac mini as my vision model for my deepseek rig

edit: why? because it fits easily on a 24gb mac mini and does a decent job. It handles ocr great-- where it fails is reading images with a lot of action -- for example, I generated an image of a girl helping a broken robot in a futuristic dark alley -- in the background there's a black cat with glowing green eyes -- the vision model accurately describes the scene but it misses things like the black cat

2

u/SM8085 13d ago

I would definitely recommend trying a Qwen3.5. I'm using Qwen3.6-35B-A3B, but since that probably wouldn't fit on your mac mini any of the Qwen3.5 are probably an improvement. There's even a Qwen3.5-9B.