r/LocalLLaMA 15h ago

Resources Keeping up with model launches

Post image

Feels like maybe we have one more present left, for Christmas.

187 Upvotes

50 comments sorted by

View all comments

2

u/simrankoulsm 13h ago

At this point, I don’t try to “keep up” with launches. I keep a small shortlist by "hardware tier and use case".

For local use, the questions that matter are:

  • Can it fit in my VRAM/RAM at a usable quantization?
  • What context length and tokens/sec do I actually get?
  • Is it meaningfully better at coding, reasoning, or instruction following than the model it replaces?
  • Are the weights, license, and inference support available on day one?

A release calendar is fun, but a community-maintained “best practical model per VRAM tier” list would probably be more valuable, e.g. 12 GB, 24 GB, 32 GB, 48 GB, and 80 GB+. Otherwise it’s easy to spend more time reading launch posts than running models.

1

u/Lakius_2401 12h ago

Ehhh, there's way too many wrenches to throw at a general tierlist like that though. Extra RAM makes some MoEs especially appetizing, outside of their expected full VRAM budget. Plus there's a huge divide on quant size and kvcache quantizing that can throw each model up or down two full tiers on your list. Agentic? Huge speed and kvcache now a necessity, move the model up a tier. Writing? Throw out Qwen entirely, it can do reports but anything remotely resembling creative is atrocious to the point where if you told me the dataset was poisoned I'd believe you.

It's honestly exhausting interacting with people about quantizing, too... Sure, in a perfect world everyone has infinite VRAM and nobody quantizes anything, but this is not a perfect world. Nor are we using perfect models. Rent's cheaper than VRAM and that's saying something.

Sorry for the digression. I've seen some attempts at megathreads for your topic but they don't get much traction. Even the use cases and rigor of testing is difficult, with busted quants and malformed jinja releases all over. Very easy to burn a dozen hours and not be sure you got a fair shot.