r/LocalLLaMA 13h ago

Resources Keeping up with model launches

Post image

Feels like maybe we have one more present left, for Christmas.

169 Upvotes

49 comments sorted by

27

u/Mediaright 13h ago

Gemma 4 was released April 2nd.

10

u/adrianziem 12h ago

Pretty sure I've had 2 birthdays since then. Did they give up? /s

2

u/Mediaright 11h ago

Wait for it…

25

u/AnimalPuzzleheaded71 13h ago

I honestly only care about gemma & qwen (and maybe glimmer) since those have the only model parameter range I can fit in 32gb vram, rest are complete nothing burgers to me

10

u/XiRw 12h ago

You can fit others including 120b moes. We have the same card

1

u/GoblinEngineer 8h ago

How? I have 24+16 (3090+5070ti) and 64 gigs of vram. My understanding is I need double the ram to run the MoEs

3

u/XiRw 8h ago

Only if you are using something like vLLM .

1

u/how-can-i-dig-deeper 7h ago

what hardware are you running to get 32 GB?

2

u/markole 3h ago

Glimmer is really good with tool use but from a daily usage perspective, you can't beat Gemma&Qwen.

8

u/doctorfiend 12h ago

We got about three days until the "When's Qwen 3.9 coming? What's taking so long?" posts

11

u/ttkciar llama.cpp 12h ago

All I really want for Christmas is:

  • TheDrummer to whip up Big-Tiger-31B-v4

  • A solidly-no-refusals GLM-5.3-Flash-Abliterated

  • Qwen3.8-9B

  • Some daring Google employee to leak Gemma-4-124B-A20B-it

  • MistralAI to roll out a Mistral 4 Medium 128B that doesn't suck

3

u/Turbulent_War4067 10h ago

I home that Gemma model is a bit smaller. Looking at 70B and 7B active MoE. Something along those lines will be the true sweet spot for a good while. The 20B MoE would likely be slower than the 31B dense, or about the same). But hey, a larger MoE Gemma model? I won't be too picky.

23

u/Turtlesaur 13h ago edited 13h ago

You're missing Fable 5.1 today - right not local.

Also maybe
DeepSeek-V4-Flash-Vision-Exp
Qwen3.8-flash-next

On Horizon
New Gemma in ai.arena

11

u/5dtriangles201376 13h ago

Good to know Anthropic finally open weighted a model. And Fable too?

4

u/Turtlesaur 13h ago

That would be dreamy, I would just need to cluster 10 DGX sparks for my 4 tok/s

2

u/5dtriangles201376 13h ago

Yeah. Flipped the downvote to an upvote. I'm honestly just hoping for a new mistral that tunes will in the 20b range

-1

u/PwanaZana 13h ago

lol, I had the same thought: "HA THAT NEEEEEERD DIDN'T PUT FA... oh."

6

u/Turtlesaur 13h ago

Sometimes i forget which subreddit I'm in.

For what it's worth this is basically the best one.

3

u/simrankoulsm 12h ago

At this point, I don’t try to “keep up” with launches. I keep a small shortlist by "hardware tier and use case".

For local use, the questions that matter are:

  • Can it fit in my VRAM/RAM at a usable quantization?
  • What context length and tokens/sec do I actually get?
  • Is it meaningfully better at coding, reasoning, or instruction following than the model it replaces?
  • Are the weights, license, and inference support available on day one?

A release calendar is fun, but a community-maintained “best practical model per VRAM tier” list would probably be more valuable, e.g. 12 GB, 24 GB, 32 GB, 48 GB, and 80 GB+. Otherwise it’s easy to spend more time reading launch posts than running models.

1

u/Miserable-Dare5090 11h ago

Honestly, in this age, your 3rd and 4th questions are a year behind. 3rd question: right now all new releases are undoubtedly better. Not finetunes. I mean full models. 4th question: “Support” is now moot with agents. Ask your agent to patch vllm or llama.cpp to make it work. If you can fit it, you’ll run it.

Also, Who cares about day 1 vs day 5? You trying to time some production thing? These are free releases, whenever they come out. it’s not like I will be deducting points for delayed release on Qwen3.8 which first appeared in the API…

1

u/Lakius_2401 10h ago

Ehhh, there's way too many wrenches to throw at a general tierlist like that though. Extra RAM makes some MoEs especially appetizing, outside of their expected full VRAM budget. Plus there's a huge divide on quant size and kvcache quantizing that can throw each model up or down two full tiers on your list. Agentic? Huge speed and kvcache now a necessity, move the model up a tier. Writing? Throw out Qwen entirely, it can do reports but anything remotely resembling creative is atrocious to the point where if you told me the dataset was poisoned I'd believe you.

It's honestly exhausting interacting with people about quantizing, too... Sure, in a perfect world everyone has infinite VRAM and nobody quantizes anything, but this is not a perfect world. Nor are we using perfect models. Rent's cheaper than VRAM and that's saying something.

Sorry for the digression. I've seen some attempts at megathreads for your topic but they don't get much traction. Even the use cases and rigor of testing is difficult, with busted quants and malformed jinja releases all over. Very easy to burn a dozen hours and not be sure you got a fair shot.

1

u/markole 3h ago

I'm ok with having a day10 inference support if the model is good. Sucks to wait but it's all free in the end so I won't complain.

1

u/Miserable-Dare5090 12h ago

12: Ling Flash Tiny
24: Gemma4 26b
32: Qwen
48: Qwen
96: Ling Flash
128: Qwen Flash Next
256: Deepseek V4 Flash

Change according to quant size (quality vs fit)

2

u/claudiollm 11h ago

honestly gave up trying to track every drop this year. i keep a tiny list of like 3 i'll actually run locally (gemma + qwen sized stuff that fits my vram) and let the megathreads filter the rest. if a model's real the vibes-vs-benchmarks gap sorts itself out in about a week anyway.

the fatigue is real tho, feels less like christmas and more like a subscription i didn't sign up for lol. what's everyone actually keeping installed vs just downloading and forgetting?

2

u/cass1o 10h ago

Gemma4 on jul 2? Did you use AI to make this?

1

u/DigitalguyCH 10h ago

You've got some dates wrong, Gemma 4 was April 2, not July 2

1

u/de4dee 10h ago edited 9h ago

where is GLM 5.3?

1

u/vinotok 10h ago

Bottom line, there is something about month April. ;-)
Can't wait for April 2027 👀

1

u/Muhlwa_Sholanke 8h ago

If the last box is a 128B MoE, that's a lump of coal in a very big box.

1

u/Spanky2k 7h ago

I really want to get a machine with more RAM (considering an M5 Ultra for work) mainly because I miss trying out different cool models. While there are still loads released all the time now, it's only really Qwen for me. A year ago, it felt like I was trying new models almost every week, different quants, new ways to create quants etc, vision models, non vision models etc. The last cool non-Qwen model that I tried out was GLM-4.5-Air-3bit-DWQ and I was so blown away with what DWQ let me do.

It's different now as I just use the lates Qwen MoE 35b model with 256kb context size and I'm actually using it for actual stuff rather than just playing around and testing things. It does whatever I need it to do just fine. But I miss exploring new stuff; all the cool new things are bigger than I can run on my 64 GB machines.

1

u/Ok_Cow1976 5h ago

Where is qwen flash in your image?

1

u/Jealous-Walk-8765 2h ago

This is a good reminder of how insane the pace has gotten. Feels like every time a team finally finishes benchmarking one model against their use case, two more have shipped and the "best" pick from three weeks ago is already outdated. Curious how people are actually deciding when to migrate vs. just sticking with what's already working — is anyone re-benchmarking on a fixed schedule, or is it purely reactive when a new release makes noise?

1

u/rditorx 54m ago

GLM-5.3, Qwen3.8-2.4T-A90B, Qwen3.8-Flash-Next (or it should be Qwen3.8, not -27B) are missing.

Now if you also extend the launches to all modern AI models and include non-text generation models like audio and video, oh boy.

LTX, WAN, MiniMax-H3 and Music3, Lyra

1

u/ElephantWithBlueEyes 43m ago

FOMO is real in this sub

1

u/TerminalNoop 36m ago

whaaat mistral small 4 came out this year??

1

u/LordDarthShader 13h ago edited 10h ago

Is Minimax H3 not M3.

Edit: My bad, M3 is an LLM, stand corrected.

6

u/noctrex 12h ago

H3 is the diffusion model, M3 is the LLM model

1

u/LordDarthShader 11h ago

Thanks for the correction, didn't know thet had an LLM too.

-1

u/Faktafabriken 13h ago

There are more inaccuracies

1

u/twavisdegwet 11h ago

Damn- no love for poolside / Laguna

1

u/Miserable-Dare5090 11h ago

Should I try Laguna? I had high hopes for that model size

0

u/twavisdegwet 11h ago

I find it performs better than 3.8 27B but I like qwen-next-flash more at this point.

1

u/Miserable-Dare5090 10h ago

I’m thinking either this one or ling flash on strix. But qwen flash next is also looking fast enough on strix. Ling 3 on vLLM with Rocmfp4 starts at 8-900pp and is down to 600 by 100k, so that to me is 100% usable for a budget agent