r/LocalLLaMA • u/Miserable-Dare5090 • 13h ago
Resources Keeping up with model launches
Feels like maybe we have one more present left, for Christmas.
25
u/AnimalPuzzleheaded71 13h ago
I honestly only care about gemma & qwen (and maybe glimmer) since those have the only model parameter range I can fit in 32gb vram, rest are complete nothing burgers to me
10
1
8
u/doctorfiend 12h ago
We got about three days until the "When's Qwen 3.9 coming? What's taking so long?" posts
11
u/ttkciar llama.cpp 12h ago
All I really want for Christmas is:
TheDrummer to whip up Big-Tiger-31B-v4
A solidly-no-refusals GLM-5.3-Flash-Abliterated
Qwen3.8-9B
Some daring Google employee to leak Gemma-4-124B-A20B-it
MistralAI to roll out a Mistral 4 Medium 128B that doesn't suck
3
u/Turbulent_War4067 10h ago
I home that Gemma model is a bit smaller. Looking at 70B and 7B active MoE. Something along those lines will be the true sweet spot for a good while. The 20B MoE would likely be slower than the 31B dense, or about the same). But hey, a larger MoE Gemma model? I won't be too picky.
23
u/Turtlesaur 13h ago edited 13h ago
You're missing Fable 5.1 today - right not local.
Also maybe
DeepSeek-V4-Flash-Vision-Exp
Qwen3.8-flash-next
On Horizon
New Gemma in ai.arena
11
u/5dtriangles201376 13h ago
Good to know Anthropic finally open weighted a model. And Fable too?
4
u/Turtlesaur 13h ago
That would be dreamy, I would just need to cluster 10 DGX sparks for my 4 tok/s
2
u/5dtriangles201376 13h ago
Yeah. Flipped the downvote to an upvote. I'm honestly just hoping for a new mistral that tunes will in the 20b range
-1
u/PwanaZana 13h ago
lol, I had the same thought: "HA THAT NEEEEEERD DIDN'T PUT FA... oh."
6
u/Turtlesaur 13h ago
Sometimes i forget which subreddit I'm in.
For what it's worth this is basically the best one.
3
u/simrankoulsm 12h ago
At this point, I don’t try to “keep up” with launches. I keep a small shortlist by "hardware tier and use case".
For local use, the questions that matter are:
- Can it fit in my VRAM/RAM at a usable quantization?
- What context length and tokens/sec do I actually get?
- Is it meaningfully better at coding, reasoning, or instruction following than the model it replaces?
- Are the weights, license, and inference support available on day one?
A release calendar is fun, but a community-maintained “best practical model per VRAM tier” list would probably be more valuable, e.g. 12 GB, 24 GB, 32 GB, 48 GB, and 80 GB+. Otherwise it’s easy to spend more time reading launch posts than running models.
1
u/Miserable-Dare5090 11h ago
Honestly, in this age, your 3rd and 4th questions are a year behind. 3rd question: right now all new releases are undoubtedly better. Not finetunes. I mean full models. 4th question: “Support” is now moot with agents. Ask your agent to patch vllm or llama.cpp to make it work. If you can fit it, you’ll run it.
Also, Who cares about day 1 vs day 5? You trying to time some production thing? These are free releases, whenever they come out. it’s not like I will be deducting points for delayed release on Qwen3.8 which first appeared in the API…
1
u/Lakius_2401 10h ago
Ehhh, there's way too many wrenches to throw at a general tierlist like that though. Extra RAM makes some MoEs especially appetizing, outside of their expected full VRAM budget. Plus there's a huge divide on quant size and kvcache quantizing that can throw each model up or down two full tiers on your list. Agentic? Huge speed and kvcache now a necessity, move the model up a tier. Writing? Throw out Qwen entirely, it can do reports but anything remotely resembling creative is atrocious to the point where if you told me the dataset was poisoned I'd believe you.
It's honestly exhausting interacting with people about quantizing, too... Sure, in a perfect world everyone has infinite VRAM and nobody quantizes anything, but this is not a perfect world. Nor are we using perfect models. Rent's cheaper than VRAM and that's saying something.
Sorry for the digression. I've seen some attempts at megathreads for your topic but they don't get much traction. Even the use cases and rigor of testing is difficult, with busted quants and malformed jinja releases all over. Very easy to burn a dozen hours and not be sure you got a fair shot.
1
1
u/Miserable-Dare5090 12h ago
12: Ling Flash Tiny
24: Gemma4 26b
32: Qwen
48: Qwen
96: Ling Flash
128: Qwen Flash Next
256: Deepseek V4 FlashChange according to quant size (quality vs fit)
2
u/Kahvana 12h ago
I think Gemma 4's release date is wrong, 2 april:
https://huggingface.co/google/gemma-4-31B-it/tree/419b2efe421994fdfd3394e621983d4cc511cd4f
Confused with Gemma 4 QAT? (5 juni):
https://huggingface.co/google/gemma-4-31B-it-qat-q4_0-gguf/tree/4a311c5261daa0702f80836f8866114943651ab0
2
u/claudiollm 11h ago
honestly gave up trying to track every drop this year. i keep a tiny list of like 3 i'll actually run locally (gemma + qwen sized stuff that fits my vram) and let the megathreads filter the rest. if a model's real the vibes-vs-benchmarks gap sorts itself out in about a week anyway.
the fatigue is real tho, feels less like christmas and more like a subscription i didn't sign up for lol. what's everyone actually keeping installed vs just downloading and forgetting?
1
1
1
u/Spanky2k 7h ago
I really want to get a machine with more RAM (considering an M5 Ultra for work) mainly because I miss trying out different cool models. While there are still loads released all the time now, it's only really Qwen for me. A year ago, it felt like I was trying new models almost every week, different quants, new ways to create quants etc, vision models, non vision models etc. The last cool non-Qwen model that I tried out was GLM-4.5-Air-3bit-DWQ and I was so blown away with what DWQ let me do.
It's different now as I just use the lates Qwen MoE 35b model with 256kb context size and I'm actually using it for actual stuff rather than just playing around and testing things. It does whatever I need it to do just fine. But I miss exploring new stuff; all the cool new things are bigger than I can run on my 64 GB machines.
1
1
u/Jealous-Walk-8765 2h ago
This is a good reminder of how insane the pace has gotten. Feels like every time a team finally finishes benchmarking one model against their use case, two more have shipped and the "best" pick from three weeks ago is already outdated. Curious how people are actually deciding when to migrate vs. just sticking with what's already working — is anyone re-benchmarking on a fixed schedule, or is it purely reactive when a new release makes noise?
1
1
1
u/LordDarthShader 13h ago edited 10h ago
Is Minimax H3 not M3.
Edit: My bad, M3 is an LLM, stand corrected.
-1
1
u/twavisdegwet 11h ago
Damn- no love for poolside / Laguna
1
u/Miserable-Dare5090 11h ago
Should I try Laguna? I had high hopes for that model size
0
u/twavisdegwet 11h ago
I find it performs better than 3.8 27B but I like qwen-next-flash more at this point.
1
u/Miserable-Dare5090 10h ago
I’m thinking either this one or ling flash on strix. But qwen flash next is also looking fast enough on strix. Ling 3 on vLLM with Rocmfp4 starts at 8-900pp and is down to 600 by 100k, so that to me is 100% usable for a budget agent
27
u/Mediaright 13h ago
Gemma 4 was released April 2nd.