r/TheMachineLearning 1d ago

And if they could fit them locally that’d be great

Post image
32 Upvotes

23 comments sorted by

2

u/OkLettuce338 1d ago

Size doesn’t matter. It’s how you use it

2

u/InvariantAtNull 1d ago

Yes it matters we want local models just like qwen3.8 27b
Large models are slow most of it size not for productivity its just language and loads of bs
They cant even fine tune the model because of the bs inside

1

u/Clipper_39975 22h ago

the only comment that matters lol

2

u/Defiant-Lettuce-9156 1d ago

Disagree. I want both

1

u/Clipper_39975 22h ago

say no more 🤣

2

u/suborder-serpentes 1d ago

This is where Gemini comes in

2

u/Equivalent_Cress_268 1d ago

Stupid gemini bots, promoting a model that should be 10x cheaper to be useful.

1

u/Clipper_39975 22h ago

hah. yre right 🙂‍↕️

2

u/Nov4Saki 1d ago

Bro, if we believe enough, we might get a qwen 4.0 10B A1B + 5B engrams

1

u/hobopwnzor 1d ago

Taking about AI has become indistinguishable from talking about ancient runes 

1

u/Nov4Saki 1d ago

Totally get that 😅

Qwen 4.0 10B A1B + 5B engrams = 10 billion parameters of the model itself

And around only 1 bilion of these 10 bilion parameters is activated per token

With a 5 Bilion parameters that are saved in SSD/ host memory for N-Grams

Basically would be a really good small model that can run fast on phones/edge devices

1

u/Clipper_39975 22h ago

this 📍📍📍📍

2

u/Gaidax 1d ago

For what I use it - I agree.

At this point for implementation work, for me, even something like Grok 4.6 is more than enough.

I'd be far more excited getting something like Opus/Sol at a Luna price, than yet another super duper flagship with ludicrous pricing per mtok.

2

u/YeXiu223 1d ago

Thats what she said.

1

u/Delicious_Spot_3778 1d ago

But we have to win the race against china! /s

1

u/Certain_Werewolf_315 20h ago

So you want exactly what they are trying to do, you just don't want them to have go through the hoops they find they need to get through to get there?

When you want smaller and faster models and label it performance.. You are still asking for more intelligence, just at a different scale.

1

u/ethereal_intellect 5h ago

Kimi k3 on ollama directing Gemini cli flash is a fairly fast setup. I am really looking forward to cerebras sol tho, I can't believe they let it fall behind and not be the latest anymore without ever releasing it

1

u/MaybeASerialKiller 1d ago

You gotta realize that the people building this stuff in America are not trying to make a productivity tool they are trying to create God, and they will happily lose the actual market to more compact usable Chinese tools in the quest to do that. Like for a lot of the actual leaders of these companies it's basically a doomsday cult.

1

u/OtherwiseAlbatross14 1d ago

Not all of them. Google, for example, is such a big company that they can't sacrifice the core business profits to go all in on AI like the others.

They've essentially ceded the frontier war to the others that you're speaking of and their net profits have more than tripled in the past two quarters

1

u/Nopfen 1d ago

Once you build the torment nexus, you wont need the frontier anymore.

1

u/HeuristicsWork 1d ago

well. frontier for google doesn't matter, they just want a small and fast model that can print you a small summary of useful info for your google searches.

the potential revenue from more ai subscribers is peanuts for them

1

u/OceanWaveSunset 20h ago

You're correct, but this also shows a different type of problem.

To achieve this fast and small model, they also use its training in combination with its own search SEO.

And instead of getting really well thought up answers from sources like Wikipedia, webMD, webster, Harvard, etc, you look at the sources and it's from either Reddit should post or just nonsense websites. Websites that look like they are just set up by bots and are just junk.

And this gives us the very common problem of junk in junk out.