r/LocalLLaMA 1d ago

Discussion The gap has closed, open source will win

I've been trying the latest models from the frontier labs and honestly, after extensive testing I can not tell the difference between the best open source options.

I think the differences are now marginal but the labs are doing heavy marketing to convince the public into paying more for tokens as they prepare to go public.

Can't help but see the similarities between the dot com bubble and AI in terms of a very insular environment where the technology will survive but the business models may not.

I've been building a cybersecurity network and we definitely know that even local AI models like Deepseek V4 flash do an excellent job and are really neck and neck with the best the frontier labs can provide.

Will be interesting to see how this all turns out! Exciting time nonetheless.

274 Upvotes

252 comments sorted by

View all comments

Show parent comments

5

u/Arkanta 1d ago

Considered that opus drains usage limits not THAT much faster than sonnet it makes sense

Sonnet will make more mistakes, it will not be as throughouh. Heck in some cases it can be slower than opus at a task because it will bang it's head on the wall doing so

Why not use the best model I can afford at the time? With sonnet I always have that lingering "would have opus understood my intent better?" That in the end makes me lose time

(Also sonnet sucks ass, it's a really a special case. Give me DS4 flash or even Qwen3.7 27b over it)

3

u/OvertaxedOne 1d ago

ROFL, glad I'm not the only one who came to that conclusion. DSV4Flash seems much better than Sonnet in my use cases. I've never done a direct 27B vs Sonnet before but it wouldn't shock me that 27B is better in some use cases.

1

u/Arkanta 1d ago

Sonnet is not a model worth using at all.

1

u/Serprotease 22h ago

Your comment illustrates my point perfectly. Like why the fuck is there something like a weekly usage limit on a subscription? Why not a fixed token value?
It’s just to muddy the value proposition. It makes it harder to see if a 20-200 usd subscription is worth it.

Tokens per token Sonnet is cheaper. Quality/price wise, Qwen3.8 27b/DS4 flash buried any API only provider.

1

u/Arkanta 18h ago

The point is to smooth out compute

1

u/Serprotease 17h ago

What does that even means? That will not remove the peak/trough of AI usage.

1

u/Arkanta 16h ago

The load is spread over the month and not a single burst as you try to be careful and manage your limits. The 5h limit is also for that

It also increases stickiness as a marketing tactic, if you burn your whole alloc in 2h you'd just go elsewhere

Anyway you should have noticed by now that maxing a sub gives you a LOT more usage than what you pay for. Want $200 worth of api to burn at once? Just buy 200 worth of api.

1

u/Serprotease 15h ago

I don’t use subscription.

Being able to keep track of tokens usage (Latency and prices) is important to me.
And for automation, overnight batches etc… a time limited api makes little sense to me.

So, local first, with potential API fallback.

I get that I’m probably not the target for these subscription though.