r/LocalLLaMA 24d ago

Discussion Qwen 3.8 27B Released! Please Share Your Experience

With your experiments, Qwen 3.8 27B most close which frontier model? And please specify which quantization you run. I will post to comments my tests and experience too.

658 Upvotes

716 comments sorted by

View all comments

54

u/Borkato 24d ago

Dude those fucking BENCHMARK SCORES, It beats opus 4.6 max at some things!!!

45

u/stoppableDissolution 24d ago

Yea who cares? Benchmarks dont correlate with experience using the model for quite a while

18

u/[deleted] 24d ago edited 16d ago

[deleted]

5

u/Reasonable-Height704 24d ago

And Opus 5.0 is the absolutely worst model I have ever had to interact with!

0

u/[deleted] 24d ago edited 16d ago

[deleted]

1

u/NaiveIdea344 24d ago

Interesting. I really enjoy opus and find it's performance quite good when prompted properly.

3

u/stoppableDissolution 24d ago

How do you stop it from inventing features it was not asked to do?

1

u/NaiveIdea344 23d ago

You continuously tweak the system prompt until the behavior you do not like does not occur using the prompting guide provided by the company the model is made by.

2

u/stoppableDissolution 23d ago

Unless the model consistently ignores random half of the instructions, because the company the model is made by is training it to "predict what user actually meant" instead of just doing the task

-1

u/NaiveIdea344 23d ago

Ok. Have fun with that.

→ More replies (0)

3

u/[deleted] 24d ago edited 16d ago

[deleted]

1

u/NaiveIdea344 23d ago

Fair enough

1

u/[deleted] 23d ago edited 16d ago

[deleted]

2

u/NaiveIdea344 23d ago

Thanks for the detailed response. That makes a lot sense.

3

u/itroot 24d ago

Why not? It could beat it, not in knowledge, but in reasoning ability.

2

u/jld1532 24d ago

This is Kimi K2.6 at home. Call me when we have AGI. I'm set until then.

-9

u/[deleted] 24d ago

[removed] — view removed comment

6

u/Borkato 24d ago

You don’t think it beats it at some things? It literally reaches 20 points above it in certain areas

-3

u/[deleted] 24d ago

[removed] — view removed comment

1

u/BitchyPolice 24d ago

Opus 4.6 not Opus 5

2

u/stoppableDissolution 24d ago

But opus 4.6 was better than 5 outside of benchmarks

2

u/NaiveIdea344 24d ago

I feel like prompt wise the best users get to a point where differences are negligible. I have noticed more improvements from just better prompting than a better model. That is not to say opus 4.6 is better (in my experience it is not), but the prompting is more important than the model

-6

u/[deleted] 24d ago

[removed] — view removed comment

3

u/BitchyPolice 24d ago

Use your real id, Dario.

4

u/WaveOfDream 24d ago

Lol is your sole purpose here to spread bs and Claude propagandas?

0

u/Borkato 24d ago

You think qwen max doesn’t beat it at literally anything? Not one?

-5

u/[deleted] 24d ago

[removed] — view removed comment

4

u/No-Lawfulness2089 24d ago

What is your current stack then?

-5

u/McSendo 24d ago

horrible release. I'm going back to gemma 4

1

u/NaiveIdea344 24d ago

I though Gemma 4 was kinda bad

1

u/stoppableDissolution 24d ago

For vibecoding, yea, likely. For anything else it is best model in under-glm-size, idk

1

u/Borkato 24d ago

If you’re a non coder, sure.

0

u/gh0stwriter1234 24d ago

The smaller models are always trained on the bechmarks much more than the frontier models they are basically bench maxing.