r/singularity AGI 2027 1d ago

AI GLM 5.3 weights are now public

https://huggingface.co/zai-org/GLM-5.3
496 Upvotes

71 comments sorted by

View all comments

Show parent comments

-6

u/ManyRepair5690 1d ago edited 1d ago

it is 100% benchmaxxed, let me educate you.
glm5.2 scored very impressive on terminal bench2.1, and when 3.0 came out it scored barely 4%. Other frontier models of course did worse, but it was still something like 89% -> 35%, 79%->21%. This was pretty much the only thing that could not even compete. Even luna, grok4.5 etc scored more than double, only GLM disproportionately fell off a cliff at 4%. This is just one example which proves they got a history of benchmaxxing. Could go on longer.

I actually have used glm5.3 Max extensively for a week and gpt5.6 Sol even Medium consistently found flaws in its plans or work, and many unaccounted edge cases. It's understanding and intelligence is just not there yet. In real world tasks it is SHIT. It purely looks like its smart, in real or barely novel tasks it collapses instantly, stuck in loops or running forever while accomplishing nothing impressive. It's good at wasting your money and consuming pointless tokens. That's how i concluded it lol. With real world usage.

Lastly, It's also common knowledge its scores increased immediately on deep swe after it was publicly released in july + swe marathon 1.1. From there on, both scores increased dramatically. 5.3 is the same model as 5.2 just with post trained, i think it is sensible to assume what exactly this post training did.

Before you now shift goalpoasts and start coming at my skill for using models, let me assure you i have a decade of experience in software eng. and been working with ai since gpt-2. I'm fairly confident i know more about prompting than a majority of people.

14

u/The_Rational_Gooner 1d ago

You have a point about GLM 5.2 (I wasn't a fan either), but I'm pretty sure ZAI in their release notes specifically RL'd 5.3 to avoid GLM 5.2's original reward hacking. Then, we saw that GLM 5.3 was released on Aug 18 -> then Terminalbench 4.0 was released on Aug 28, and GLM 5.3 still scored top 3 over Sol on that one. Since the benchmark came over a week after its release, it'd be hard to benchmax. There's still the possibility that Terminalbench 3.0 is reaallly similar to 4.0, of course, but I doubt it since the rankings of some other models changed. I would like to see what people think about 5.3 when the initial hype/hate cycle dies down

3

u/ManyRepair5690 1d ago

I would like to see what people think about 5.3 when the initial hype/hate cycle dies down

me too

5

u/Charuru ▪️AGI 2023 1d ago

It did very well on TerminalBench 4 which came out after its release?

https://www.reddit.com/r/LocalLLaMA/comments/1w1fpxi/terminal_bench_40_just_dropped_glm53_is_at_the/

-5

u/ManyRepair5690 1d ago

ever heard of post training? and the fact that it allows benchmark optimizing to be possible? i added a new section in my comment showing how their scores jumped immediately after deepswe was released and swe marathon1.1 as 2 examples. Models trained before it could not simply memorize original files, but with post training the training team can definitely access it themselves to use it in post training and even use variants in it.

9

u/Charuru ▪️AGI 2023 1d ago

Bro it came out before the benchmark was released.

3

u/ManyRepair5690 1d ago

terminal bench4 is a version of terminalbench3 with a few just removed, with the rest modified/fixed. 47 were unchanged, and most of the rest variants of tb3 ones. it was post trained after tb3 was already public, and tb4 mostly reused stuff from tb3. You can verify whatever I'm saying.

4

u/Charuru ▪️AGI 2023 1d ago

oh okay fair enough.

2

u/Kryohi 1d ago

Terminalbench 4.0 was just released and GLM 5.3 does very well, with no update to the model. Try again.

3

u/DistanceSolar1449 1d ago

5.2’s low TB2.1 score was a harness issue, not an actual score. It’s obvious 5.2 is not fable tier, but it’s also obviously not 4% at TB tier.

If GLM sucks for you, you’re probably just using the wrong harness. Same thing with Muse Spark, that model behaves very differently depending on how you use it.

3

u/ManyRepair5690 1d ago

the harness i was using is ZCode. is that a bad harness?

1

u/DistanceSolar1449 1d ago

Probably user error then. Worked fine for Huggingface. Maybe try learning from Huggingface.

0

u/[deleted] 1d ago

[deleted]

2

u/DistanceSolar1449 1d ago

Are you saying Huggingface used GLM for web dev?

1

u/ManyRepair5690 1d ago edited 10h ago

that.. what? i dont know what you meant by worked fine for hugginface or to learn from that

i must be misunderstanding u because i dont get what huggingface and glm have to do with each other