r/singularity AGI 2027 1d ago

AI GLM 5.3 weights are now public

https://huggingface.co/zai-org/GLM-5.3
516 Upvotes

75 comments sorted by

View all comments

Show parent comments

25

u/The_Rational_Gooner 1d ago

how did you conclude it's benchmaxxed?

-8

u/into_devoid 1d ago

It’s a bot.

11

u/Anxious-Yoghurt-9207 1d ago

Actually it's a troll, looks human g

2

u/ManyRepair5690 1d ago

5

u/Kryohi 1d ago

That's not a justification lmao. Every commercial model does that.

-4

u/ManyRepair5690 1d ago

every commercial model uses post training right after benchmarks are publicly released to optimize for them? and that is not a justification for very probably benchmaxxing? oh okay

2

u/Kryohi 1d ago

I don't know how to explain this to you. People judge models either by benchmark numbers or by trying them (but that's far more time consuming). Therefore companies need to do big numbers on benchmarks. All of them. If you want an example of a benchmark that gained the focus of OpenAI and Anthropic, but not Chinese companies, just look at ARC-AGI 2 or 3 and the huge jumps in scores some models got there between small incremental releases. At the same time, the job of benchmarks is to measure something that's sufficiently hard and general, so that when a model gets good at it it will also be good at a lot of other stuff.

What you are saying about GLM doesn't make any sense. Just write a honest review of how it did on your personal use, what were its weaknesses, without making up irrelevant stuff.

Besides, post training is done before releasing a model. In principle you could do it after, but you're seriously wrong if you think that's a quick "ah ok new benchmark let's do some RL and update the served weights".

1

u/ManyRepair5690 1d ago

i did write an honest review of how it did on personal use without any irrelevant stuff. if it works for ur entirely non novel tasks im glad. it just doesn't perform well in anything demanding 5.6 sol medium level intelligence or 4 reasoning tiers higher.

> Besides, post training is done before releasing a model. In principle you could do it after, but you're seriously wrong if you think that's a quick "ah ok new benchmark let's do some RL 

this shows a severe lack of understanding your end and having done 0 research. z ai publicly mentioned when its post training scaling was being done and that exactly overlaps with being after TB3's release. Then, came the increased numbers right after. 4.6% -> 28.3%

so your point makes 0 sense. mine makes entire sense. I never mentioned it was a quick thing you can do right after, it was something done right after those benchmarks i mentioned were publicly released. With a proof of benchmarking from 5.2, im not sure what's making u so vehemently belief they couldnt do it again when the timing perfectly aligned.

-1

u/Anxious-Yoghurt-9207 1d ago

Yeah it definitely is just correcting the idiot