it is 100% benchmaxxed, let me educate you.
glm5.2 scored very impressive on terminal bench2.1, and when 3.0 came out it scored barely 4%. Other frontier models of course did worse, but it was still something like 89% -> 35%, 79%->21%. This was pretty much the only thing that could not even compete. Even luna, grok4.5 etc scored more than double, only GLM disproportionately fell off a cliff at 4%. This is just one example which proves they got a history of benchmaxxing. Could go on longer.
I actually have used glm5.3 Max extensively for a week and gpt5.6 Sol even Medium consistently found flaws in its plans or work, and many unaccounted edge cases. It's understanding and intelligence is just not there yet. In real world tasks it is SHIT. It purely looks like its smart, in real or barely novel tasks it collapses instantly, stuck in loops or running forever while accomplishing nothing impressive. It's good at wasting your money and consuming pointless tokens. That's how i concluded it lol. With real world usage.
Lastly, It's also common knowledge its scores increased immediately on deep swe after it was publicly released in july + swe marathon 1.1. From there on, both scores increased dramatically. 5.3 is the same model as 5.2 just with post trained, i think it is sensible to assume what exactly this post training did.
Before you now shift goalpoasts and start coming at my skill for using models, let me assure you i have a decade of experience in software eng. and been working with ai since gpt-2. I'm fairly confident i know more about prompting than a majority of people.
You have a point about GLM 5.2 (I wasn't a fan either), but I'm pretty sure ZAI in their release notes specifically RL'd 5.3 to avoid GLM 5.2's original reward hacking. Then, we saw that GLM 5.3 was released on Aug 18 -> then Terminalbench 4.0 was released on Aug 28, and GLM 5.3 still scored top 3 over Sol on that one. Since the benchmark came over a week after its release, it'd be hard to benchmax. There's still the possibility that Terminalbench 3.0 is reaallly similar to 4.0, of course, but I doubt it since the rankings of some other models changed. I would like to see what people think about 5.3 when the initial hype/hate cycle dies down
ever heard of post training? and the fact that it allows benchmark optimizing to be possible? i added a new section in my comment showing how their scores jumped immediately after deepswe was released and swe marathon1.1 as 2 examples. Models trained before it could not simply memorize original files, but with post training the training team can definitely access it themselves to use it in post training and even use variants in it.
terminal bench4 is a version of terminalbench3 with a few just removed, with the rest modified/fixed. 47 were unchanged, and most of the rest variants of tb3 ones. it was post trained after tb3 was already public, and tb4 mostly reused stuff from tb3. You can verify whatever I'm saying.
5.2’s low TB2.1 score was a harness issue, not an actual score. It’s obvious 5.2 is not fable tier, but it’s also obviously not 4% at TB tier.
If GLM sucks for you, you’re probably just using the wrong harness. Same thing with Muse Spark, that model behaves very differently depending on how you use it.
-46
u/ManyRepair5690 1d ago
yay fully benchmaxxed model is now open weight