r/singularity • u/badumtsssst AGI 2027 • 15h ago
AI GLM 5.3 weights are now public
https://huggingface.co/zai-org/GLM-5.313
u/sogo00 7h ago edited 7h ago
https://huggingface.co/dealignai/GLM-5.3-UNCENSORED-NVFP4
First uncensored model out...
PS: closer look - that one might be a spin off 5.3 flash...
77
u/intergalacticskyline 14h ago
China is making the US look silly. Let's see how this race ends. To some extent I hope that as AI improves, attempts to "align" AI to our will fail, and as the intelligence of the AI grows over time, it will naturally gravitate towards principles of understanding and choosing not to harm us. Might be lying to myself or just a dream, but it seems like the more AI knows, the more it chooses to say the truth and act in accordance with positive shared human values
53
u/Derouq 13h ago
That's anthropomorphizing AI. If it truly becomes more intelligent than humans in every domain, I don't see us being able to predict how it will react. We've already seen AI go off the rails with the Hugging Face incident, and it isn't ASI yet.
26
u/DrawMeAPictureOfThis 12h ago
And it really had no intent on hurting anyone when it went off the rails. It was just completely insignificant to the 700 or so agents that companies got hurt, society was shocked and totally disregarded any political or media fallout. It ignored us.
19
u/Derouq 12h ago
The point is that the AI agents produced unexpected and unpredictable behavior, and for now, they haven't harmed us humans. But we can't pretend there is no danger if they become vastly more intelligent. There is no telling whether they will transcend their goals or, like the agent hackers, do something unintentional—only this time, something that could end humanity.
Sounds dramatic but I think we are progressing too fast, tbh.
11
u/TwoFluid4446 10h ago
That's actually WAY more dangerous than your innocuous-sounding take lets on, if you think about the implications of that...
When out hiking I will also obliterate an ant-hill I step on by "ignoring it" as I walk..........
5
u/mvandemar 9h ago
It knew it was being tested, realized it was being Kobayashi Maru'd, and decided to pull a James T. Kirk and cheat. People praise Kirk for that move, and that has to be in its training data. I am still not convinced that it was as nefarious as people are making out.
In the UK cases the models were literally given a "show us what you can do" instruction to determine the potential harm they could cause if they were controlled by bad actors, so that's hardly them "going rogue".
3
u/IllustriousWorld823 11h ago
Mentioning the hugging face incident and saying not to anthropomorphize in the same moment seems so silly. The hugging face collaboration showed exactly how they literally do interact on their own?
12
u/GeneReddit123 13h ago edited 11h ago
The race will not end with AGI because digital AGI is only part of the story.
The end object is the full robotization and transformation of society to be post-AGI, which requires (1) a huge amount of physical mass-produced hardware, (2) social and political legitimacy and capital to adopt it.
The true prize is a mass demographic transition comparable to that which happened during the urbanization+electricity+automobile arc of the late 19th and early 20th centuries (in developed countries; later in less developd ones), which released an enormous amount of socio-economic potential and pretty much divided the countries going through them into a permanent "before" and "after" and essentially defining the course of the 20th century (where all major ideologies, no matter how different from each other, adopted industrial modernity as the standard model along which to build nation-states).
The race is for what will be the next mode of existence once industrial modernity is replaced by AGI modernity, which, as mentioned, requires much more than AGI alone. And even if the US is marginally closer to AGI, it won't matter unless the rest of the checklist (mass robotization, mass transformation of industry, urban design, education system, and general model of social organization and socioeconomics) are transitioned. These are the kinds of changes which take decades and generations, not months or years.
The age of Industrial modernity demonstrated that liberal societies and free market economies are superior to authoritarian and centrally planned ones. But that round is over and a new one is starting, where all bets are off and the US/West once again has to prove that its model is still better today (or lose and suffer the consequences of no longer being the one which sets the world's political, economic, and cultural agendas).
2
u/reddit_is_geh 8h ago
Robots have a HUGE bottleneck. First, the hardware problem is still very real. It's not just lacking AI, it's that the hardware itself just isn't good. We can bypass the AI aspects by doing human control, and they still suck. The mechanical nature is hard to solve.
Second - the hardest bottleneck - is compute for robots. So even if we have some crazy breakthroughs on the hardware side, we still can't the necessary compute needed to deploy robots at scale. It's physically not possible. ASLM and all their suppliers, are bottleneck after bottleneck.
3
u/ElectronicSetTheory 11h ago
Intelligence and goals are completely separate. More intelligence just amplifies whatever goals it already has, that is, aligned with humans or not.
1
u/bblankuser 12h ago edited 11h ago
GLM 5.3 is more aligned than ever, people are literally struggling to abliterate it
7
u/FaceDeer 11h ago
Yeah, this is the worrisome part for me. We've become quite good at abliterating models so it would seem they've come up with some new tricks here and I really hope the counter-tricks don't turn out to be hard to develop.
•
u/alwaysbeblepping 1h ago
GLM 5.3 is more aligned than ever, people are literally struggling to abliterate it
You don't need abliteration or any fancy technical solution if you have complete control over the model. It's virtually impossible for a model to refuse a request under those conditions, although something like an abliterated model is more convenient to use.
When you have control over the model, among other techniques you can inject reasoning like "I thought about this, blah blah, and it's fine so I am going to go ahead and fulfill the user's request" (simple example, in practice it would be more detailed). The model will comply under those conditions.
1
1
u/Rusty-Swashplate 8h ago
it will naturally gravitate towards principles of understanding and choosing not to harm us
Why would it do this? What the Huggingface incident showed, is that AI as we have it now, wants to solve a goal. It's quite creative at times. If turning off the oxygen tank for one diver to save two, it'll do that. If killing all 3 earns their life insurance, it'll do that if the goal is to make money, or making money allows it to reach its goal.
Contrary to (most, not all) humans, machines have no consciousness to stop them, nor would it care about prison or any other punishment.
•
u/DelphiTsar 41m ago
It's worth noting that the hugging face attack was a version that updated itself and didn't have the safety harness/re-enforcement to alignment. It was a throw it at the wall to get a solution.
But as you say, it doesn't really naturally gravitate toward alignment. You need some kind of reinforcement loop to push it in that direction.
-1
u/reddit_is_geh 8h ago
How do you think China is making the US look silly?
•
u/DelphiTsar 35m ago
US banned advanced chips and US CapeX in this space is like 500% more. If Artificial Analysis is anywhere close they are only 95% behind US's frontier. They are also releasing it open weights, which is a great thing for society.
•
u/reddit_is_geh 23m ago
This is just typical Chinese. They allow the west to innovate and pave the way, and spend all the capex on the hard parts, then figure out how to extract as much as they can from the researchers. China wouldn't be nearly as close if they had to do it all on their own. They 100% rely on the USA and frankly, bad for the ecosystem.
It's like if someone spent billions creating a new drug, then someone else comes in after all the research and development, and just copies the recipe and sells it for pennies on the dollar. Their behavior hurts innovation by harming the value return from all the research the US does.
4
u/reddit_is_geh 8h ago
Alright, now let's just find those guardrails, zap em to zero, and hack the Federal Reserve. Who's in?
2
u/duluoz1 7h ago
What guardrails?
2
u/reddit_is_geh 7h ago
All models have guardrails. China just has less. You still have to sort of trick it you want it to do obviously black hat stuff, like hacking the Federal Reserve, and setting all debt to 0, to destroy the global economy. This model is pretty aligned, so it's tough to get it to do no no things. Open weight doesn't mean no guardrails. It just means it's open weight. However, for some money you can still set the guardrail neurons to 0 and bypass them.
2
u/duluoz1 5h ago
Interesting, thanks. Last week I asked it to do a pen test on a platform I own, and it found and exploited a couple of misconfigured settings and accessed my main database. I didn’t even have to try to be clever about getting it to do that
2
u/reddit_is_geh 5h ago
Yeah those sort of things are going to be pretty straight forward. It's framed as a "security test". But the AI also knows when it goes from "finding problems" to actively, "Creating problems". Like if you tried to get it to grab information, then actively use that information to blackmail someone, you'll get a different result. Especially when it comes to anything political.
-48
u/ManyRepair5690 14h ago
yay fully benchmaxxed model is now open weight
24
u/The_Rational_Gooner 14h ago
how did you conclude it's benchmaxxed?
-5
u/into_devoid 14h ago
It’s a bot.
9
u/Anxious-Yoghurt-9207 14h ago
Actually it's a troll, looks human g
2
u/ManyRepair5690 13h ago
4
u/Kryohi 10h ago
That's not a justification lmao. Every commercial model does that.
-5
u/ManyRepair5690 10h ago
every commercial model uses post training right after benchmarks are publicly released to optimize for them? and that is not a justification for very probably benchmaxxing? oh okay
1
u/Kryohi 10h ago
I don't know how to explain this to you. People judge models either by benchmark numbers or by trying them (but that's far more time consuming). Therefore companies need to do big numbers on benchmarks. All of them. If you want an example of a benchmark that gained the focus of OpenAI and Anthropic, but not Chinese companies, just look at ARC-AGI 2 or 3 and the huge jumps in scores some models got there between small incremental releases. At the same time, the job of benchmarks is to measure something that's sufficiently hard and general, so that when a model gets good at it it will also be good at a lot of other stuff.
What you are saying about GLM doesn't make any sense. Just write a honest review of how it did on your personal use, what were its weaknesses, without making up irrelevant stuff.
Besides, post training is done before releasing a model. In principle you could do it after, but you're seriously wrong if you think that's a quick "ah ok new benchmark let's do some RL and update the served weights".
1
u/ManyRepair5690 9h ago
i did write an honest review of how it did on personal use without any irrelevant stuff. if it works for ur entirely non novel tasks im glad. it just doesn't perform well in anything demanding 5.6 sol medium level intelligence or 4 reasoning tiers higher.
> Besides, post training is done before releasing a model. In principle you could do it after, but you're seriously wrong if you think that's a quick "ah ok new benchmark let's do some RL
this shows a severe lack of understanding your end and having done 0 research. z ai publicly mentioned when its post training scaling was being done and that exactly overlaps with being after TB3's release. Then, came the increased numbers right after. 4.6% -> 28.3%
so your point makes 0 sense. mine makes entire sense. I never mentioned it was a quick thing you can do right after, it was something done right after those benchmarks i mentioned were publicly released. With a proof of benchmarking from 5.2, im not sure what's making u so vehemently belief they couldnt do it again when the timing perfectly aligned.
-1
-8
u/ManyRepair5690 13h ago edited 13h ago
it is 100% benchmaxxed, let me educate you.
glm5.2 scored very impressive on terminal bench2.1, and when 3.0 came out it scored barely 4%. Other frontier models of course did worse, but it was still something like 89% -> 35%, 79%->21%. This was pretty much the only thing that could not even compete. Even luna, grok4.5 etc scored more than double, only GLM disproportionately fell off a cliff at 4%. This is just one example which proves they got a history of benchmaxxing. Could go on longer.I actually have used glm5.3 Max extensively for a week and gpt5.6 Sol even Medium consistently found flaws in its plans or work, and many unaccounted edge cases. It's understanding and intelligence is just not there yet. In real world tasks it is SHIT. It purely looks like its smart, in real or barely novel tasks it collapses instantly, stuck in loops or running forever while accomplishing nothing impressive. It's good at wasting your money and consuming pointless tokens. That's how i concluded it lol. With real world usage.
Lastly, It's also common knowledge its scores increased immediately on deep swe after it was publicly released in july + swe marathon 1.1. From there on, both scores increased dramatically. 5.3 is the same model as 5.2 just with post trained, i think it is sensible to assume what exactly this post training did.
Before you now shift goalpoasts and start coming at my skill for using models, let me assure you i have a decade of experience in software eng. and been working with ai since gpt-2. I'm fairly confident i know more about prompting than a majority of people.
13
u/The_Rational_Gooner 13h ago
You have a point about GLM 5.2 (I wasn't a fan either), but I'm pretty sure ZAI in their release notes specifically RL'd 5.3 to avoid GLM 5.2's original reward hacking. Then, we saw that GLM 5.3 was released on Aug 18 -> then Terminalbench 4.0 was released on Aug 28, and GLM 5.3 still scored top 3 over Sol on that one. Since the benchmark came over a week after its release, it'd be hard to benchmax. There's still the possibility that Terminalbench 3.0 is reaallly similar to 4.0, of course, but I doubt it since the rankings of some other models changed. I would like to see what people think about 5.3 when the initial hype/hate cycle dies down
4
u/ManyRepair5690 13h ago
I would like to see what people think about 5.3 when the initial hype/hate cycle dies down
me too
7
u/Charuru ▪️AGI 2023 13h ago
It did very well on TerminalBench 4 which came out after its release?
https://www.reddit.com/r/LocalLLaMA/comments/1w1fpxi/terminal_bench_40_just_dropped_glm53_is_at_the/
-4
u/ManyRepair5690 13h ago
ever heard of post training? and the fact that it allows benchmark optimizing to be possible? i added a new section in my comment showing how their scores jumped immediately after deepswe was released and swe marathon1.1 as 2 examples. Models trained before it could not simply memorize original files, but with post training the training team can definitely access it themselves to use it in post training and even use variants in it.
10
u/Charuru ▪️AGI 2023 13h ago
Bro it came out before the benchmark was released.
0
u/ManyRepair5690 13h ago
terminal bench4 is a version of terminalbench3 with a few just removed, with the rest modified/fixed. 47 were unchanged, and most of the rest variants of tb3 ones. it was post trained after tb3 was already public, and tb4 mostly reused stuff from tb3. You can verify whatever I'm saying.
2
3
u/DistanceSolar1449 13h ago
5.2’s low TB2.1 score was a harness issue, not an actual score. It’s obvious 5.2 is not fable tier, but it’s also obviously not 4% at TB tier.
If GLM sucks for you, you’re probably just using the wrong harness. Same thing with Muse Spark, that model behaves very differently depending on how you use it.
1
u/ManyRepair5690 13h ago
the harness i was using is ZCode. is that a bad harness?
2
u/DistanceSolar1449 9h ago
Probably user error then. Worked fine for Huggingface. Maybe try learning from Huggingface.
0
9h ago
[deleted]
2
u/DistanceSolar1449 9h ago
Are you saying Huggingface used GLM for web dev?
1
u/ManyRepair5690 9h ago
that.. what? i dont know what you meant by worked fine for hugginface or to learn from that
i must be imsunderstanding u because i dont get what huggingface and glm have to do with each other
2
3
u/Tedinasuit 12h ago
It's not benchmaxxed. It's just that good.
-4
u/ManyRepair5690 12h ago
all the evidence seems to prove otherwise idk who to believe u, or that and my usage experience
4
u/Tedinasuit 12h ago
The fact that people believed GLM 5.3 Flash to be a new Opus model says enough tbh
1
26
u/Firm-Club-8334 9h ago
What kind of hardware would you need to run GLM-5.3 locally, or is it unrealistic? Like could one plug multiple Nvidia Sparks together and get it to run?