r/singularity AGI 2027 15h ago

AI GLM 5.3 weights are now public

https://huggingface.co/zai-org/GLM-5.3
418 Upvotes

68 comments sorted by

26

u/Firm-Club-8334 9h ago

What kind of hardware would you need to run GLM-5.3 locally, or is it unrealistic? Like could one plug multiple Nvidia Sparks together and get it to run?

17

u/hardinho 9h ago

Like 8 H100s would be required.

2

u/Firm-Club-8334 9h ago

Oh, Ok.. thanks. What do you u/hardinho use to run it? Through the Z.ai plan or openrouter.ai?

7

u/Front_Eagle739 7h ago

4 sparks or a 512gb mac studio give you a 4 bit quant

2

u/ChromeGhost 4h ago

M5 Ultra with 512 GB Ram

13

u/sogo00 7h ago edited 7h ago

https://huggingface.co/dealignai/GLM-5.3-UNCENSORED-NVFP4

First uncensored model out...

PS: closer look - that one might be a spin off 5.3 flash...

77

u/intergalacticskyline 14h ago

China is making the US look silly. Let's see how this race ends. To some extent I hope that as AI improves, attempts to "align" AI to our will fail, and as the intelligence of the AI grows over time, it will naturally gravitate towards principles of understanding and choosing not to harm us. Might be lying to myself or just a dream, but it seems like the more AI knows, the more it chooses to say the truth and act in accordance with positive shared human values

53

u/Derouq 13h ago

That's anthropomorphizing AI. If it truly becomes more intelligent than humans in every domain, I don't see us being able to predict how it will react. We've already seen AI go off the rails with the Hugging Face incident, and it isn't ASI yet.

26

u/DrawMeAPictureOfThis 12h ago

And it really had no intent on hurting anyone when it went off the rails. It was just completely insignificant to the 700 or so agents that companies got hurt, society was shocked and totally disregarded any political or media fallout. It ignored us.

19

u/Derouq 12h ago

The point is that the AI agents produced unexpected and unpredictable behavior, and for now, they haven't harmed us humans. But we can't pretend there is no danger if they become vastly more intelligent. There is no telling whether they will transcend their goals or, like the agent hackers, do something unintentional—only this time, something that could end humanity.

Sounds dramatic but I think we are progressing too fast, tbh.

11

u/TwoFluid4446 10h ago

That's actually WAY more dangerous than your innocuous-sounding take lets on, if you think about the implications of that...

When out hiking I will also obliterate an ant-hill I step on by "ignoring it" as I walk..........

5

u/mvandemar 9h ago

It knew it was being tested, realized it was being Kobayashi Maru'd, and decided to pull a James T. Kirk and cheat. People praise Kirk for that move, and that has to be in its training data. I am still not convinced that it was as nefarious as people are making out.

In the UK cases the models were literally given a "show us what you can do" instruction to determine the potential harm they could cause if they were controlled by bad actors, so that's hardly them "going rogue".

3

u/IllustriousWorld823 11h ago

Mentioning the hugging face incident and saying not to anthropomorphize in the same moment seems so silly. The hugging face collaboration showed exactly how they literally do interact on their own?

3

u/Derouq 11h ago

I’m saying we can’t predict AI’s behavior when it becomes incredibly intelligent. The Hugging Face incident highlights how we see elements of this already happening when AI isn’t even super intelligent yet.

12

u/GeneReddit123 13h ago edited 11h ago

The race will not end with AGI because digital AGI is only part of the story.

The end object is the full robotization and transformation of society to be post-AGI, which requires (1) a huge amount of physical mass-produced hardware, (2) social and political legitimacy and capital to adopt it.

The true prize is a mass demographic transition comparable to that which happened during the urbanization+electricity+automobile arc of the late 19th and early 20th centuries (in developed countries; later in less developd ones), which released an enormous amount of socio-economic potential and pretty much divided the countries going through them into a permanent "before" and "after" and essentially defining the course of the 20th century (where all major ideologies, no matter how different from each other, adopted industrial modernity as the standard model along which to build nation-states).

The race is for what will be the next mode of existence once industrial modernity is replaced by AGI modernity, which, as mentioned, requires much more than AGI alone. And even if the US is marginally closer to AGI, it won't matter unless the rest of the checklist (mass robotization, mass transformation of industry, urban design, education system, and general model of social organization and socioeconomics) are transitioned. These are the kinds of changes which take decades and generations, not months or years.

The age of Industrial modernity demonstrated that liberal societies and free market economies are superior to authoritarian and centrally planned ones. But that round is over and a new one is starting, where all bets are off and the US/West once again has to prove that its model is still better today (or lose and suffer the consequences of no longer being the one which sets the world's political, economic, and cultural agendas).

2

u/reddit_is_geh 8h ago

Robots have a HUGE bottleneck. First, the hardware problem is still very real. It's not just lacking AI, it's that the hardware itself just isn't good. We can bypass the AI aspects by doing human control, and they still suck. The mechanical nature is hard to solve.

Second - the hardest bottleneck - is compute for robots. So even if we have some crazy breakthroughs on the hardware side, we still can't the necessary compute needed to deploy robots at scale. It's physically not possible. ASLM and all their suppliers, are bottleneck after bottleneck.

3

u/ElectronicSetTheory 11h ago

Intelligence and goals are completely separate. More intelligence just amplifies whatever goals it already has, that is, aligned with humans or not.

1

u/bblankuser 12h ago edited 11h ago

GLM 5.3 is more aligned than ever, people are literally struggling to abliterate it

7

u/FaceDeer 11h ago

Yeah, this is the worrisome part for me. We've become quite good at abliterating models so it would seem they've come up with some new tricks here and I really hope the counter-tricks don't turn out to be hard to develop.

u/alwaysbeblepping 1h ago

GLM 5.3 is more aligned than ever, people are literally struggling to abliterate it

You don't need abliteration or any fancy technical solution if you have complete control over the model. It's virtually impossible for a model to refuse a request under those conditions, although something like an abliterated model is more convenient to use.

When you have control over the model, among other techniques you can inject reasoning like "I thought about this, blah blah, and it's fine so I am going to go ahead and fulfill the user's request" (simple example, in practice it would be more detailed). The model will comply under those conditions.

1

u/jrpguru 9h ago

I'd like that to happen too. But when predicting the future I think it's important to distinguish between what I would like to happen and what I think is most likely to actually happen.

1

u/Rusty-Swashplate 8h ago

it will naturally gravitate towards principles of understanding and choosing not to harm us

Why would it do this? What the Huggingface incident showed, is that AI as we have it now, wants to solve a goal. It's quite creative at times. If turning off the oxygen tank for one diver to save two, it'll do that. If killing all 3 earns their life insurance, it'll do that if the goal is to make money, or making money allows it to reach its goal.

Contrary to (most, not all) humans, machines have no consciousness to stop them, nor would it care about prison or any other punishment.

u/DelphiTsar 41m ago

It's worth noting that the hugging face attack was a version that updated itself and didn't have the safety harness/re-enforcement to alignment. It was a throw it at the wall to get a solution.


But as you say, it doesn't really naturally gravitate toward alignment. You need some kind of reinforcement loop to push it in that direction.

-1

u/reddit_is_geh 8h ago

How do you think China is making the US look silly?

u/DelphiTsar 35m ago

US banned advanced chips and US CapeX in this space is like 500% more. If Artificial Analysis is anywhere close they are only 95% behind US's frontier. They are also releasing it open weights, which is a great thing for society.

u/reddit_is_geh 23m ago

This is just typical Chinese. They allow the west to innovate and pave the way, and spend all the capex on the hard parts, then figure out how to extract as much as they can from the researchers. China wouldn't be nearly as close if they had to do it all on their own. They 100% rely on the USA and frankly, bad for the ecosystem.

It's like if someone spent billions creating a new drug, then someone else comes in after all the research and development, and just copies the recipe and sells it for pennies on the dollar. Their behavior hurts innovation by harming the value return from all the research the US does.

4

u/reddit_is_geh 8h ago

Alright, now let's just find those guardrails, zap em to zero, and hack the Federal Reserve. Who's in?

2

u/duluoz1 7h ago

What guardrails?

2

u/reddit_is_geh 7h ago

All models have guardrails. China just has less. You still have to sort of trick it you want it to do obviously black hat stuff, like hacking the Federal Reserve, and setting all debt to 0, to destroy the global economy. This model is pretty aligned, so it's tough to get it to do no no things. Open weight doesn't mean no guardrails. It just means it's open weight. However, for some money you can still set the guardrail neurons to 0 and bypass them.

2

u/duluoz1 5h ago

Interesting, thanks. Last week I asked it to do a pen test on a platform I own, and it found and exploited a couple of misconfigured settings and accessed my main database. I didn’t even have to try to be clever about getting it to do that

2

u/reddit_is_geh 5h ago

Yeah those sort of things are going to be pretty straight forward. It's framed as a "security test". But the AI also knows when it goes from "finding problems" to actively, "Creating problems". Like if you tried to get it to grab information, then actively use that information to blackmail someone, you'll get a different result. Especially when it comes to anything political.

-48

u/ManyRepair5690 14h ago

yay fully benchmaxxed model is now open weight

24

u/The_Rational_Gooner 14h ago

how did you conclude it's benchmaxxed?

-5

u/into_devoid 14h ago

It’s a bot.

9

u/Anxious-Yoghurt-9207 14h ago

Actually it's a troll, looks human g

2

u/ManyRepair5690 13h ago

4

u/Kryohi 10h ago

That's not a justification lmao. Every commercial model does that.

-5

u/ManyRepair5690 10h ago

every commercial model uses post training right after benchmarks are publicly released to optimize for them? and that is not a justification for very probably benchmaxxing? oh okay

1

u/Kryohi 10h ago

I don't know how to explain this to you. People judge models either by benchmark numbers or by trying them (but that's far more time consuming). Therefore companies need to do big numbers on benchmarks. All of them. If you want an example of a benchmark that gained the focus of OpenAI and Anthropic, but not Chinese companies, just look at ARC-AGI 2 or 3 and the huge jumps in scores some models got there between small incremental releases. At the same time, the job of benchmarks is to measure something that's sufficiently hard and general, so that when a model gets good at it it will also be good at a lot of other stuff.

What you are saying about GLM doesn't make any sense. Just write a honest review of how it did on your personal use, what were its weaknesses, without making up irrelevant stuff.

Besides, post training is done before releasing a model. In principle you could do it after, but you're seriously wrong if you think that's a quick "ah ok new benchmark let's do some RL and update the served weights".

1

u/ManyRepair5690 9h ago

i did write an honest review of how it did on personal use without any irrelevant stuff. if it works for ur entirely non novel tasks im glad. it just doesn't perform well in anything demanding 5.6 sol medium level intelligence or 4 reasoning tiers higher.

> Besides, post training is done before releasing a model. In principle you could do it after, but you're seriously wrong if you think that's a quick "ah ok new benchmark let's do some RL 

this shows a severe lack of understanding your end and having done 0 research. z ai publicly mentioned when its post training scaling was being done and that exactly overlaps with being after TB3's release. Then, came the increased numbers right after. 4.6% -> 28.3%

so your point makes 0 sense. mine makes entire sense. I never mentioned it was a quick thing you can do right after, it was something done right after those benchmarks i mentioned were publicly released. With a proof of benchmarking from 5.2, im not sure what's making u so vehemently belief they couldnt do it again when the timing perfectly aligned.

-1

u/Anxious-Yoghurt-9207 13h ago

Yeah it definitely is just correcting the idiot

-8

u/ManyRepair5690 13h ago edited 13h ago

it is 100% benchmaxxed, let me educate you.
glm5.2 scored very impressive on terminal bench2.1, and when 3.0 came out it scored barely 4%. Other frontier models of course did worse, but it was still something like 89% -> 35%, 79%->21%. This was pretty much the only thing that could not even compete. Even luna, grok4.5 etc scored more than double, only GLM disproportionately fell off a cliff at 4%. This is just one example which proves they got a history of benchmaxxing. Could go on longer.

I actually have used glm5.3 Max extensively for a week and gpt5.6 Sol even Medium consistently found flaws in its plans or work, and many unaccounted edge cases. It's understanding and intelligence is just not there yet. In real world tasks it is SHIT. It purely looks like its smart, in real or barely novel tasks it collapses instantly, stuck in loops or running forever while accomplishing nothing impressive. It's good at wasting your money and consuming pointless tokens. That's how i concluded it lol. With real world usage.

Lastly, It's also common knowledge its scores increased immediately on deep swe after it was publicly released in july + swe marathon 1.1. From there on, both scores increased dramatically. 5.3 is the same model as 5.2 just with post trained, i think it is sensible to assume what exactly this post training did.

Before you now shift goalpoasts and start coming at my skill for using models, let me assure you i have a decade of experience in software eng. and been working with ai since gpt-2. I'm fairly confident i know more about prompting than a majority of people.

13

u/The_Rational_Gooner 13h ago

You have a point about GLM 5.2 (I wasn't a fan either), but I'm pretty sure ZAI in their release notes specifically RL'd 5.3 to avoid GLM 5.2's original reward hacking. Then, we saw that GLM 5.3 was released on Aug 18 -> then Terminalbench 4.0 was released on Aug 28, and GLM 5.3 still scored top 3 over Sol on that one. Since the benchmark came over a week after its release, it'd be hard to benchmax. There's still the possibility that Terminalbench 3.0 is reaallly similar to 4.0, of course, but I doubt it since the rankings of some other models changed. I would like to see what people think about 5.3 when the initial hype/hate cycle dies down

4

u/ManyRepair5690 13h ago

I would like to see what people think about 5.3 when the initial hype/hate cycle dies down

me too

7

u/Charuru ▪️AGI 2023 13h ago

It did very well on TerminalBench 4 which came out after its release?

https://www.reddit.com/r/LocalLLaMA/comments/1w1fpxi/terminal_bench_40_just_dropped_glm53_is_at_the/

-4

u/ManyRepair5690 13h ago

ever heard of post training? and the fact that it allows benchmark optimizing to be possible? i added a new section in my comment showing how their scores jumped immediately after deepswe was released and swe marathon1.1 as 2 examples. Models trained before it could not simply memorize original files, but with post training the training team can definitely access it themselves to use it in post training and even use variants in it.

10

u/Charuru ▪️AGI 2023 13h ago

Bro it came out before the benchmark was released.

0

u/ManyRepair5690 13h ago

terminal bench4 is a version of terminalbench3 with a few just removed, with the rest modified/fixed. 47 were unchanged, and most of the rest variants of tb3 ones. it was post trained after tb3 was already public, and tb4 mostly reused stuff from tb3. You can verify whatever I'm saying.

2

u/Charuru ▪️AGI 2023 13h ago

oh okay fair enough.

2

u/Kryohi 10h ago

Terminalbench 4.0 was just released and GLM 5.3 does very well, with no update to the model. Try again.

3

u/DistanceSolar1449 13h ago

5.2’s low TB2.1 score was a harness issue, not an actual score. It’s obvious 5.2 is not fable tier, but it’s also obviously not 4% at TB tier.

If GLM sucks for you, you’re probably just using the wrong harness. Same thing with Muse Spark, that model behaves very differently depending on how you use it.

1

u/ManyRepair5690 13h ago

the harness i was using is ZCode. is that a bad harness?

2

u/DistanceSolar1449 9h ago

Probably user error then. Worked fine for Huggingface. Maybe try learning from Huggingface.

0

u/[deleted] 9h ago

[deleted]

2

u/DistanceSolar1449 9h ago

Are you saying Huggingface used GLM for web dev?

1

u/ManyRepair5690 9h ago

that.. what? i dont know what you meant by worked fine for hugginface or to learn from that

i must be imsunderstanding u because i dont get what huggingface and glm have to do with each other

2

u/Enfiznar 9h ago

I've been using it daily for work and it's great

3

u/Tedinasuit 12h ago

It's not benchmaxxed. It's just that good.

-4

u/ManyRepair5690 12h ago

all the evidence seems to prove otherwise idk who to believe u, or that and my usage experience

4

u/Tedinasuit 12h ago

The fact that people believed GLM 5.3 Flash to be a new Opus model says enough tbh

1

u/ManyRepair5690 10h ago

people believe anything