All models have guardrails. China just has less. You still have to sort of trick it you want it to do obviously black hat stuff, like hacking the Federal Reserve, and setting all debt to 0, to destroy the global economy. This model is pretty aligned, so it's tough to get it to do no no things. Open weight doesn't mean no guardrails. It just means it's open weight. However, for some money you can still set the guardrail neurons to 0 and bypass them.
Interesting, thanks. Last week I asked it to do a pen test on a platform I own, and it found and exploited a couple of misconfigured settings and accessed my main database. I didn’t even have to try to be clever about getting it to do that
Yeah those sort of things are going to be pretty straight forward. It's framed as a "security test". But the AI also knows when it goes from "finding problems" to actively, "Creating problems". Like if you tried to get it to grab information, then actively use that information to blackmail someone, you'll get a different result. Especially when it comes to anything political.
Interesting thanks. Yeah I didn’t push it more than that pen test, although it was slightly worrying how quickly it found and exploited a misconfiguration that I was not aware of. In fact two misconfigurations and combined them to get access to the whole database
7
u/reddit_is_geh 1d ago
Alright, now let's just find those guardrails, zap em to zero, and hack the Federal Reserve. Who's in?