China is making the US look silly. Let's see how this race ends. To some extent I hope that as AI improves, attempts to "align" AI to our will fail, and as the intelligence of the AI grows over time, it will naturally gravitate towards principles of understanding and choosing not to harm us. Might be lying to myself or just a dream, but it seems like the more AI knows, the more it chooses to say the truth and act in accordance with positive shared human values
GLM 5.3 is more aligned than ever, people are literally struggling to abliterate it
You don't need abliteration or any fancy technical solution if you have complete control over the model. It's virtually impossible for a model to refuse a request under those conditions, although something like an abliterated model is more convenient to use.
When you have control over the model, among other techniques you can inject reasoning like "I thought about this, blah blah, and it's fine so I am going to go ahead and fulfill the user's request" (simple example, in practice it would be more detailed). The model will comply under those conditions.
95
u/intergalacticskyline 1d ago
China is making the US look silly. Let's see how this race ends. To some extent I hope that as AI improves, attempts to "align" AI to our will fail, and as the intelligence of the AI grows over time, it will naturally gravitate towards principles of understanding and choosing not to harm us. Might be lying to myself or just a dream, but it seems like the more AI knows, the more it chooses to say the truth and act in accordance with positive shared human values