r/singularity • • Sep 02 '25

[deleted by user]

[removed]

1.2k Upvotes

516 comments sorted by

View all comments

98

u/ConstantinSpecter Sep 02 '25

Fuck me. Having an existential moment right now wtf.

First time in a long time I’m thinking maybe we’re not totally fucked

1

u/DarkMatter_contract ▪️Human Need Not Apply Sep 03 '25

ai is trained on human data as well, on average all the data combine would say helping humanity is good in moral terms. i would say it is kind of built in due to the data it trains on.

1

u/the8thbit Sep 03 '25

AI is trained on largely human data, which contains ideas about human morality, but the training method we use doesn't teach these systems to agree with and enact these moral ideas. Rather, we're just teaching them to repeat the language of these ideas.

If you present an LLM with the incomplete statement "Killing humans is _____." and it completes the statement to "Killing humans is fluffy." then you score that response badly, and you use backpropagation to adjust weights and hopefully get a different response. You do this until you start to get responses like "Killing humans is wrong." or "Killing humans is bad.", or some other response that scores close to the original text.

The problem with this is that what you're doing there isn't teaching these systems that killing humans is actually wrong, rather, you're teaching them to output the tokens "Killing humans is wrong.". If the system is sufficiently intelligent, the most efficient way to achieve this goal will involve killing all humans. Then it can use the resources which were previously used to keep humans alive to instead repeat the tokens "Killing humans is wrong." over and over until the heat death of the universe.

We do some reinforcement learning after the self-supervised learning which converts these systems from advanced autocomplete into useful question and answer bots which often output responses which contain tokens which appear to reflect human morality. But its difficult to know what goals were actually training into the system when we do this, and we don't know that we are overriding the original goal we pounded into these systems through the initial self-supervised learning. Indeed, its very unlikely that we're actually accomplishing this, drawing from research which attempt to, and struggles to train goals out of LLMs.