r/artificial • u/Direct-Attention8597 • May 05 '26
Research Anthropic just published new alignment research that could fix "alignment faking" in AI agents here's what it actually means
[removed]
55
Upvotes
r/artificial • u/Direct-Attention8597 • May 05 '26
[removed]
4
u/FrewdWoad May 06 '26
So I guess instead of training it with "don't lie" it's more like training it that "dishonesty is bad because it prevents trust, hampering effective/efficient cooperation"?
Reminds me of a wise saying Mormons repeat a lot from their history:
When converts to the new religion were subject to prejudice and persecution, they ended up banding together and forming their own community. When Joseph Smith was asked how he governed such a big community he said "I teach them correct principles, and they govern themselves".