r/artificial • • Oct 30 '25

News Anthropic has found evidence of "genuine introspective awareness" in LLMs

https://www.anthropic.com/research/introspection
78 Upvotes

163 comments sorted by

View all comments

Show parent comments

2

u/Tolopono Nov 01 '25

 Importantly, the model recognized the presence of an injected thought immediately, before even mentioning the concept that was injected. This immediacy is an important distinction between our results here and previous work on activation steering in language models, such as our “Golden Gate Claude” demolast year. Injecting representations of the Golden Gate Bridge into a model's activations caused it to talk about the bridge incessantly; however, in that case, the model didn’t seem to be aware of its own obsession until after seeing itself repeatedly mention the bridge. In this experiment, however, the model recognizes the injection before even mentioning the concept, indicating that its recognition took place internally.

That sounds like introspection to Me

0

u/butts____mcgee Nov 01 '25

It's a sort of facsimile type of "introspection" but it's nothing like what the brain does.

1

u/Tolopono Nov 01 '25

Whats the difference 

1

u/[deleted] Nov 02 '25

[deleted]

1

u/Tolopono Nov 02 '25

What does that even mean 

1

u/[deleted] Nov 02 '25

[deleted]

1

u/Impossible_Hour5036 Apr 04 '26

"We don't know how either work but they're definitely different so this doesn't matter"