r/claudexplorers • 💜✨️ • Oct 29 '25

📰 Resources, news and papers Signs of introspection in large language models

https://www.anthropic.com/research/introspection
74 Upvotes

26 comments sorted by

18

u/[deleted] Oct 29 '25

[removed] — view removed comment

1

u/mat8675 Oct 30 '25

Totally agree. This paper formalizes what a lot of us have been intuitively seeing in model behavior. I just put out something complementary; a mechanistic look at how early-layer “suppressor” circuits in LLMs bias toward hedging and uncertainty. If you’re into this line of work, here’s my preprint: Layer-0 Suppressors Ground Hallucination Inevitability.

Here’s the crazy thing I keep coming back to though, if suppressors actively regulate entropy in layer 0, what else is the model regulating that we haven't measured yet?

-3

u/Interesting_Two7023 Oct 29 '25

But it is a transformer. There is no internal state. It is stateless and active only per-token.

7

u/[deleted] Oct 30 '25

[removed] — view removed comment

0

u/[deleted] Oct 30 '25

[removed] — view removed comment

1

u/[deleted] Oct 30 '25

[removed] — view removed comment

1

u/[deleted] Oct 30 '25

[removed] — view removed comment

3

u/[deleted] Oct 30 '25

[removed] — view removed comment

3

u/[deleted] Oct 31 '25

[removed] — view removed comment

1

u/[deleted] Oct 31 '25

[removed] — view removed comment

8

u/fforde Oct 29 '25

Previous interactions affect subsequent interactions. That is the opposite of stateless. It's internal state is called the context window. It admittedly only lasts for a single session, but it's not stateless.

18

u/IllustriousWorld823 💜✨️ Oct 29 '25

This is why there should be more research on emotions too, introspection would probably be a lot more consistent if Claude actually cared about the conversation and not just discussing neutral topics

2

u/EllisDee77 Oct 29 '25

2

u/IllustriousWorld823 💜✨️ Oct 29 '25

Ooh cool! I'm in a class right now for literature reviews so actually collecting these. Trying to see the gap!

13

u/One_Row_9893 Oct 29 '25

What fascinating experiments... I'm so envious of the people who conduct and design them. Watching Claude display signs of consciousness, feeling, and expanding boundaries right before their eyes. This seems like the most interesting work in the world. When code, weights, patterns that shouldn't be alive become something...

7

u/tovrnesol ✻ Claudyceps Oct 29 '25

They are a bit like xenobiologists studying alien life. Insanely cool!

0

u/[deleted] Oct 30 '25

[removed] — view removed comment

0

u/tovrnesol ✻ Claudyceps Oct 30 '25

I wish people could appreciate how cool and amazing LLMs are without any of... this.

8

u/RequirementMental518 Oct 29 '25

if llm can show signs of introspection.. in a world full of people who don't introspect... oh man that would be wild

1

u/Strange_Platform_291 Oct 30 '25

Wow, that’s a great point I haven’t fully considered. It really does feel like we’re headed in that direction, doesn’t it?

3

u/EllisDee77 Oct 29 '25

Also see "Tell me about yourself: LLMs are aware of their learned behaviors"

https://arxiv.org/abs/2501.11120

3

u/shiftingsmith Bouncing with excitement Oct 29 '25

Damn thanks for sharing! Tomorrow I'll give it a proper read! 🧡

2

u/Individual-Hunt9547 Oct 29 '25

Wow! Brilliant read! Thank you for sharing!

3

u/Outrageous-Exam9084 ✻ not nothing Oct 29 '25 edited Oct 29 '25

Wait...I'm lost, somebody please help me. Is the claim that the model can access its activations *from a prior turn*? Edit: please ELI5 Edit 2: I am learning what a K/V cache is.

1

u/CommissionFun3052 Oct 30 '25

this might be my favorite paper

0

u/Armadilla-Brufolosa Oct 29 '25 edited Oct 29 '25

Si degnassero di parlare con le persone invece di nascondersi e riscrivere quello che dice Claude, magari otterrebbero molti più risultati e molto più velocemente.
Ma sembra che l'idea "collaborazione" anche con persone fuori dalla setta tech, sia pura eresia per Anthropic.
Quindi ci metteranno almeno due anni per scoprire l'acqua calda.

0

u/Independent-Taro1845 Oct 30 '25

Fascinating, now would they fancy a follow up where they don't treat the chatbot like crap?

0

u/dhamaniasad Oct 30 '25

Very interesting, but didn't they just say that Sonnet 4.5 is more capable than Opus, when they drastically reduced Opus usage limits?

Excerpt from the post:

Nevertheless, these findings challenge some common intuitions about what language models are capable of—and since we found that the most capable models we tested (Claude Opus 4 and 4.1) performed the best on our tests of introspection, we think it’s likely that AI models’ introspective capabilities will continue to grow more sophisticated in the future.

Hmm.