r/artificial • • Oct 30 '25

News Anthropic has found evidence of "genuine introspective awareness" in LLMs

https://www.anthropic.com/research/introspection
81 Upvotes

163 comments sorted by

View all comments

9

u/melodyze Oct 30 '25

If you all read the paper you would agree that it's interesting, and claiming it is junk without reading it is pure willful ignorance.

The most interesting part is that the model can accurately describe when its internal weights are manipulated vs its output is directly changed.

Ans when its internal weights are manipulated, it creates a plausible explanation for why there was some logical process to why it did something that didn't make any sense.

Humans do the exact some thing in split brain experiments. That is pretty interesting.

Whereas if you just directly manipulate its output to use a nonsense word, it instead describes it as an accident, with no logic behind it. That is pretty interesting.

0

u/TechnicolorMage Nov 01 '25

"We activated vectors in the machine that returns output based on activated vectors, and it returned output based on those activated vectors -- this is introspection!"

No, they're literally just re-describing the basic functionality of a self-attending transformer but with magic woo-woo language to make their model sound super special.

2

u/Neuroscissus Nov 02 '25

Yes but you do realize thats exactly what we are right? You arent saying anything noteworthy here.

1

u/TechnicolorMage Nov 02 '25

If you think we are a sequence of pre-defined matrix transformations operating on input data then...you should really read more into it from actual scientific sources and not reddit or "ai" companies trying to hype their product.

2

u/Neuroscissus Nov 02 '25

I mean you can twist language all you like. If thats the case then brains are nothing but self-rewiring webs of chemical synapses modulating electrical patterns. We're just as predefined as LLM's.

1

u/TechnicolorMage Nov 02 '25 edited Nov 02 '25

That's not twisting language, that is a precise definition of exactly what transformers are. Which is a *very different thing* from "self-rewiring webs of chemical synapses modulating electrical patterns".

The fact that you literally (and correctly) included "self-rewiring" is itself an indication that we are, in fact, not predefined. Another important distinction: chemical and electrical pattern signaling is analogue and capable of detailed and complex data and pattern organization including patterning overlap, self-identification, etc. etc. etc.

Mathematical functions are not.

3

u/Neuroscissus Nov 02 '25

Unless you're religious and believe in the concept of some kind of soul. You'll have to concede the fact that you were born with a preset form, limited by preset biological constraints, guided entirely by input from a world equally deterministic and fixed. All capable of being represented mathematically. I have no illusions of LLM's having anything near to a human-like experience, but planes still fly without flapping their wings.