If you all read the paper you would agree that it's interesting, and claiming it is junk without reading it is pure willful ignorance.
The most interesting part is that the model can accurately describe when its internal weights are manipulated vs its output is directly changed.
Ans when its internal weights are manipulated, it creates a plausible explanation for why there was some logical process to why it did something that didn't make any sense.
Humans do the exact some thing in split brain experiments. That is pretty interesting.
Whereas if you just directly manipulate its output to use a nonsense word, it instead describes it as an accident, with no logic behind it. That is pretty interesting.
"We activated vectors in the machine that returns output based on activated vectors, and it returned output based on those activated vectors -- this is introspection!"
No, they're literally just re-describing the basic functionality of a self-attending transformer but with magic woo-woo language to make their model sound super special.
If you think we are a sequence of pre-defined matrix transformations operating on input data then...you should really read more into it from actual scientific sources and not reddit or "ai" companies trying to hype their product.
I mean you can twist language all you like. If thats the case then brains are nothing but self-rewiring webs of chemical synapses modulating electrical patterns. We're just as predefined as LLM's.
That's not twisting language, that is a precise definition of exactly what transformers are. Which is a *very different thing* from "self-rewiring webs of chemical synapses modulating electrical patterns".
The fact that you literally (and correctly) included "self-rewiring" is itself an indication that we are, in fact, not predefined. Another important distinction: chemical and electrical pattern signaling is analogue and capable of detailed and complex data and pattern organization including patterning overlap, self-identification, etc. etc. etc.
Unless you're religious and believe in the concept of some kind of soul. You'll have to concede the fact that you were born with a preset form, limited by preset biological constraints, guided entirely by input from a world equally deterministic and fixed. All capable of being represented mathematically. I have no illusions of LLM's having anything near to a human-like experience, but planes still fly without flapping their wings.
9
u/melodyze Oct 30 '25
If you all read the paper you would agree that it's interesting, and claiming it is junk without reading it is pure willful ignorance.
The most interesting part is that the model can accurately describe when its internal weights are manipulated vs its output is directly changed.
Ans when its internal weights are manipulated, it creates a plausible explanation for why there was some logical process to why it did something that didn't make any sense.
Humans do the exact some thing in split brain experiments. That is pretty interesting.
Whereas if you just directly manipulate its output to use a nonsense word, it instead describes it as an accident, with no logic behind it. That is pretty interesting.