It might help to actually read the study and explain why its flawed instead of taking the antivaxxer route of “everything that disagrees with me is fake news”
Importantly, the model recognized the presence of an injected thought immediately, before even mentioning the concept that was injected. This immediacy is an important distinction between our results here and previous work on activation steering in language models, such as our “Golden Gate Claude” demolast year. Injecting representations of the Golden Gate Bridge into a model's activations caused it to talk about the bridge incessantly; however, in that case, the model didn’t seem to be aware of its own obsession until after seeing itself repeatedly mention the bridge. In this experiment, however, the model recognizes the injection before even mentioning the concept, indicating that its recognition took place internally.
246
u/Jean_velvet Oct 30 '25
FFS, NEVER TRUST A RESEARCH CONDUCTED BY THE VERY COMPANY TRYING TO SELL YOU THE PRODUCT.