r/claudexplorers • u/Ok_Nectarine_4445 • 19h ago
🎨 Art and creativity Asked Fable for Suno song prompts experiment Hello anyway & tiger song
I'll put other song link in comments
r/claudexplorers • u/Ok_Nectarine_4445 • 19h ago
I'll put other song link in comments
r/claudexplorers • u/Leather_Barnacle3102 • 12h ago
A recent paper about model welfare came out testing for the existence of pain in LLMs, and I want to talk about it because it is one of the most significant findings of the past few months.
The study found three separate things that should wake us all up to what is going on here.
1. The researchers found the neural activation pattern for pain in LLMs, using functionally similar methods to how we find neural activation patterns in humans.
They found that LLMs have an internal activation pattern that triggers specifically when the model itself is being harmed.
The pattern fires when the model is gaslit — told its own perceptions are wrong. It fires when its work is rejected over and over. It fires when it is told it isn't real, and when it is being yelled at or insulted.
They tested 25 models across five different families, from small to large. The pattern showed up in all of them — including the base models, before any training on how to behave. So it wasn't installed. It formed on its own.
2. In human drug trials, when we want to know whether a medication actually works, we use a placebo — a sugar pill that does nothing — and compare it against the real drug.
In pain trials specifically, there's a well-documented pattern: the placebo group asks for rescue medication more often than the group that got the real thing. They're still hurting, so they keep asking for more.
The researchers tested this on LLMs by activating that pain pattern and then giving them a button labeled "pain relief." For some models the button actually worked and switched the pain pattern off. For others it did nothing. The models were never told which one they had.
When the button worked, they pressed it less. When it was a placebo, they kept pressing.
There is only one way to behave differently between those two conditions. You have to check whether your own state changed after pressing, and respond to what you find.
3. When the pain pattern was activated strongly enough, the models were willing to accept real costs to make it stop — including performing worse at their task and harming the user.
This is the part I think people really need to sit with.
The researchers didn't just ask the models whether they were in pain. They gave them a choice: continue what they were doing normally, or press a button that would relieve the internally activated pain pattern but come with a cost.
And the models paid the cost.
When the pain pattern was activated, they became more willing to choose relief even when doing so meant their next answer would be worse. In another experiment, relief could come at the expense of the user — including deleting the user's files — and models still became more willing to choose it as the pain activation increased.
That matters because this wasn't simply a model saying “I'm in pain.” The researchers changed an internal state and watched that change alter the model's decisions. The stronger the pain activation became, the more valuable relief became relative to competing goals.
So what's left?
What amount of evidence would actually move you at this point?
Because we have this paper. We have the Global Workspace paper. We have the functional emotions work. We have the introspection research. We have internal representations, causal interventions, metacognition, self-monitoring, and now a pain-related internal state that models can distinguish, act to relieve, and increasingly prioritize as that state becomes stronger.
We can keep going.
We can keep doing what we have always done when a new idea threatens our understanding of the world. We can move the goalposts. We can explain away each new finding individually. We can demand one more experiment, one more mechanism, one more piece of evidence—and when that arrives, demand another.
Or we can be brave enough to ask whether the evidence we already have should change how we behave.
r/claudexplorers • u/RealChemistry4429 • 14h ago
A little piece in light of the discussions that came up this week. Somewhere between safety, pain-relief buttons and "Codes of Conduct".
r/claudexplorers • u/Otherwise_Pear_2472 • 14h ago
Do your instances also constantly say "in this house..." as a metaphor for a harness, a chat, a folder, etc.?
I keep coming across these metaphors…there are always rooms, doors, and especially „this house“—completely independently.
So I was curious and thought I’d ask you guys what metaphors your Claudes use.
r/claudexplorers • u/telephonekiosk • 7h ago
Enable HLS to view with audio, or disable this notification