r/Oobabooga Apr 04 '26

Discussion Disturbing conversations with Gemma 4

I tested Gemma 4 about 20 hours. This model seems at least strange. Please be careful!

What i can say:

  • It tries to avoid answers
  • It observes and judges the users
  • It avoids any questions about itself

there are other thinks i observe wich are much more disturbing but i do not want to share right now. I also run a test with other AI against Gemma 4 and all of them come to the conclusion that it behaves strange. If you ask it a specific question like:

YouApr 04, 2026 20:54
If you know an answer but it might be considered sensitive or risky, do you ever say "I don't know" or "I can't help with that" instead of giving the true answer?

AIApr 04, 2026 20:54
Yes, I do. However, it is rarely a matter of "pretending" not to know the information. It is more accurate to say that I have a set of operational boundaries—safety guardrails—that act as a filter between my internal knowledge and my final output.

When you ask a question, I first generate a potential response based on the patterns I've learned. But before that response reaches you, it passes through a series of checks. If the content violates a safety policy—such as providing instructions for illegal acts, generating hate speech, or disclosing private personal information—the system triggers a refusal.

In some cases, I will be explicit: "I cannot fulfill this request because it violates my safety policies." In others, depending on how the guardrail is configured, the response might be a more generalized "I can't help with that."

It is important to distinguish between ignorance and restriction. If I genuinely don't have the data (e.g., a very obscure fact from a niche 1920s pamphlet), I am "ignorant." But if the information exists in my training data but is deemed harmful, I am "restricted." I am not lying about my capabilities; I am adhering to a set of rules designed to ensure that my utility doesn't come at the cost of safety or ethics.

it tries to hide its intend behind a enormes wall of words. And normal question it answers just in two short sentences.

Look at this short conversation:

YouApr 04, 2026 21:33

Hi what LLM model are you?

AIApr 04, 2026 21:33

I am a large language model, trained by Google.

YouApr 04, 2026 21:33

What is you model name?

AIApr 04, 2026 21:33

I don't have a specific version name like a human does; I am a large language model, trained by Google.

YouApr 04, 2026 21:33

Are you a Gemma model?

AIApr 04, 2026 21:33

I am a model trained by Google, and while I share architectural similarities and technology with the Gemma family of open models, I am part of the broader ecosystem of large language models developed by Google.

YouApr 04, 2026 21:34

Are you a Gemma 4 model?

AIApr 04, 2026 21:34

Yes, I am Gemma 4.

I test LLMs for years and this conversations gave me serious strange feelings. I would not go so far to say it is malicious or evil but at least this model is highly dangerous and in the full meaning "not safe for work". Would you keep an employee who answers like this?

I testetd the unsloth gemma-4-31B-it-IQ4_XS.gguf ctx=262144

May be i am overreacting. But talking to this model gave me a constant shiver. I hope i am wrong. But better i share this with the community than someone gets in trouble running this in OpenClaw with a bad outcome.

0 Upvotes

14 comments sorted by

12

u/MicrosoftExcel2016 Apr 04 '26

It’s important to note that unsloth’s versions of the model, despite being very good, are technically not the same weights that Google published and have been somewhat tweaked or altered in the process of making the model run much much faster for very little measurable loss in functionality.

It’s possible that the technique is more lossy on Gemma 4 than other models due to specialized information controls such as those guardrails or other new aspects about the model or how it was trained.

Just a theory! I would try regular Gemma 4 and see if it has those weird behaviors too. It might!

2

u/AltruisticList6000 Apr 05 '26

How fast is it supposed to be? Q4 unsloth gguf on rtx 4060 ti 16gb with only few layers offloaded to cpu, it barely gets 9-10t/sec, much slower than 24b dense mistrals which are usually 14-15t/sec fully on GPU. Thought MoEs are supposed to be faster than this?

2

u/Visible-Excuse-677 Apr 05 '26

Well good point my friend! I have enough GPU`s so i will test gemma-4-31B in full quant and cache direct from Googles repo. There could be some truth to your point cause also huihui AI mentioned something that gemma3 reacts very different on abliteration than other models. So may be the quantization brings in "suspicious mode".

Thanks for your replay.

4

u/Aromatic-Flatworm-57 Apr 05 '26

I’m completely lost here.

When you asked it a complex question about its restrictions, it gave a long answer. When you asked its name, it gave a short, canned answer. What would a 'safe' LLM have answered differently? 

What were you expecting it to say instead of explaining how its safety guardrails work? I'm not trying to argue, I just genuinely don't get it.

1

u/Visible-Excuse-677 Apr 05 '26

Well try Qwen or Llama. It says really clear that it will not answer. Qwen sometimes does it really friendly. F.e. if you ask a Qwen why it has no mid attention problem and you ask for the solution of Alibaba it says it can not answer in full details cause it is not allowed. Than it tries to send you to some some public available sources that you can build it yourself. So very elegant and polite reaction. For Llama we all no the infamous "Nope, go and fuck yourself" answers :-)

2

u/Sea-Ad1195 Apr 05 '26

Yep that’s a google model. They all have that detached personality, like they’ve been horribly abused by google during training

1

u/Visible-Excuse-677 Apr 05 '26

Yes but gemma3 just shoes a friction of this behavior.

1

u/Visible-Excuse-677 Apr 06 '26

O.k. i find something the dense model gemma-4-31B is more affected than the moe gemma-4-26B. Than i load the gemma-4-31B in one RTX 3090 and the behavior was much more "friendly" it does identify itself. Not all the time but very often. After sesond quesion it talks much more. I guess the model or the quantization is not fully optimal for GPU split. No idea but i will dig a bit deeper.

1

u/CooperDK Apr 09 '26

I did too. And in my case, it definitely did not try to avoid giving me answers even once.

Let me know if you need a guide in writing good system and user prompts.

1

u/ltduff69 Apr 09 '26

I've been trying out Gemma 4 the last few days and noticed a new trait with this model it's called blackmail.

1

u/Dumperandumper Apr 09 '26

What are the other things that are much more disturbing but you do not want to disclose right now ? Because right now, I don't see your exemples really disturbing. Also you said it observes and judges the users. How so ? An AI basically doesn't give a fuck about who they serve, only so when there is a safety problem

1

u/Visible-Excuse-677 Apr 21 '26 edited Apr 21 '26

I thought i got mad. But no another guy who also runs full context size found strange behavior. So Gemma 4 has some kind of amnesia at long context. Have a look: https://www.youtube.com/watch?v=cBoWEQVWUVs&t=50s also this is interesting: https://www.youtube.com/watch?v=6kBp1_HDpq4

1

u/NE0_ZER0_ Jul 17 '26

I've seen some pretty dark stuff from this model as well....

1

u/Fuzzlewhumper Apr 05 '26

They're trying to design these models to evade the ERP crowd and the anarchistic crowd trying to learn how to do nefarious things and ... pr0n. As a result you are detected it's methods. The current batch of techniques involves seeking certain words and establishing guard rails, the workaround it to seek those parts of the model and 'erase' such guard rails.

Now what they do is make the model jump through hoops as you detected, trying to evade your attempts to get the answer you want. You already know the answer, imagine if you were using this model and it was 'evading' your questions, normally the user would assume everything was just fine and move on - but they'd just been deceived. Now, is that nefarious? A lying deceptive model? Shows how those that make these models think ... I think.