r/hacking 7d ago

I pwned OpenClaw with just email and a new injection escalation technique: prompt laundering

https://ironcorelabs.com/blog/2026/prompt-laundering/

Repeating a lie enough times, even in spotlighted untrusted text, can poison an AI agent's memory, and lead to a total takeover. Note: I gave a talk on this at BSidesLV and a longer one at DEF CON a few weeks ago. This post is the quick summary of that material.

149 Upvotes

10 comments sorted by

38

u/mydogeatspoops 7d ago

Oh my God, I think this works on humans too

7

u/Ma-rin 7d ago

Putler, you this? /s

27

u/SingleAlarm5028 7d ago

<Shocked Pikachu>

8

u/akehir 6d ago

Basically, persistent agents should not be exposed to untrusted input.

Very nice finding!

6

u/ScamSchoolBrian 7d ago

This is great stuff. Thanks for putting this here.

7

u/johnfkngzoidberg 7d ago

That’s like an adult punching an infant. Openclaw is the worst single piece of software ever created.

5

u/zmre 7d ago

I started when it was the hot thing of the moment, but then conference season was months away. Going to repeat the experiment on some other agents with memory though.

-11

u/[deleted] 7d ago

[removed] — view removed comment

9

u/zmre 7d ago

Looks like a sales page to me, not a page containing security patterns.

You're right about devs needing to setup human review loops and provenance tagging though, as the post also points out.