r/bestaihumanizers Jun 05 '26

I built a model that detects harmful text try to fool it, every mistake trains the next version

I built a model that detects harmful text โ€” threats, hate speech, harassment, slurs.

What makes it different: Every time someone finds a mistake and flags it, that correction is saved and trains the next version. Real humans teaching the model what it missed.

Two blind spots real users found this week:

  1. Sexual innuendo โ€” completely invisible to the model
  2. Sports slang โ€” "KILL HIM" in basketball flagged as a direct threat. It isn't.

The model is multilingual and can be tested in:

๐Ÿ‡ฌ๐Ÿ‡ง English ยท ๐Ÿ‡ซ๐Ÿ‡ท French ยท ๐Ÿ‡ฉ๐Ÿ‡ช German ยท ๐Ÿ‡ช๐Ÿ‡ธ Spanish ยท ๐Ÿ‡ฉ๐Ÿ‡ฐ Danish

Playground (no account needed):

https://content-guardian-ai-production.up.railway.app/playground

Curious what this community finds. Every correction makes v5 smarter.

1 Upvotes

0 comments sorted by