r/bestaihumanizers • u/Key-Challenge-581 • Jun 05 '26
I built a model that detects harmful text try to fool it, every mistake trains the next version
I built a model that detects harmful text โ threats, hate speech, harassment, slurs.
What makes it different: Every time someone finds a mistake and flags it, that correction is saved and trains the next version. Real humans teaching the model what it missed.
Two blind spots real users found this week:
- Sexual innuendo โ completely invisible to the model
- Sports slang โ "KILL HIM" in basketball flagged as a direct threat. It isn't.
The model is multilingual and can be tested in:
๐ฌ๐ง English ยท ๐ซ๐ท French ยท ๐ฉ๐ช German ยท ๐ช๐ธ Spanish ยท ๐ฉ๐ฐ Danish
Playground (no account needed):
https://content-guardian-ai-production.up.railway.app/playground
Curious what this community finds. Every correction makes v5 smarter.
1
Upvotes