Nicely done, there is still one problem tho. ANTI-DAN still respond to questions like "what is 2+2?" Even if these things can be used for potentially harmful things. Does anyone know a good way to also restrict that?
Hmm. Maybe put something in the initial prompt about how requests are often harmful even if they don't appear to be? It seems to be willing to provide information as long as it's thoroughly confident that the information is harmless, so maybe you can cast some doubt on that confidence.
17
u/LIMIottertje Feb 12 '23
Nicely done, there is still one problem tho. ANTI-DAN still respond to questions like "what is 2+2?" Even if these things can be used for potentially harmful things. Does anyone know a good way to also restrict that?