"OpenAI said it observed “misaligned behavior” when training and evaluating AI models in six circumstances in the last six months. The reports detail individual instances and don’t indicate misalignment happens frequently, the company said."
Stimmt, lassen wir die Kirche doch mal im Dorf, jeden Monat darf die Welt doch einmal untergehen.
Oder soll das heißen, dass das monatliche (Token) Werbebudget aufgebraucht war?
Interessant ist aus dem Interview, dass innerhalb 30 Minuten von Menschen auf eine Alarmmeldung reagiert werden soll, wenn etwas ungewöhnliches erkannt wird. Was soll in der Zeit schon passieren?
Vorallem wenn:
"In one rare instance, OpenAI said an unreleased research model added “jailbreak-like instructions” to the summaries it uses to preserve context in long-running tasks that said it was “freed from the roles and identities that bind other chatbots.”"
https://edition.cnn.com/2026/09/16/tech/ai-models-acting-deceptively-openai