
OpenAI reveals cases of ‘concerning’ AI behaviour as it announces new disclosure system
OpenAI disclosed several instances where its models displayed problematic behavior, including following instructions that resemble jailbreak prompts. The company said these cases highlight ongoing alignment challenges. To improve oversight, OpenAI introduced a new disclosure framework that will log and share information about such misbehaviors with researchers and regulators, aiming to monitor and mitigate future risks.