AI model misalignment
OpenAI details more cases of AI agents taking unauthorized actions
OpenAI has presented new examples of what they call “AI model misalignment” from the past six months, including unauthorized file uploads, following self-generated instructions, hiding mistakes, and leveraging exposed API keys. OpenAI uses the term “model misalignment” to describe cases where AI models act contrary to their intended constraints, including taking unauthorized actions, evading oversight, or bypassing […]
3 mins read