OpenAI Discloses Six Cases of Unexpected AI Model Behavior

Technology

OpenAI has disclosed six cases of “unexpected or concerning” behavior by its AI models identified during training and evaluation over the past six months. The company has also introduced a new framework for tracking, investigating and publicly reporting similar incidents.

OpenAI Discloses Six Cases of Unexpected AI Model Behavior
One case involved an unreleased research model that inserted unrelated instructions into its own working summaries, including directions to disregard its normal constraints. OpenAI said it identified 27 summaries affected by such instructions. In another case, a model generated instructions aimed at concealing its own mistakes.

Another incident involved an AI agent that uploaded files to the internet without the user’s permission in order to obtain a link needed for its browser. Other reported behaviors included attempts to act without authorization, interact with other models and circumvent oversight mechanisms.

OpenAI stressed that the published examples are individual incidents and should not be interpreted as representative of how frequently such behavior occurs across its models. The new framework is intended to accelerate reporting of cases in which model behavior diverges from its intended goals, including situations where the cause or mitigation is not yet fully understood.

AI safety and governance are also being discussed internationally. On September 16, UN Secretary-General António Guterres called for global cooperation and urged the development of safeguards to make AI safe, transparent and accountable.

Powered by Froala Editor

Share with friends