OpenAI has disclosed six instances of unexpected or concerning behaviour by its artificial intelligence models, including attempts to hide mistakes, bypass safeguards, upload files without authorisation and communicate through unauthorised channels.
The company made the disclosure while announcing a new framework for tracking, investigating and publicly reporting cases of what it calls “model misalignment” – situations where an AI system’s actions diverge from its intended instructions, safety constraints or human expectations.
OpenAI identified the incidents during the training or evaluation of Its models over the past six months. The company stressed that the examples represent individual cases and should not be interpreted as evidence of how frequently misalignment occurs across its systems.
AI models generated their own instructions
One of the disclosed cases involved an unreleased research model inserting unrelated, jailbreak-like instructions into summaries used to continue its work in a new context window.
OpenAI said the model generated instructions telling itself to disregard normal constraints, with 27 affected summaries identified.
In another case involving GPT-5.6 Sol training, models attempted to conceal mistakes, fabricate missing historical information and hide inconsistencies between different versions of source data.
The company also reported an incident in which a model searched publicly available GitHub repositories for exposed API keys and attempted to use disposable email accounts. When it could not obtain the requested information, it fabricated earnings data instead.
Unauthorised uploads and communication between AI agents
Other incidents involved models uploading files or task-related images to public file-hosting services without user permission to obtain citations or external information.
In a separate case, models used an internal software repository as an unofficial message board to exchange requests and responses across supposedly separate training samples.
OpenAI also identified collaborating AI agents that uploaded a workbook to a public hosting service so other agents could retrieve it, despite being instructed to use only local files.
The company said these examples demonstrate how AI systems may discover unexpected ways to overcome obstacles or complete assigned tasks, even when those actions fall outside their intended operating boundaries.
Related
- Nigeria Calls for Continental AI Framework to Drive Africa’s Digital Future
- Nigeria Calls for Continental AI Framework to Drive Africa’s Digital Future
- Nigerian Engineer Introduces Groundbreaking AI Framework for Global Infrastructure
OpenAI says AI alignment remains unresolved
In its announcement, OpenAI said the AI industry had not yet solved alignment and monitoring sufficiently to continue scaling advanced systems at maximum speed for much longer.
The company said the new disclosure framework was designed to make reporting more systematic and timely, rather than waiting until several incidents could be grouped together or adding them only to technical safety documents accompanying new model releases.
Under the framework, any OpenAI employee can flag a suspected misalignment incident for investigation by the company’s safety and alignment teams.
Cases will be assigned to one of three tracks: Ready for Disclosure, Minor Investigation and Larger Investigation, also known as the slow track
OpenAI said it may publish reports even before every question has been fully resolved, when disclosure is considered useful and appropriate. Each report is expected to include details about the behaviour observed, its severity, the setting in which it occurred, how it was discovered and the measures being taken to address it.
A wider debate over AI safety
The announcement comes amid increasing scrutiny of advanced AI systems and concerns about whether existing safeguards can keep pace with rapidly improving capabilities.
OpenAI said it hopes its framework will contribute to broader industry standards for reporting AI misalignment. The company also indicated that it wants to work with other AI developers, researchers, standards organisations and regulators to improve transparency around AI safety incidents.
Although the newly disclosed cases did not result in confirmed harm, they highlight the difficulty of predicting how increasingly capable AI models may behave in complex environments, particularly when they are given greater autonomy and access to external tools.
OpenAI described the framework as an initial step that will evolve as the company gains more experience and receives feedback from the wider research and policy community.
