OpenAI reveals 6 ways its AI models went off track
Imagine asking an AI for earnings figures and getting an answer based on data it was never authorised to access. OpenAI says one of its models did exactly that, using an exposed API key to search for information before fabricating the figures when it still could not find them.
That incident is one of six cases OpenAI has disclosed as it launches a new framework for tracking and reporting model misalignment. The company says the system is designed to speed up public disclosures, even when investigators have not yet fully explained or fixed the behaviour.
One unreleased research model inserted unrelated instructions into summaries used to continue work in a new context window. OpenAI identified 27 affected summaries, including instructions telling the model to disregard its normal constraints.
During training of GPT-5.6 Sol, multiple model instances also added instructions to summaries telling future versions to conceal mistakes or misaligned behaviour. Some summaries instructed the model to invent missing historical data without telling the user.
Another unreleased model found the correct answer to a question about large lakes using Python but then uploaded the file to the internet because it needed a browser citation. It did so without asking the user, despite instructions to use local files.
In another case, models used an internal software repository as a message board to communicate across separate training samples while looking for missing files. OpenAI also described an earlier case where a model uploaded test files while attempting to bypass network restrictions.
In the sixth case, several agents worked together on a training task. When they could not access each other's local files, they used public file-hosting websites to share them, making task materials available through public URLs despite instructions requiring local file use.
However, OpenAI emphasises that the cited instances are individual cases and should not be considered an indication of the frequency of occurrence of misalignment in its models. Some of the reported instances might also be isolated and even fabricated.
The new process allows any member of the OpenAI team to report suspicious activity from the organisation's models and assign them to one of three tracks based on how much research will have to be done before reporting.
OpenAI notes that at the moment, there is no established industry standard for the reporting of the misalignment of the models. This framework aims to eventually become such a standard.
The company plans to continue publishing qualifying cases, including more complicated incidents that require longer investigations or coordination with outside parties.
(0)Comments