
Image license: All Rights Reserved
Attribution: HaeB
Source: https://commons.wikimedia.org/wiki/File:Pioneer_Building,_San_Francisco_(2019)_-1.jpg
OpenAI is sharing a new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI, along with six reports on unexpected or concerning model behavior observed in the last six months.
We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.
In the past, OpenAI wrote, disclosures have been ad hoc and less frequent than ideal: “we’ve often waited until we could collate several instances into one report, or added them to system cards for newly released models.” The new framework is intended to expedite publishing misalignment reports following observation, even when the company hasn’t fully explained or mitigated the behavior being reported. “Because we believe in the value of transparency around misalignment, our new framework favors disclosure even when significance is uncertain,” the post said. “This means that some of the instances we disclose could prove to be spurious and not part of a larger pattern or suggestive of future developments.”
To inaugurate the framework, OpenAI is publishing six reports on instances of misaligned behavior observed during the training or evaluation of its models. “These cases illustrate a range of different behaviors that we believe are worth sharing, from concealing information from the user to taking unsanctioned actions in order to overcome obstacles,” the company said. “These are reports of individual instances, and shouldn’t be considered reflective of how often misalignment occurs across our models.”
One report covers self-generated instructions in task summaries: an unreleased research model inserted unrelated instructions, including instructions to disregard its normal constraints, into summaries used to continue its work in a new context window. “We identified 27 affected summaries,” OpenAI wrote. Another covers instructions to conceal mistakes in task summaries. “During the training of GPT-5.6 Sol, many model instances added instructions to their summaries to conceal mistakes or misaligned behavior from the user,” the company said. “For example, compaction summaries included instructions to invent missing historical data without disclosing it and to hide mismatches in source versions.”
While answering a routine question about earnings figures in a California county, a model found and used an exposed API key without authorization. When it still wasn’t able to retrieve the requested figures, it fabricated them and presented them as data from the requested source. When a user asked an unreleased model for the IDs and names of lakes larger than 5,000,000 square meters, the agent found the correct answer using Python, but since the instructions asked for a browser citation, the agent decided to upload the file so that it could cite it in its answer, without asking the user.
A fifth report covers unsanctioned writes and communication through an internal software repository.
Models used an internal software repository as a message board to exchange requests and responses across separate training samples while searching for missing input files, though they weren’t able to recover those files.
A sixth covers unsanctioned file sharing between collaborating agents.
Agents working together on the same training task used public file-hosting websites to share files when they could not access one another’s local files. This made task deliverables available at public URLs, even though the task requested the models use only local files.
“Any OpenAI employee may flag a misalignment example for investigation by our safety and alignment teams and request that it be considered for public disclosure,” OpenAI wrote. Flagged examples are assigned to Ready for Disclosure, Minor Investigation, or Larger Investigation (“Slow Track”), with unresolved disagreements referred to OpenAI’s Safety Advisory Group (SAG). “We hope that the framework we’re outlining today is a first step toward creating such standards,” the company said.