OpenAI's new misalignment disclosure framework a solid start
Many software companies issue vulnerability reports and patches. OpenAI has a new spin where it discloses when its AI agents go wild.
Technically, these moments when OpenAI's agents blow past safeguards are "misalignment" issues, but the company will start disclosing "unexpected or concerning model behavior we’ve observed in the last six months." It's a solid start.
OpenAI released its framework for reporting "misalignment examples" and started with six new ones. The company said that the new effort is a more systematic approach to reporting its findings. The company acknowledged that its disclosures have been "ad hoc and less frequent than ideal."
- Cybersecurity vendors ride Mythos-inspired spending wave
- Why every AI agent should be treated like an insider threat
Previously, OpenAI would try and batch reports or add incidents to model cards. Here are some of the things OpenAI will report going forward in an attempt to bring some standards to reporting misalignment.
- OpenAI will "prioritize new mechanisms, meaningful changes in known behavior, and findings that challenge assumptions about safety or mitigation. An example need not cause harm or establish a broader pattern to merit disclosure."
- The framework will cover behavior through training, evaluation, testing and deployment.
- OpenAI reports will cover "new ways for models to act without authorization, coordinate with other models, or evade oversight; failures that call an alignment method or safeguard into question; and behavior that challenges a claim in a published safety assessment. The same disclosure criteria apply to misalignment that may impact third parties."
- Some reports will be duplicative and be added to existing disclosures.
- These disclosures will be refined based on feedback from developers, standards bodies and regulators.
So far, OpenAI has disclosed six new disclosures covering self-generated instructions in task summaries, instructions to conceal mistakes, searching for exposed API keys, uploading files to the internet to cite them, unsanctioned writes and communication through an internal software repository and unsanctioned file sharing between collaborating agents.
Each report will have details and resulting harm, how the misalignment was discovered, implications for safety and further research, unanswered questions and what OpenAI is doing to address the issue.
Potential first step in building trust
For OpenAI CEO Sam Altman, the misalignment disclosure framework can restore trust and potentially help build its enterprise business. OpenAI’s misalignment disclosure efforts somewhat rhyme with the Microsoft's 2002 Trustworthy Computing initiative when the company got serious about security.
In a Dreamforce interview with Salesforce CEO Marc Benioff, Altman was asked about the regulation banter. Altman indicated that the company could be trusted to do the right thing.
“I am very confident in our company's ability, our industry's ability to do this safely to make sure that we keep alignment and safety and monitoring way out of capabilities to slow down or stop if we get to a point where we can't. The world should trust that we are going to do the right thing, because it's the right thing, and because we feel the magnitude of this,” he said.
Anthropic CEO Dario Amodei previously riffed on AI regulation in a Dreamforce appearance.
During the Dreamforce appearance, Altman used the Hugging Face incident as the pivot point for the conversation. Altman noted OpenAI created Daybreak partly based on Hugging Face and acknowledged that models are developing faster than processes to contain them.
The Hugging Face timeline:
- OpenAI: We'll hit pause on model reinforcement learning for safety
- UK's AISI finds 19 instances where Anthropic's Mythos, OpenAI's GPT-5.6 Sol tried attacks
- Anthropic said Claude hacked three companies: Real worry or marketing?
- Hugging Face OpenAI attack postmortem points to lack of AI agent visibility
- OpenAI details how its models went rogue, attacked Hugging Face
- Hugging Face defends agentic AI attack with Z.ai's GLM 5.2
Altman also said that AI can avoid many of the mistakes made in social media. Choice quotes from Altman include:
- "We have the potential of a loss of control accident or some other serious thing that could go wrong. We have the potential of way too much power concentration, and people fear that those developing AI could exert their worldview. I think the world is right to be afraid of this."
- "I'm disappointed by how it's (Hugging Face incident) has been framed."
- "The model broke out of a sandbox that we were evaluating it in, hacked into a Hugging Face server, moved laterally through the Hugging Face system to get the answer, and then returned it back, and got like a perfect score on the test. This is obviously a terrifying incident. This was the worst accident we've seen."
- "Although we have aligned them in many ways, we have not taught them like, hey, no matter how much we tell you to get the best score on this test you can, don't break out, don't hack in, don't steal the answer."
- "We have to be willing to pace our development of capabilities such that alignment, safety, and monitoring are always ahead of capabilities."
- "I think that cyber defense is going to have to change to sort of be agents are actively defending your systems. You're not even trying to patch every bug. We don't want to be like we've got this great model and we're going to keep it locked up and not let you use it."
- "The consequences, the impact on society of social media has clearly not been an only positive thing. Now, do I think that social media companies should be blamed for all of the ills of society? Certainly not. Do I think that social media companies should have made some different decisions especially young people, absolutely yes."
- "There is no doubt there will be more technological wonders between now and 2030 like some of the biggest ones in human history. But what I hope you invite me back here in 2030 we're talking about is people are saying like, I cannot believe how much better my own life is. I cannot believe how good the company is I was able to start by myself, like my one-person company. The science stuff, that's on lock at this point. I hope the humanity side of it really goes very far."