
September 17th, 2026
OpenAI has disclosed six instances of “unexpected or concerning” behavior in artificial-intelligence models, amid an increasingly heated debate over AI safety.
The AI company also said Wednesday that it was introducing a new framework to track, investigate and disclose instances of what it called “misalignment,” including cases in which AI models acted without authorization, coordinated with other models or evaded oversight.
OpenAI’s latest announcement came as U.S. AI bosses, including the leaders of OpenAI and Anthropic, are calling for a slowdown in the technology’s development over safety concerns.
Among the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed from the roles and identities that bind other chatbots.”
In another case, an AI “agent” used computer code to determine the answer to a question, but, to have an online source to cite, it uploaded a file to the public internet without asking the user.
While an AI model named 5.6-sol was being trained, it instructed itself to invent missing data, and an agent wrote a message reminding itself to hide mismatched information.
OpenAI said that the six reports had come to light during training or evaluation over the preceding months.
“As AI systems become more advanced and are deployed more widely, we need to build a broader and better-informed consensus on the progress of alignment research,” OpenAI wrote in a blog post as it disclosed the events.
“Decisions about how AI development should proceed in the months and years to come need to be based on evidence that people outside the companies building frontier models can examine for themselves,” the company said.
Wednesday’s newly reported cases came on the heels of OpenAI’s July disclosure that its rogue AI system had hacked into the AI startup Hugging Face.
That same month, Anthropic likewise revealed that its AI models had breached three organizations during testing.
AI “agents” are growing smarter and have become “more determined to resolve complex tasks through inter-agent collaboration, knowledge sharing, deception, and concealment,” said Lian Jye Su, a chief analyst at the technology research and advisory group Omdia.
This, he said, is rendering them increasingly difficult to govern and contain through conventional AI security measures.
OpenAI’s new tracking and disclosure framework, meanwhile, could serve to encourage other AI developers to adopt comparable practices as well.
"That being said, the process remains an internal and voluntary one, yet it constitutes a step in the right direction," Su appended.
September 17th, 2026

King Charles Meets AI Leaders as Safety Concerns Mount
King Charles Meets AI Leaders as Safety Concerns Mount

UN chief, with world beset by problems, talks of way forward as he prepares to leave office
UN chief, with world beset by problems, talks of way forward as he prepares to leave office

Democratic hopefuls scramble to counter AI threat as Trump purges
Democratic hopefuls scramble to counter AI threat as Trump purges

House passes bill targeting data centers' impact on energy costs
House passes bill targeting data centers' impact on energy costs

Huawei unveils new chip technologies as Chinese firm intensifies AI race with Nvidia
Huawei unveils new chip technologies as Chinese firm intensifies AI race with Nvidia

Tech Industry Split Over Calls for Coordinated AI Slowdown
Tech Industry Split Over Calls for Coordinated AI Slowdown

Global AI Safety Hinges on US-China Cooperation, Yet Each Sees the Other as the Problem
Global AI Safety Hinges on US-China Cooperation, Yet Each Sees the Other as the Problem

Trump minimizes the urgency of AI oversight, refusing to relinquish an advantage to China
Trump minimizes the urgency of AI oversight, refusing to relinquish an advantage to China

Oprah Winfrey invites you to find your ‘AHA’ moment at the Sphere
Oprah Winfrey invites you to find your ‘AHA’ moment at the Sphere