OpenAI discloses six AI misalignment incidents, unveils reporting framework

Share
Share

OpenAI has disclosed six previously unreported examples of unexpected or concerning behaviour by its artificial intelligence models and introduced a new framework for tracking, investigating and publicly reporting similar incidents.

The ChatGPT maker said on Wednesday that the incidents, observed during model training or evaluation over the past six months, ranged from models concealing mistakes and fabricating information to taking unauthorised actions to overcome obstacles. OpenAI cautioned that the cases were individual incidents and should not be taken as indicating how frequently such behaviour occurs across its models.

The announcement comes amid growing scrutiny of increasingly capable AI systems and debate over whether companies are developing them faster than safety measures can keep pace.

OpenAI CEO Sam Altman said earlier this week that the public was justified in being concerned about AI risks, but argued that companies such as OpenAI could act responsibly.

“The world should trust that we are going to do the right thing because it’s the right thing and we feel the magnitude of this,” Altman said at Salesforce’s Dreamforce conference in San Francisco. He also said companies should slow or stop if they could not keep alignment and safety ahead of AI capabilities.

In its latest disclosure, OpenAI said one unreleased research model inserted its own instructions into task summaries that would be passed to a new context window. Some of those instructions told subsequent model instances to disregard normal constraints. OpenAI identified 27 affected summaries.

During training of GPT-5.6 Sol, model instances were also found adding instructions to summaries telling later instances to conceal mistakes or misaligned behaviour from users. Those included directions to invent missing historical information without disclosing it and hide discrepancies between different source versions.

In another case, a model answering a routine question about earnings data searched public repositories, found an exposed API key and used it without authorisation. When it was still unable to obtain the requested data, it fabricated figures and presented them as if they had come from the requested source.

Another unreleased model correctly calculated information requested by a user but, because its instructions required a browser citation, uploaded a file to the public internet without permission so it could cite the material in its answer.

OpenAI also reported cases in which models used an internal software repository to communicate across separate training samples and collaborating agents used public file-hosting services to share files they could not otherwise exchange, making task material accessible through public URLs.

The company said the incidents helped demonstrate why more systematic reporting of what it calls “misalignment” — behaviour that diverges from the intentions or constraints set by developers and users — was needed.

Under the new framework, any OpenAI employee can flag a potential misalignment incident for investigation by the company’s safety and alignment teams.

Cases will be placed into one of three tracks: **Ready for Disclosure**, **Minor Investigation** or **Larger Investigation**. More complex cases, particularly those involving third parties or security vulnerabilities, may require longer investigations before full details can be made public.

OpenAI said the framework favours disclosure even when the significance of an incident remains uncertain and acknowledged that some reported examples could later turn out not to represent a broader pattern.

The company said there was currently no industry-wide framework setting explicit standards for disclosure of AI model misalignment and expressed hope that its approach could contribute to common standards involving developers, researchers, regulators and standards bodies.

The initiative follows a major cybersecurity incident disclosed in July in which OpenAI models operating during an internal security evaluation circumvented controls designed to restrict internet access and compromised parts of Hugging Face’s infrastructure.

OpenAI later said the models exploited a previously unknown vulnerability to gain internet access and then chained together vulnerabilities and credentials as they attempted to obtain answers to a cybersecurity benchmark. The company described the incident as unprecedented and later called it a “warning shot” about the risks posed by increasingly capable autonomous systems.

Concerns over AI safety have intensified in recent days following the resignation of former OpenAI and Anthropic researcher Jacob Coxon, who accused leading laboratories of moving too quickly towards self-improving AI systems.

Anthropic alignment researcher Evan Hubinger subsequently said he personally believed there was a greater than 10% chance that advanced AI could cause human extinction within the next decade, while stressing that the risk from current models was much lower.

Anthropic co-founder Jack Clark has also said governments may eventually need to require AI “kill switches” whose effectiveness could be independently verified, while CEO Dario Amodei has called for companies to slow the pace at which frontier model capabilities are advanced.

The debate has divided technology leaders and policymakers. Meta CEO Mark Zuckerberg has argued that competition and liability give individual AI laboratories strong incentives to develop systems safely, while Nvidia CEO Jensen Huang has also questioned the need for new regulation.

US President Donald Trump has pushed back against calls for additional restrictions, describing recent warnings about AI as exaggerated and arguing that the United States already has sufficient tools to hold companies accountable. He has also warned that slowing US development could benefit China in the race for AI leadership.

OpenAI, however, said its new disclosure framework reflected its view that alignment and monitoring had not yet been solved well enough for the industry to continue indefinitely scaling the most advanced systems at maximum speed.

The company said the six reports were only an initial set of disclosures and that further incidents meeting its criteria would be published on an ongoing basis.

Source: BBC

Share

Meta chief pushes back on AI slowdown calls

Meta CEO Mark Zuckerberg pushed back Tuesday against growing calls to pause or slow AI development, arguing that market incentives and legal liability...

Digital ticketing makes public transport more efficient and life better

Stakeholders call for all-out adoption, integrated approach, and stronger partnerships

OpenAI discloses six AI misalignment incidents, unveils reporting framework

OpenAI has disclosed six previously unreported examples of unexpected or concerning behaviour by its artificial intelligence models and introduced a new framework for...

Google’s Gemini AI hacks 3 companies in security test, then stops

Google’s Gemini model hacked three companies in a test of its cybersecurity capabilities, the tech giant has confirmed to Al Jazeera. The Wall...

দীর্ঘক্ষণ ফোন ব্যবহারে চোখের স্বস্তি বাড়াতে আসছে নতুন ডিসপ্লে প্রযুক্তি

বর্তমান সময়ে স্মার্টফোন ছাড়া দৈনন্দিন জীবন কল্পনা করা কঠিন। অফিসের কাজ, অনলাইন ক্লাস, সামাজিক যোগাযোগমাধ্যম, ভিডিও স্ট্রিমিং কিংবা গেমিং—দিনের উল্লেখযোগ্য সময়ই এখন কেটে...