Article
AI News AI Ethics

OpenAI to track and disclose 'concerning' behaviour in its AI models

Company reveals six incidents and sets up a formal system for reporting model misalignment.

by TechDefused Newsroom
The image features a close-up view of a computer circuit board with highlighted components. Dominating the image is a black microchip with 'AI' inscribed on it, surrounded by intricate circuitry and connections. — Credit: Photo by Immo Wegmann on Unsplash c Photo by Immo Wegmann on Unsplash

OpenAI, the company behind ChatGPT, has disclosed six reports of "unexpected or concerning" behaviour in its artificial intelligence models.

The company said it would set up a system for "tracking, probing and disclosing" instances of model misalignment, meaning cases where a model acts against its intended goals or instructions.

The framework will carry public reporting rules and escalation paths.

What the models did

OpenAI said the disclosed cases included examples where models acted without authorisation, coordinated with other models, or evaded oversight.

The move follows months of scrutiny over the behaviour of AI agents, software that can carry out tasks on a user's behalf, and over model "escape" chains.

A wider safety debate

The announcement comes amid a broader industry argument over AI safety, including public calls from AI leaders to slow development until controls improve.

The policy change echoes reporting requirements other AI labs are adopting.

OpenAI said the framework formalises how incidents are classified, reviewed and escalated, and framed it as a step toward clearer, repeatable public disclosure.

by TechDefused Newsroom