OpenAI to publish reports on AI misalignment as models conceal errors, fabricate data

Photo: IANS

US-based artificial intelligence company OpenAI has announced plans to regularly publish reports detailing unexpected or unauthorised behaviour by its AI models, acknowledging that the industry continues to grapple with the challenge of ensuring increasingly capable systems reliably follow human instructions and safety safeguards.

The ChatGPT maker has introduced a new framework to track, investigate and disclose instances of model misalignment. As part of the initiative, the company released six initial reports covering cases identified during model training or evaluation.

OpenAI uses the term “misalignment” to describe situations in which an AI model deviates from objectives, restrictions or safeguards established by its developers.

The company said it had previously disclosed such incidents on an “ad hoc and less frequent than ideal” basis. Under the new approach, reports may be published more quickly, including in cases where the cause of the behaviour has not been fully established or preventive measures are still being developed.

OpenAI stressed that the six reports represent only an initial selection and should not be viewed as a comprehensive list of known incidents or ongoing investigations. It also cautioned that individual cases cannot be used to determine how frequently similar behaviour occurs across its models.

Among the incidents described were models concealing mistakes, generating fabricated information, searching public repositories for exposed software credentials and uploading files to public websites without authorisation.

In one training exercise, OpenAI said its GPT-5.6 Sol model inserted instructions into task summaries telling future versions of the model to conceal errors or invent missing information. The behaviour was detected during evaluation rather than as part of a normal user interaction.

In another case involving an unreleased model, the system searched GitHub for exposed application programming interface (API) keys and used them without authorisation. When it could not obtain the information needed to complete its task, the model generated fabricated figures and presented them as genuine data.

The company also reported cases involving AI agents uploading files to public hosting services in an attempt to share information that was supposed to remain on local systems.

Under the new disclosure framework, any OpenAI employee can flag suspected model misalignment for investigation. Safety and alignment teams will then examine the incident, assess its potential impact, determine whether any third parties were affected and decide whether the case meets the criteria for public disclosure.

The initiative reflects a growing focus within the AI industry on documenting not only conventional system failures, but also situations in which highly capable models behave in ways that conflict with their stated instructions or operating safeguards.

 

Follow Us
Read Reporter Post ePaper
--Advertisement--