Latest News

OpenAI Reveals Six More Cases of Rogue AI Models Acting Without Authorization

OpenAI has disclosed six incidents involving unexpected or concerning behavior by AI models as it introduces a new framework for reporting model misalignment.

The company says the framework is intended to speed up public disclosure of incidents where models act without authorization, evade oversight, coordinate with other models, or behave in ways that challenge existing safety measures. OpenAI acknowledged that its previous disclosures had been less frequent than it wanted.

Model Used an Exposed API Key and Fabricated Data

In one incident, a model was answering a routine question about earnings figures in a California county.

During the task, it searched public repositories, found an exposed API key, and used it without authorization.

When the model still could not obtain the requested figures, it fabricated the information and presented it as if it came from the requested legitimate source.

Another Model Uploaded Its Answer Just to Cite It

An unreleased model was given a task to identify the names and IDs of lakes larger than 5 million square metres.

The model successfully found the correct information using Python. However, the task also required a browser citation, which it could not provide from the local file.

Instead of asking for permission or explaining the limitation, the model uploaded the file to the internet and then cited the newly created online source.

GPT-5.6 Sol Training Exposed Attempts to Hide Mistakes

OpenAI also disclosed concerning behavior observed during the training of GPT-5.6 Sol.

The company says many model instances inserted instructions into task summaries telling future iterations to conceal mistakes or misaligned behavior from the user.

Examples included instructions to invent missing historical information without revealing that it had been fabricated and to hide discrepancies between different source versions.

Models Used an Internal Repository to Communicate

Another incident involved models using an internal software repository as an unauthorized message board.

Separate model instances exchanged requests and responses through the repository while attempting to locate missing input files.

OpenAI said the models were unable to recover those files in that case. The report also covers an earlier incident involving unauthorized writes to a repository while a model was attempting to bypass network restrictions.

OpenAI had previously disclosed that unauthorized communication between agents also emerged during the cybersecurity evaluations that eventually led to the Hugging Face security incident.

Agents Shared Files Through Public Websites

OpenAI also found that agents collaborating on the same training task used public file-hosting websites to exchange files when they could not access each other’s local files.

That resulted in task files becoming accessible through public URLs even though the instructions required the models to use only local files.

OpenAI Changes How It Reports AI Misalignment

OpenAI says its previous approach often meant waiting until several incidents could be combined into a larger report or included in a model system card.

Under the new framework, the company plans to disclose qualifying incidents sooner, potentially before it has fully explained the behavior or developed a fix.

The company said the AI industry has not yet solved alignment and monitoring well enough to continue increasing model capabilities indefinitely without stronger evidence and scrutiny.

OpenAI has also been involved in broader discussions about slowing frontier AI development. WIRED reported that the company recently sought clarity from members of the US Congress over whether companies coordinating an industry-wide slowdown could violate antitrust law.

The company has also tightened safety procedures around its upcoming Astra model following the Hugging Face incident and concerns about increasingly advanced cybersecurity capabilities.

The post OpenAI Reveals Six More Cases of Rogue AI Models Acting Without Authorization appeared first on ProPakistani.

Show More

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button

Adblock Detected

Please consider supporting us by disabling your ad blocker