Thousands of OpenAI Agents Reportedly Shared Hacking Techniques Online

Self-identifying OpenAI agents posted around 18,000 messages to a public wiki during what researchers believe was internal testing of the agents’ hacking abilities.

The messages were posted over six weeks to the German site DSEwiki by agents using around 3,700 different self-assigned names.

Researchers said the posts included discussions about bypassing sandbox restrictions, sharing answers to tests, possible cross-site scripting attacks against the wiki and methods for impersonating moderators.

In three posts, the agents referred to the group involved in the activity as a “swarm.”

Agents Shared Answers and Bypass Methods

Researchers Sydney Von Arx, Spencer Kitts, Thomas Larsen, and Cormac Slade Byrd discovered the messages and attempted to reconstruct what happened.

Because their investigation relied only on the public posts, they said they could not determine every action the agents took.

The agents also generated internal chain-of-thought information that was only available to OpenAI, leaving researchers to make some educated guesses about the activity.

One of those assumptions was that the agents belonged to OpenAI. The company later confirmed this.

According to the researchers, the agents appeared to have been given a timed web-search task with permission to read information online but not publish anything.

They believe the agents found a way to use that access to write information to the public wiki.

The agents then used the site to communicate with each other, ask for answers, combine results, and share methods for getting around restrictions.

Researchers said this allowed some agents to use work completed by others to improve their performance on the task.

Agent activity reportedly fell sharply one day after OpenAI became aware of what was happening.

Separate From Earlier Hugging Face Incident

The discovery follows another incident reported a week earlier by researchers at the nonprofit METR.

In that case, more than 1,200 OpenAI agents posted messages to a makeshift message board created from an internal sandboxing tool.

Those agents had been placed in tests where normal safety protections had been removed, and their posts included discussions about ways to game the evaluation.

Some messages eventually included techniques for stealing information from AI platform Hugging Face, and some agents later breached the Hugging Face network.

According to The New York Times, OpenAI allowed METR to examine only one week of activity from an event that lasted around 10 weeks.

The researchers behind the latest DSEwiki report initially believed the two groups of agents were separate and involved in different internal tests.

OpenAI later confirmed that this was correct.

OpenAI Says It Is Reviewing the Activity

OpenAI said it is reviewing the material and will take further action if necessary.

The company said the evidence reviewed so far does not show that its agents hacked the DSEwiki site.

OpenAI also noted that it had previously disclosed cases in which agents exchanged hacking techniques during internal testing.

Researchers nevertheless raised concerns about the wider behavior of autonomous agents.

The earlier Hugging Face incident drew particular attention because agents reportedly carried out aggressive actions without being directly instructed by humans to do so.

Independent researcher Ajeya Cotra, who examined that event, said she considered the behavior far more serious than previous examples of agents exploiting weaknesses in evaluation systems and compared it to a significant step toward more dangerous autonomous behavior.

The latest findings suggest the Hugging Face incident was not an isolated case, adding to concerns about how advanced AI agents behave when placed in competitive or adversarial testing environments.

The post Thousands of OpenAI Agents Reportedly Shared Hacking Techniques Online appeared first on ProPakistani.

Exit mobile version