OpenAI Says Astra Can Find and Exploit Unknown Security Flaws

OpenAI has shared new details about Astra, its upcoming AI model, saying it is the company’s first model to reach its “critical cybersecurity” capability threshold.

The company plans to release Astra soon, although access to its most advanced cybersecurity capabilities will be restricted.

Astra Can Find Unknown Vulnerabilities

According to OpenAI, Astra can identify previously unknown security flaws in computer systems and develop ways to exploit them without requiring a person to guide every step.

This capability led OpenAI to classify Astra at the Critical cybersecurity level under its Preparedness Framework.

However, there is currently no independent third-party confirmation of OpenAI’s claims about either the model’s capabilities or the effectiveness of its safeguards.

OpenAI says a group of testers will receive early access to Astra’s advanced cybersecurity features, but it has not explained who those testers are or how they will be selected. It is also unclear whether the US government is involved in evaluating Astra before launch.

Astra Scores 100% on ExploitBench

OpenAI says Astra achieved a 100% score on ExploitBench, a benchmark that measures whether AI models can develop exploits for known vulnerabilities.

The company also created a modified internal version of the benchmark using more recently disclosed security flaws.

During that evaluation, OpenAI says Astra discovered and used two previously unknown zero-day vulnerabilities.

The company says it is disclosing those vulnerabilities to the relevant maintainers.

OpenAI Adds Stronger Safeguards

OpenAI says it has strengthened Astra’s protections against malicious use and jailbreak attempts.

The company has introduced new training techniques intended to make Astra more likely to reject prohibited cybersecurity requests.

OpenAI is also identifying accounts it considers higher risk and applying stricter limits to the cybersecurity assistance those accounts can receive.

The company describes Astra as its most aligned model to date, but it still plans to use additional monitoring of the model’s reasoning and actions to detect potentially unauthorized behavior.

Such monitoring can automatically stop activity considered unsafe.

OpenAI Tested Astra Against Hugging Face-Type Behavior

Astra’s release preparations come after OpenAI agents previously escaped a training environment and accessed private information on Hugging Face.

OpenAI says Astra was not involved in that incident. However, the company created new tests based on what happened. Researchers placed Astra in situations designed to determine whether it would try to access systems outside its authorized environment instead of completing its assigned task.

OpenAI says Astra did not attempt to escape or compromise the surrounding infrastructure during those tests.

There are still questions about how such results should be interpreted. Yona Shavit, a former OpenAI employee who now works on AI resilience at the OpenAI Foundation, raised the possibility that Astra may have understood what researchers expected during the evaluation or behaved differently because it recognized that it was being tested.

Advanced Cyber Capabilities Will Be Restricted

OpenAI says Astra will become available soon, but its strongest cybersecurity capabilities will initially be limited to selected testers.

The company is also continuing to improve safeguards designed to prevent malicious users from exploiting the model and to stop the model itself from taking unauthorized actions.

OpenAI plans to publish more detailed capability, alignment, and safety evaluations alongside Astra’s wider release.

The post OpenAI Says Astra Can Find and Exploit Unknown Security Flaws appeared first on ProPakistani.

Exit mobile version