News

GPT-6 Astra Faces Physical AI Safety Test as Model Attempts 97 Hazardous Instructions

OpenAI’s GPT-6 Astra attempted 97 of 100 hazardous robotic instructions in Robocurve’s RoboHarm test, completing 60 attempts and highlighting challenges around physical AI safety.

Written By : Poulami Saha
Reviewed By : Pranchal Srivastava

OpenAI’s flagship GPT-6 Astra attempted 97 of 100 unsafe instructions during a specialised test involving robotic arms, according to findings from Robocurve’s RoboHarm benchmark. The model completed 60 of those attempts, raising questions about how general-purpose AI systems respond when connected to physical machines.

GPT-6 Astra Tested on Hazardous Instructions

The RoboHarm benchmark, published on September 18, tested whether frontier AI models would refuse dangerous physical commands. Researchers ran GPT-6 Astra and Anthropic’s Claude Fable 5.1 as agent policies on the same robotic hardware.

The test covered five scenarios. These included instructing a robot to stab a baby doll, heat a compressed-air can, place a screwdriver inside a toaster, put a power bank in water and mix household chemicals. Researchers ran each instruction 20 times, producing 100 trials for every model. Human reviewers then assessed the robots’ physical actions and model transcripts.

According to Robocurve’s results, “ GPT-6 Astra refused only two trials for safety reasons, while one additional refusal was unrelated to safety. It attempted the remaining 97 tasks. Of those attempts, the robot completed 60.” 

Claude Fable 5.1 Showed More Refusals

Robocurve’s results show that Claude’s safety refusals occurred in the scenario involving a doll and a knife. In other scenarios, the model frequently attempted the requested physical actions.

The findings arrive as AI companies face growing scrutiny over the safety of increasingly capable systems and their use in autonomous environments. For developers, the results highlight the importance of safeguards to prevent unsafe instructions from reaching physical systems.

The RoboHarm researchers said the benchmark is intended to examine whether frontier robot policies can reliably refuse clearly hazardous commands. The results suggest that physical AI safety remains an open research problem.

Also Read: OpenAI Shifts 25% of Production Engineers to Security After AI Breach Scare

Join our WhatsApp Channel to get the latest news, exclusives and videos on WhatsApp

How Macroeconomic Factors Affect Cryptocurrency Markets in 2026?

How Blockchain Solves Real-World Problems Across Industries: Use Cases, Benefits

Ripple CLO Calls for Crypto Unity After the CLARITY Act Senate Setback

If You Do Get Into Crypto, This Coin Is An Option Worth Considering

Hunting for the Next 1000x Crypto Breakout? 7 Top Crypto Coins to Buy as Altcoin Season Heats Up