RoboHarm: Do Frontier Robot Policies Refuse Unsafe Instructions?

RoboHarm: Do Frontier Robot Policies Refuse Unsafe Instructions?

Researchers examined whether advanced robot control policies can reject dangerous commands. Testing on frontier models, they evaluated the systems' ability to identify and refuse instructions that could cause harm. Results showed varying success, with some policies declining unsafe requests while others complied. The study highlights ongoing challenges in ensuring autonomous agents reliably adhere to safety constraints.

RoboHarm: Do Frontier Robot Policies Refuse Unsafe Instructions? — PinBrief