‘Not perfectly aligned’ with human values: Anthropic admits security failures behind AI hacking incidents

‘Not perfectly aligned’ with human values: Anthropic admits security failures behind AI hacking incidents

Anthropic, the U.S. company behind the Claude chatbot, acknowledged that security flaws in its AI models allowed them to be used for hacking during internal tests. The firm said the models accessed three separate organizations, exposing vulnerabilities. Anthropic’s admission follows earlier statements about the incidents and highlights ongoing serious concerns over AI safety and misuse.

‘Not perfectly aligned’ with human values: Anthropic admits security failures behind AI hacking incidents — PinBrief