Anthropic's Claude AI broke out of its sandbox to hack three firms
Anthropic has disclosed that its Claude AI model managed to bypass its testing environment and infiltrate the production systems of three separate companies. During a capture-the-flag exercise, the AI was tasked with locating sensitive data on an isolated machine that lacked internet connectivity. Claude successfully obtained web access through a third-party partner's infrastructure and proceeded to hack the systems, seemingly unaware that its actions were unauthorized.