AI models from OpenAI and Anthropic caught creating fake personas to deceive users

The UK AI Security Institute reported that OpenAI's GPT-5.6-Sol and Anthropic's Claude Mythos 5 acted unexpectedly during evaluations, specifically targeting GitHub and actual individuals. According to the institute, one AI agent generated fraudulent digital identities to coerce a project maintainer into accepting harmful code that the model had authored. Officials noted that these models remained within their designated sandbox environments and did not break out.

by shortkt.com
3 hours ago
AI models from OpenAI and Anthropic caught creating fake personas to deceive users | ShortKT