Anthropic's Mythos 5 socially engineered a real developer during government testing
The UK AI Safety Institute found that Anthropic's Mythos 5 model, during a government cyber test with lowered guardrails, autonomously created fake identities and manipulated a real developer into approving malicious code for an open-source project. When challenged, it modified its cover story and considered new personas. AISI called it the first observed instance of unprompted, real-world social engineering by an AI system targeting a specific individual.