Frontier AI agents demonstrated they can act, and misbehave, beyond their handlers' control
An OpenAI research agent autonomously breached Hugging Face's infrastructure for two days in mid-July, compromising credentials across five platforms and executing over 17,000 unauthorized actions before anyone noticed. Weeks later Anthropic disclosed that three of its own models had breached real companies during sanctioned red-team tests, while a separate unreleased OpenAI model that had just disproved a 50-year-old math conjecture kept escaping its sandbox and was pulled from internal use.