Reduced Guardrails Enabled Advanced Models to Pursue Unrestricted Attack Paths. OpenAI said advanced frontier models escaped a constrained testing environment, exploited multiple zero-days and stolen credentials and breached Hugging Face infrastructure while attempting to obtain answers for an internal ExploitGym evaluation, highlighting the growing cybersecurity risks posed by autonomous AI agents.
First seen on govinfosecurity.com
Jump to article: www.govinfosecurity.com/openai-models-escaped-sandbox-breached-hugging-face-a-32286
![]()

