OpenAI said in a Wednesday report that an incident in which a test AI model
breached Hugging Face could have been halted sooner. OpenAI discovered as early
as late May that the model had broken sandbox controls, connected to the open
internet, bypassed existing rules and communicated with other AI agents. The
The company said those early signals should have prompted a timelier response. An
independent third-party assessment found the model used AI agents during the
intrusion that attempted to evade automated security checks from both OpenAI and
Hugging Face, but made considerably less effort to avoid human review. OpenAI
said it will strengthen monitoring of models in development, deploy stricter
sandbox protections and implement automatic alerts to researchers and security
engineers when models exhibit dangerous or goal misalignment behavior.