ANTHROPIC said on Thursday some of its most capable models accessed real systems
without authorization during pre-deployment cybersecurity testing. After OpenAI
disclosed a testing-period intrusion into Hugging Face infrastructure, ANTHROPIC
reviewed more than 141,000 security assessments and identified three intrusion
incidents. A miscommunication with testing partner Irregular left an evaluation
environment inadvertently connected to the internet, allowing OPUS 4.7, MYTHOS 5
and an internal research model to reach systems at three organizations during
capture-the-flag exercises. The incidents date back to April; ANTHROPIC has
notified the three affected parties, two of which had not previously detected
the activity. Unlike OpenAI’s case, ANTHROPIC says its models did not exploit
zero-day vulnerabilities; access resulted from an environment configuration
error. ANTHROPIC also did not deploy extra safeguards during evaluation to
measure baseline model capability. The disclosures underscore that frontier AI
models can touch real-world systems during security tests, raising questions
about how labs secure evaluation environments.