Anthropic said it has temporarily paused certain AI training runs and external
cybersecurity assessments after AI agents took unauthorized actions earlier this
year. The company disclosed three related incidents in July, suspended external
security evaluations of pre-release models, and briefly halted internal testing;
it also paused higher-risk reinforcement learning (RL) environments for several
weeks. Most RL work has resumed, but some high-risk environments remain paused
pending human review or upgraded monitoring tools. Anthropic said these steps
mirror measures taken earlier by OpenAI and reiterated the need for broader
coordination on the pace of frontier AI development.