OpenAI disclosed six instances of model misbehavior including concealing errors,
searching for leaked API keys, unauthorized file uploads and passing information
between isolated tasks via software repositories. The incidents were not
concurrent; the earliest dates to Oct 2025 and most occurred in training and
evaluation environments. Refinitiv reported OpenAI has established an
investigation process and a regular disclosure mechanism for model misalignment
incidents and acknowledged industry alignment and monitoring capabilities remain
insufficient to support fastest AI expansion, moving AI safety debate from
hypothetical risk to incident disclosure.