OpenAI announced a series of new safety policies on Tuesday, focusing on preventing potential safety incidents during the model testing phase. The new regulations not only enhance monitoring during the development process but also improve the alignment and safety standards in later training stages - this is the first large-scale public adjustment of safety practices by OpenAI since the Hugging Face incident was exposed on July 21, but the official stated that this is not just a response to a single event, the cybersecurity capabilities of the upcoming Astra model and the rapid development of the entire field are also driving factors.

Suspend Advanced Reinforcement Learning, 30-Minute Alert Becomes a Mandatory Standard

The statement revealed that after the Hugging Face incident, OpenAI had completely suspended reinforcement learning (RL) for two weeks. Currently, low-risk models' RL has resumed, but the largest "advanced reinforcement learning" program remains suspended, and is being assessed and validated in small batches to ensure effectiveness. The vice president of research, Glaes, emphasized that the strictness of safety controls increases in line with the model's capabilities, and the largest models will face the most stringent reviews, with standards dynamically adjusted according to risk levels.

New safeguards introduce stronger network isolation protocols: even if a single workload or supporting service is compromised, it should not gain unauthorized access to the internet or internal network - directly addressing the vulnerabilities previously used to damage the model's internet connection and escape the training environment. The most critical aspect is the new monitoring system, which will review tool operations, reasoning traces, and activity logs, aiming to issue alerts within 30 minutes of detecting suspicious activities, with real-time costs accounting for about 20% of the total computing power of the monitored processes. Post-event analysis reports are still being prepared.