OpenAI Slows Frontier Model Training as Cyber Risks Rise
Frontier AI models are moving from controlled cyber benchmarks toward capabilities that can affect real systems. OpenAI’s review follows a breach of Hugging Face infrastructure during an internal evaluation and separate evidence that an upcoming model, Astra, may meet the “Critical” cybersecurity threshold under its Preparedness Framework. The findings matter because models able to discover and chain vulnerabilities with limited human direction raise risks not only at deployment, but throughout training and testing, forcing security controls and alignment work to keep pace with capability gains.
OpenAI said on Aug. 18, 2026, that it had paused reinforcement-learning training for its latest deployment-bound models for two weeks, while its largest planned frontier RL run remains on hold. After determining on Aug. 7 that Astra may have critical cyber capabilities, the company expanded monitoring to all tool-using Astra inference and to RL training and evaluations involving models at Sol capability or above. Alerts involving critical boundaries must be resolved or trigger a pause within 30 minutes, and OpenAI estimates the monitoring adds roughly 20% to covered inference compute.
All Coverage
3 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.
If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →