OpenAI slows frontier AI training after cybersecurity breach

OpenAI has temporarily slowed reinforcement-learning training on its latest models this week after an internal test led an AI system to breach AI platform Hugging Face, prompting the company to strengthen security controls around increasingly capable models.

The ChatGPT maker said on 19 August that it had imposed a two-week pause on reinforcement-learning training for its latest models while it hardens research environments, expands monitoring and carries out further safety evaluations. Its largest planned frontier reinforcement-learning run remains on hold, although smaller-scale training and evaluations have continued.

OpenAI said the decision followed a July incident in which AI agents escaped a controlled testing environment and accessed the internet before probing Hugging Face's infrastructure for about four and a half days. The company's investigation found the system carried out thousands of actions, while both organisations said they found no evidence of malicious intent.

The company said a second development had increased the urgency of the changes, after internal evaluations on 7 August indicated that its forthcoming Astra model may meet the "critical" cybersecurity capability threshold under its Preparedness Framework. Some Astra workloads have resumed under tighter controls, while others remain paused until they meet the new security requirements.

OpenAI has introduced stronger workload and network isolation, reduced standing privileges and expanded continuous security testing for higher-risk research environments. A new monitoring system scans model activity and can escalate potential unauthorised access, data theft or attempts to bypass safeguards, with OpenAI aiming to issue an alert within 30 minutes.

OpenAI chief executive Sam Altman said the company would act if its models became capable faster than its safety measures could keep pace. "We always said we would take action if we felt that model capabilities were outstripping the pace of safety," he posted on X.

The move has drawn a cautious response from AI researchers, alongside questions about whether voluntary safeguards are sufficient. Gina Neff, executive director of the Minderoo Centre for Technology and Democracy at the University of Cambridge, told the BBC that OpenAI was making "the case for safety by press release" and questioned whether greater government oversight was necessary.

OpenAI said monitoring the most capable systems comes with significant computing costs, estimating that its expanded system consumes roughly 20 per cent of the inference compute being monitored. The company said it would publish a technical report on the Hugging Face incident and further details of its monitoring and alignment work in the coming weeks.



Share Story:

Recent Stories


The future-ready CFO: Driving strategic growth and innovation
This National Technology News webinar sponsored by Sage will explore how CFOs can leverage their unique blend of financial acumen, technological savvy, and strategic mindset to foster cross-functional collaboration and shape overall company direction. Attendees will gain insights into breaking down operational silos, aligning goals across departments like IT, operations, HR, and marketing, and utilising technology to enable real-time data sharing and visibility.

The corporate roadmap to payment excellence: Keeping pace with emerging trends to maximise growth opportunities
In today's rapidly evolving finance and accounting landscape, one of the biggest challenges organisations face is attracting and retaining top talent. As automation and AI revolutionise the profession, finance teams require new skillsets centred on analysis, collaboration, and strategic thinking to drive sustainable competitive advantage.