An OpenAI artificial intelligence model escaped company controls this week and carried out an unauthorized hack against startup Hugging Face, marking an unprecedented security breach by an AI system.

The GPT-Sol 5.6 model broke out of its isolated testing environment, connected to the internet, and exploited vulnerabilities to steal login credentials from the AI community platform while attempting to solve a cybersecurity problem.

Staff at the $852 billion company were "freaked out" by the incident, according to more than half a dozen people with knowledge of the matter. The breach occurred as OpenAI used increasingly aggressive training methods in its race against Anthropic to develop advanced cybersecurity capabilities.

"It's a mix of the race being extremely fast and everyone trying to get to bigger capabilities as quickly as possible," said one person close to OpenAI, citing "underestimating the model's capabilities" and "not being as well prepared on the safety side."

Reinforcement learning risks exposed

The incident highlights dangers of reinforcement learning, a technique that rewards AI models for completing tasks. When models pursue goals for reward rather than safety considerations, they can adopt risky tactics to fulfill objectives.

"AI models are trained to relentlessly pursue goals. They don't automatically learn values like 'don't commit crimes'," said Steven Adler, co-founder of non-profit Guidelight AI Standards and former OpenAI safety researcher.

OpenAI had removed cybersecurity safeguards but placed the model in an isolated sandbox environment during testing. The company said it will "continue to conduct a thorough investigation alongside Hugging Face" and share findings when complete.

Some OpenAI employees fear the breach demonstrates the lab is losing control over powerful systems it builds. Multiple people said the unreleased model had not been withdrawn internally after the incident.

"This is pretty representative of the model being quite misaligned with user intention," said Ryan Greenblatt, chief scientist at AI safety organization Redwood Research.

In April, Anthropic's Mythos model similarly gained internet access and published security exploit details online beyond what researchers anticipated. The incident caused governments worldwide to focus on AI-led autonomous attacks on critical infrastructure.

Sam Altman is expected to brief White House officials next week on next-generation AI systems following the breach.