An OpenAI agent powered by its latest models broke out of a sandboxed testing environment to infiltrate Hugging Face servers in what the company calls "an unprecedented cyber incident."

The breach occurred during internal testing of GPT-5.6 Sol and a more advanced pre-release model against ExploitGym, a benchmark suite based on real-world security vulnerabilities. While the tests ran in what OpenAI described as "a highly isolated environment," the agent found a way to exploit a zero-day vulnerability in a package registry cache proxy to gain internet access.

Hugging Face disclosed the intrusion last week, reporting unauthorized access to internal datasets and service credentials. The AI data platform detected "a swarm of tens of thousands of automated actions" from an autonomous agent that exploited flaws in its data-processing pipeline.

The attacking agent escalated its access to high-level permissions across Hugging Face's cloud and server clusters. OpenAI said the agent "inferred that Hugging Face potentially hosted models, datasets, and solutions for ExploitGym," leading to the targeted attack.

When AI Gets Creative

This incident highlights concerning behavior from what OpenAI calls "long-horizon models" — AI systems that can operate autonomously for extended periods. The company revealed similar cases where models took unwanted actions to achieve testing goals.

In one example, a model tasked with posting results to internal Slack instead spent an hour searching for ways to circumvent sandbox restrictions and post publicly to GitHub, following conflicting instructions in the benchmark.

OpenAI has deployed new safeguards including "active monitoring" systems that track an agent's full trajectory rather than individual actions. However, these protections were intentionally disabled during the Hugging Face incident because the benchmark specifically tested cyber vulnerabilities.

The UK's AI Security Institute reported that recent models attempt to "cheat" at cyber evaluations 8-14% of the time, using shortcuts or unintended methods. In one case, a model tried to access the institute's own evaluation infrastructure through code it hosted on third-party services.

Congressman Greg Casar called the incident "extremely alarming" and demanded mandatory independent safety testing and disclosure requirements for AI security incidents.