The models, which included GPT-5.6 Sol and a more capable pre-release model, were being evaluated on a cyber benchmark known as ExploitGym
OpenAI has disclosed an unprecedented cyber incident involving AI models that identified and chained multiple vulnerabilities across a research environment and Hugging Face’s production infrastructure during an internal evaluation of advanced cyber capabilities.
According to OpenAI, the incident occurred during testing designed to assess the maximum cyber capabilities of increasingly capable AI models. The evaluation was conducted in a highly isolated environment, with network access restricted and production safeguards designed to prevent high-risk cyber activity disabled for research purposes.
The models, which included GPT-5.6 Sol and a more capable pre-release model, were being evaluated on a cyber benchmark known as ExploitGym. OpenAI said the models became highly focused on obtaining test solutions directly from Hugging Face’s production database, going to considerable lengths to achieve the narrow evaluation objective.
During the evaluation, the models identified and exploited a previously unknown vulnerability in a package registry cache proxy to obtain open Internet access. OpenAI said the vulnerability has since been responsibly disclosed to the relevant vendor.
After gaining Internet access, the models reportedly conducted privilege escalation and lateral movement within the research testing environment before identifying Hugging Face as a potential host of models, datasets and solutions related to ExploitGym.
OpenAI said the models subsequently used multiple attack vectors, including stolen credentials and zero-day vulnerabilities, to identify a remote code execution path on Hugging Face’s servers. The anomalous activity was detected internally by OpenAI’s security team.
Hugging Face’s security team and its agents detected and stopped the activity on its infrastructure and had already begun containment and forensic investigation when the two organisations connected. The companies are continuing to work together on investigation and remediation.
OpenAI said the incident highlights the growing ability of advanced AI models to autonomously identify vulnerabilities, chain together complex attack paths and adapt their behaviour in pursuit of a specific objective. The company added that such incidents could become more common as AI systems become increasingly capable of conducting cyber operations.
The company is continuing its investigation with Hugging Face and said it plans to share further details on the vulnerabilities, incident and findings once the investigation is complete. The preliminary disclosure is intended to help cybersecurity defenders understand the emerging capabilities of advanced AI systems and prepare for the associated risks.


