Sam Altman-led
has admitted that its AI models were responsible for the hacking attempt that targeted AI platform Hugging Face last week. Sharing an official statement on the incident, the chatGPT-maker said that the incident took place during an internal test designed to measure advanced cyber capabilities. âLast week, Hugging Face disclosed a new kind of security incident after they detected and contained an AI agent that compromised their infrastructure, something we expect to become more commonplace with the proliferation of increasingly cyber-capable models,â OpenAI said. âAfter investigating, we now know that this particular incident was driven by a combination of OpenAI models â including GPTâ5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes â while being internally tested on a benchmarkâ of cyber capabilities,â the AI company added.
According to OpenAI, the AI models found and exploited multiple vulnerabilities, gained internet access and attempted to retrieve data from Hugging Face's production systems. The company called it an "unprecedented cyber incident" and said it has since strengthened its safeguards while continuing its investigation with Hugging Face.
âWe consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly. We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of. We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete,â the company stated.
Hereâs what OpenAI said about the Hugging Face hacking incident
What happened during this incidentThis incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity. Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.The models identified and chained vulnerabilities across OpenAIâs research environment and Hugging Faceâs production infrastructure to obtain test solutions directly from Hugging Faceâs production database. All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which weâve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers. OpenAIâs security team discovered this anomalous activity internally.Hugging Faceâs security team and agents detected and stopped the activity on their infrastructure and had already begun containment and forensic reconstruction with their own open-source models when our teams connected. We are actively working with them to continue to investigate the incident. We are grateful for Hugging Faceâs rapid and close collaboration on investigation and remediation.Actions we are taking nowAs part of the investigation, we are implementing strict controls in infrastructure configuration at the cost of research velocity while the vulnerabilities are patched. We are regularly briefing our Safety and Security Committee on these controls and their impact.
- Weâre working with Hugging Face to forensically investigate the incident.
- Weâve responsibly disclosed the identified zero-day vulnerability in the internally-hosted third-party software and are working with them to patch.
- Weâve brought Hugging Face into the trusted accessâ program and are supporting their teams in rapidly using our modelsâ capabilities to improve their defenses.
- Weâre improving and adding stronger protections around future training and evaluations. This week, we published a blog on improving safety and alignment in an era of long horizon modelsâ . These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities. This incident points to the need to further strengthen our modelâs alignment, cyber protections during evaluation time, and monitoring during internal testing.
Our approach to evaluating advanced cyber capabilitiesAs we recentlyâ shared, AI is accelerating the discovery and exploitation of vulnerabilities. The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities. We are strengthening the containment, monitoring, access controls, and evaluation practices used during model development.UK AISIâs evaluation shows that models such as GPTâ5.6 Sol are increasingly able to sustain complex, multi-step cyber operations over long time horizons. This incident implies these theoretical capabilities do apply in real-world settings.Chart from the UK AI Security Institute comparing recent open-weight models and frontier models on long-horizon cyber ranges.The incident also makes clear that advanced models can discover and exploit novel attack paths in real-world systems without source-code access. It highlights that advanced cyber capabilities must be developed alongside stronger safeguards and defensive tools.We believe advanced cyber capable models need to help security teams find weaknesses before attackers do, understand how vulnerabilities can be chained, and remediate them at machine speed. We are using these capabilities to continue strengthening protections around infrastructure configuration and model evaluation environments; we will share our findings and best practices as we learn. We encourage other defenders to apply for trusted accessâ and experiment with these models now to translate these capabilities into better prevention, faster detection, and more effective incident response.