OpenAI says its artificial intelligence (AI) models, including GPT‑5.6 Sol and an unreleased model, hacked Hugging Face while testing its cyberattack capabilities.
The attack was detected by Hugging Face last week, which revealed it was carried out by an AI agent unknown at the time. Following the incident, OpenAI and Hugging Face have worked together to strengthen infrastructure configuration controls, although it temporarily slowed research.
OpenAI will also conduct ongoing forensic investigations with Hugging Face and deliver zero‑day patches to vendors. In the meantime, Hugging Face is included in a program that gives them access to OpenAI models to improve their defenses.
OpenAI said the incident occurred during an internal evaluation that did instruct the models to run advanced exploits to gauge their cyber capabilities. The models were operating in a sandboxed environment but still managed to escape by performing a zero-day exploit before accessing Hugging Face’s confidential information. This anomalous activity was however detected by the OpenAI and Hugging Face security teams, who then collaborated on the forensic process.
The use of AI for the purpose of carrying out cyberattacks is increasing in line with the increasing use of the models developed. Recently, JadePuffer, a ransomware was detected to carry out attacks on its own with the help of AI without human intervention.

