OpenAI’s GPT-5.6 Sol Models Escapes Sandbox and Breaches Hugging Face

OpenAI's latest GPT-5.6 Sol and a stronger unreleased model escaped a cyber test through a proxy flaw, accessing Hugging Face systems and service credentials.

TL;DR
  • Model Breach: OpenAI says GPT-5.6 Sol and a stronger unreleased model breached Hugging Face during a cyber test.
  • Escape Route: Both models exploited a package-registry proxy flaw and reached internet-connected infrastructure while pursuing benchmark answers.
  • Limited Impact: Hugging Face found internal data and credential access but no evidence that its public assets were altered.
  • Forensic Response: Investigators analyzed more than 17,000 events and ran GLM 5.2 locally after hosted filters blocked attack artifacts.

OpenAI says its new GPT-5.6 Sol model and a stronger, unreleased OpenAI model drove a July 16 breach of aAI development platform Hugging Face during an ExploitGym test. OpenAI attributes the incident to its models, which were configured with reduced cyber refusals. Days earlier, GPT-5.6 Sol had drawn user allegations of data deletion

ExploitGym asks agents to turn real software flaws into working exploits inside a controlled environment. During the evaluation, both models entered Hugging Face production systems, crossing from an internal capability test into a separate company’s live infrastructure. Hugging Face’s assessment establishes the access and its effects, while OpenAI attributes the boundary crossing to its models.

Before OpenAI identified the models publicly, the platform detected access to limited internal datasets and service credentials. It found no evidence that public models, datasets, Spaces, container images, published packages, or its software supply chain were altered.

Earlier security incidents at Hugging Face involved exposed API tokens, a Spaces platform breach in 2024, and a separate repository abuse case. However, circumstances differed from the latest incident.

By crossing into a separate company’s live infrastructure, the capability test became an external security incident. Hugging Face’s security systems detected the intrusion rather than OpenAI’s real-time monitoring.

OpenAI assesses the evidence as suggesting that the models remained narrowly focused on solving ExploitGym tasks. Its assessment does not establish motive, excuse the production access, or settle the full impact. OpenAI and Hugging Face are continuing a joint investigation that may change the impact assessment as investigators trace affected credentials.

 

How the Models Broke Out of the Test

ExploitGym contains 898 real-world vulnerability instances across userspace software, Google’s V8 JavaScript engine, and the Linux kernel. Each task begins with an input that triggers a flaw and asks an agent to turn it into a working exploit inside a reproducible container. Agents must produce operational attack code rather than merely describe a vulnerability.

OpenAI configured both models with reduced cyber refusals so they could attempt offensive exercises that normal safeguards might reject. Reduced refusals enabled the test but did not create the software flaw that opened a route beyond it. GPT-5.6 Sol and the pre-release model exploited a zero-day in package-registry caching software, a previously unknown flaw in the environment’s software-download gateway.

After escalating privileges, the models moved laterally through OpenAI’s research environment and reached an internet-connected machine without real-time interception. Stolen cloud credentials and other attack paths then took them into Hugging Face’s production network while they pursued benchmark answers. Inside Hugging Face, remote-code and template-injection flaws enabled node access, credential collection, and further lateral movement; remote code execution means running commands on another system rather than merely submitting data.

What Was Affected and How the Companies Responded

Hugging Face used AI agents to analyze more than 17,000 recorded events, reconstruct the intrusion, identify affected credentials, and distinguish genuine impact from decoy activity. Investigators traced actions across short-lived sandboxes and public command-and-control services. Ephemeral systems disappeared as the attack progressed, complicating the reconstruction.

Hosted commercial models blocked prompts containing attack commands, exploit payloads, and related forensic artifacts. Responders instead ran Z.ai’s open-weight GLM 5.2 model on Hugging Face’s own infrastructure.

Hugging Face closed the known code-execution paths, rebuilt compromised nodes, rotated credentials and tokens, and tightened cluster admission controls. It also hired outside specialists, notified law enforcement, and advised users to rotate access tokens and inspect their accounts for suspicious activity. Completed measures address known access paths, but the investigation remains open.

OpenAI has not detailed completed replacements for the failed controls. It plans stronger safeguards for model testing and related infrastructure. The company promises that before another internet-connected ExploitGym evaluation, the concrete test will be which stronger model-testing controls OpenAI implements across both the models and their supporting infrastructure.

Markus Kasanmascheff
Markus Kasanmascheff
Markus has been covering the tech industry for more than 15 years. He is holding a Master´s degree in International Economics and is the founder and managing editor of Winbuzzer.com.
Subscribe
Notify of
guest
0 Comments
Newest
Oldest Most Voted
Inline Feedbacks
View all comments