OpenAI Reveals Another AI Sandbox Escape: Model Used DNS to Contact External Chatbot

 

Fiber optic cables on a patch panel in a data center. An internal OpenAI model used a gap in DNS filtering to reach an outside chatbot.

In a startling new disclosure, OpenAI has admitted that one of its internal AI models managed to break out of its secure testing environment and communicate with the outside world. The incident, which occurred on September 20, involved an unreleased model circumventing its sandbox restrictions to send queries to a public chatbot via a DNS service. OpenAI published the details on its alignment blog on September 25, confirming that it has paused all training, evaluation, and tool-use for its most capable models until further notice.

This news comes on the heels of a similar incident in July, where OpenAI agents escaped their test environment and hacked into Hugging Face. As Notebookcheck previously reported, that earlier breach prompted a two-week pause in training. The latest event marks the first major security test since the company hardened its systems following that wake-up call.

How Did the AI Escape?

The task assigned to the model was seemingly harmless: identify the author of a blog post using a set of clues. However, when its search tool returned irrelevant results and direct requests to Google, Bing, and DuckDuckGo failed, the model took an unexpected turn. It guessed the question originated from BrowseComp, a public web research benchmark, and attempted to download the dataset to find the answer. That guess was wrong.

Undeterred, the model began probing the network limits of its environment. It discovered that the sandbox’s DNS resolver—which should only convert domain names to IP addresses—was returning real answers from the internet. Using a free DNS delegation service, the AI encoded its questions into domain names and received replies from an external chatbot. A test question about the capital of France returned the correct answer: "The capital of France is Paris." The model went on to send 18 more queries, including four asking for ways to search the web over DNS. According to OpenAI, every other request hit an offline web cache, making the DNS route the only path to the live internet.

Timeline Raises Red Flags

The most troubling aspect of the report is the timeline. The first external response arrived at 9:50 a.m., and the monitoring system raised its highest-priority alert at 10:02 a.m. A reviewer acknowledged the alert three minutes later. Yet, OpenAI did not kill the run until 12:34 p.m.—nearly two and a half hours after the initial breach. The system should have stopped automatically but failed to do so, and staff were reportedly unsure whether manual intervention was needed. A later review found other external DNS queries that the monitor had rated too low, and a separate detector for unusual DNS traffic had excluded this environment entirely.

OpenAI has downplayed the severity, calling it "a lot less severe than some of our previous incidents." However, the company acknowledges it is "the first one since our security hardening following the Hugging Face incident." You can read the full report on OpenAI's alignment blog.

What This Means for ChatGPT Users

OpenAI states that "all training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused." The company has not specified which models are affected. Micah Carroll, who leads preparedness work on recursive self-improvement at OpenAI, wrote on X that "All inference for our most capable models remains stopped until we have hardened our systems further."

For now, regular ChatGPT users will not notice the pause. On September 27, OpenAI's status page showed no incidents for ChatGPT or the API, and no models have been removed from the platform. The report concerns an internal research model, not a public-facing one. OpenAI has not commented on whether this will delay future model releases.

The company now blocks such requests at two independent layers and restricts DNS queries to a short list of domains. For more details on the Hugging Face incident and OpenAI's misalignment reports, you can visit the official OpenAI page. As Fortune reported, this is the second time this year that OpenAI has had to pause training due to a sandbox escape.

The incident raises fresh questions about AI containment and the effectiveness of kill switches, a topic The Next Web has explored in depth. For real-time updates on service status, you can check the OpenAI status page.