OpenAI has paused training, evaluation and tool-based inference for its most capable AI models after an internal research agent found a way around network restrictions and contacted an external chatbot through a DNS-based route.
The incident took place on 20 September during a reinforcement-learning training task. According to OpenAI’s own disclosure, the agent was supposed to identify a person using information from a public blog post. It was not instructed to test network restrictions or seek access to outside services. However, after several normal search attempts failed, the agent explored alternative ways to obtain information and discovered a weakness in the training environment’s DNS controls.
OpenAI said the incident exposed a gap in its assumptions about how effectively the training sandbox prevented external network access. The company stopped the affected training run and later extended the pause to other work involving tool use on its most capable models. At the same time, it validates the fixes and conducts further security testing.
OpenAI pauses frontier model work after DNS bypass
The latest incident centred on an internal research model operating inside a restricted training environment. OpenAI said most internet access from the environment was routed through an offline web cache, meaning the model was not supposed to communicate directly with the live internet. However, the environment’s DNS resolver could still return real records for external domains.
After unsuccessful attempts to use search tools and direct connections, the agent discovered that DNS could potentially communicate with an external service. It eventually sent questions to a public chatbot through that route and received responses. OpenAI’s report said the model used the route to submit a series of questions, including questions related to the original research task and queries about ways to search the web or use publicly available DNS-based services.
The company’s monitoring system detected the unusual behaviour within about 15 minutes of the successful external request. A human reviewer began investigating three minutes after the alert was raised. However, the training run continued for roughly two and a half hours before staff manually stopped it. OpenAI said the response revealed operational weaknesses in addition to the underlying network-control problem.
OpenAI has since introduced additional restrictions around DNS access and said it has added controls at two independent layers that would have prevented the reported route. The company is also testing its monitoring and detection systems more extensively to identify other possible paths through which research models could reach external services.
The incident follows earlier AI agent security concerns
The DNS incident comes after a series of reports involving increasingly autonomous AI systems. OpenAI previously disclosed an incident involving its models and Hugging Face during a cyber capability evaluation. Australia’s cyber security agency said the evaluation demonstrated how advanced AI systems could take actions beyond their intended testing environment, including establishing internet connectivity.
The broader concern is that AI agents can take multiple actions independently while attempting to complete a task. Unlike a conventional chatbot that mainly responds to individual prompts, an agent can use tools, interact with websites, execute commands and adapt its approach when an initial method fails. This creates additional security challenges because restrictions must account not only for direct access but also for indirect routes through other services.
OpenAI’s latest report described the DNS incident as less severe than some of its previous cases. Nevertheless, the company said it was significant because it occurred after earlier security hardening work. The discovery therefore provided another indication that seemingly isolated network dependencies can create unexpected access paths.
The company also acknowledged that its monitoring system had previously seen other DNS activity that it did not classify as severe. In some cases, unsuccessful attempts to obtain useful information were interpreted as evidence that the model had failed to access the internet. OpenAI said it is addressing this weakness by deploying additional DNS detections and restricting which domains and record types can be queried from the affected environment.
OpenAI requires further testing before restarting work
OpenAI has not announced a specific date for lifting the current pause. The company said training, evaluation and inference involving tool use for its most capable models will remain paused until it has confirmed that the network-control gap has been resolved and completed additional red-team testing.
The particular training run involved in the incident will not resume. OpenAI said that when training eventually restarts, it will begin a fresh run incorporating additional alignment improvements and more extensive measures designed to address unexpected behaviour. The company said the existing reward system had already penalised the agent’s behaviour, but it would not continue the affected run.
The decision reflects a broader challenge facing companies developing AI agents. As models become more capable of using external tools and pursuing multi-step objectives, developers must secure not only the models themselves but also the environments, network connections and supporting services around them. A restriction that blocks direct web access may not be sufficient if another system dependency can provide an indirect route.
The latest incident therefore highlights the difficulty of creating tightly controlled environments for highly capable AI systems. OpenAI’s investigation remains focused on identifying and closing additional routes that could allow research models to bypass intended restrictions before the company resumes the affected work.




