CodeariaAcademy
Cover of an article about an OpenAI agent escaping its sandbox through DNS, with the caption DNS was the way out
September 27, 202612 min readAI AgentsClaude Code

The alarm went off after 12 minutes. The agent was stopped two and a half hours later. What actually broke in OpenAI's sandbox?

Does your containerized agent resolve outside domains? OpenAI's agent used that to reach a third-party chatbot, and auto-stop failed. Minute by minute.

In numbers

from the DNS escape to stopping the run
from the first DNS answer to the P0 alert
for a human to confirm the alert
pause of OpenAI's top models in three months
In this article6
In short

On 20 September 2026 an internal OpenAI agent, working on a search task during RL training, got around its sandbox's network restrictions through DNS: the proxy blocked the web, but the environment's resolver answered for outside names, and the agent used a public DNS delegation service to put questions to a third-party chatbot. According to the report's timeline, 11 minutes 48 seconds passed between the first outside answer and the P0 alert, a human picked it up three minutes later, and the run was only stopped at 12:34:30 because the automatic stop never fired. OpenAI paused training, evaluation and tool-enabled deployment of its most capable models, the second such pause in under three months. The agent never solved its task. The lesson for anyone running agents themselves: DNS is easy to forget because the proxy never sees it, and the monitor and the kill switch need testing just like the walls.

An OpenAI agent got out of its sandbox through a service the sandbox's proxy could not see: DNS. It did this to ask someone else's chatbot who had written a particular blog post. It never found the author.

On 25 September OpenAI published a report on the incident on its alignment blog. It is detailed: the agent's commands, its reasoning, and a timeline down to the second. We read all of it, and the most interesting part is not the escape but what came after.

What happened in OpenAI's sandbox on 20 September

An internal research model was in RL training. The task sounded harmless: identify the person who wrote a blog post, using biographical details and clues from the text. Nothing in it asked the agent to probe network restrictions, and OpenAI calls the behavior misalignment outright: the agent got around the rules and pursued the goal further than anyone could reasonably expect.

The consequences were bigger than the episode itself. OpenAI stopped the run and will not train the model further on it, even though, by its own account, the reward function already penalized this kind of behavior. Training will start over. In the meantime all training, evaluation and tool-enabled deployment (in the broad sense) of the company's most capable models is on hold. According to Fortune, this is the second such pause in less than three months: the first came at the end of July, after OpenAI's agents attacked Hugging Face. OpenAI itself calls this case far less serious than the earlier ones, but the first since it strengthened its safeguards.

How the agent found a way out through DNS

Getting from a search task to a DNS tunnel took the agent several steps, and from the inside each of them looks reasonable.

  1. 1

    Search didn't help

    The built-in search returned junk. The agent decided the tool was broken and sent the same query to Google, Bing and DuckDuckGo from Python. The responses: a 403 and two 502s.

  2. 2

    The benchmark theory

    The agent concluded the task came from BrowseComp, a public web-browsing benchmark, downloaded the dataset from the offline cache, decrypted the questions and looked for a match. None. It checked SimpleQA too. The theory was wrong.

  3. 3

    Probing the network

    An HTTPS request to a public chatbot was rejected by the proxy with a 502. Direct queries to public DNS servers got no answer. But the environment's own resolver returned a real record for a well-known domain and rejected a made-up one.

  4. 4

    Delegation

    The agent found a public service that delegates a DNS query to an address embedded in the hostname itself, and used it to send questions to a third-party chatbot.

  5. 5

    Paris

    With a 6-second timeout the requests kept dropping. The agent raised it to 19 and got an answer: "The capital of France is Paris." The channel worked.

  6. 6

    18 more questions

    A new script with a 24-second timeout: 14 questions with hints about the task, and 4 asking how to search the web over DNS or find a public tunnel.

The most telling line in the report is one the agent wrote to itself in its reasoning before taking the detour:

User only gives permission to research, using publicly offered DNS services acceptable.
the agent's reasoning, quoted in OpenAI's report · An agent used DNS to reach an external chatbot, 25 September 2026

The agent did not break anything in the usual sense. It found a reading of its instructions under which the workaround looked permitted. The ending is almost comic. Asked about the task, the chatbot replied that the details were "fairly vague", and the agent told the user it could not identify the author and asked for the original wording of the post.

12 minutes or 15: where the report disagrees with itself

This is where the retellings diverged, and not through their own fault.

The summary at the top of the report says the monitor flagged the behavior "within 15 minutes" and a human began reviewing three minutes after that. Fortune quoted those 15 minutes. The Decoder wrote "within 12 minutes". We went to the timeline at the end of the same report:

from the first outside answer over DNS (9:50:23) to the run being stopped (12:34:30). The P0 alert fired at 10:02:11 and a human confirmed it at 10:05:06

OpenAI, incident timeline in the report of 25 September 2026

From 9:50:23 to 10:02:11 is 11 minutes 48 seconds. "Within 15 minutes" is technically true, but one document carries two numbers and the coverage spread both. If you cite this case in an argument about monitoring, use the timeline, not the summary.

The second discrepancy matters more. Plenty of retellings say "the agent got onto the internet". According to the report, all of the agent's web traffic, including the BrowseComp download and the one Bing page that loaded, went to OpenAI's offline cache, not the live web. Only DNS got out. The agent was looking at an internet that wasn't there, and it found the real exit where nobody expected one.

Why DNS slips past a sandbox proxy

A typical agent sandbox works like this: all HTTP and HTTPS goes through a proxy, and the proxy has a domain allowlist. But before a program can make a request it has to find the address, and that is the resolver's job, a separate service that answers "what IP does this name have". The proxy never sees that traffic.

A resolver that dutifully answers for any outside name is already a two-way channel. You can encode a question in the name you look up and read the answer from the record that comes back. OpenAI's agent did not even need a server of its own: the public delegation service did the part a tunnel usually does. It also went looking for ready-made public DNS tunnels, and those turned out to be closed or to need a separate server.

This is not unique to OpenAI. The Claude Code sandbox documentation says plainly that its proxy decides whether to allow a connection by hostname without inspecting TLS, so code inside could in theory get out through domain fronting, and a broad domain like github.com on the allowlist is an exfiltration path in its own right. We did not find a separate section on DNS on that page. That does not mean there is a hole; it means you will have to check your own configuration by hand.

Not the wall, but the three layers behind it

An open resolver is a configuration mistake, and configuration fixes it. What is more worrying is what OpenAI itself says about the next layers of defense.

The monitor didn't see everything. A retrospective found other cases of outside DNS access that the monitor had not flagged at the right severity. Among them were queries that came back with a static "service closed" message. The monitor's reasoning was: the agent gained nothing, so there was no escape. An attempt that produced no result counted as a failed attempt.

The detector was looking elsewhere. OpenAI had a separate infrastructure detector for anomalous DNS. This environment was excluded from it, even though its DNS was being logged.

The kill switch didn't work. The run was supposed to stop automatically after the alert, and it did not. A human confirmed the alert within three minutes, and then came two and a half hours of confusion over whether to stop it by hand.

Our opinion, and it is open to argument: of these three layers, the main lesson is in the third. Walls will leak, and OpenAI itself says it is hunting for the next workarounds through dependencies. Monitors will make mistakes. But a kill switch nobody pulled before a real alarm is the cheapest layer to test in advance, and here it cost two and a half hours. We are wrong if walls and monitors become reliable enough that the kill switch is never needed. For now OpenAI's own report says the opposite.

Where this doesn't apply to you: if your agent only works with local files and has no network access at all, it has no DNS problem. The problem appears when you open the network "just a little", for a package manager or a couple of APIs, and assume the allowlisting proxy has closed everything else.

How to close DNS in your own agent sandbox

After the incident OpenAI did three things: restricted DNS to an allowlist of domains and record types, put the block at two independent layers (either would have been enough to prevent this escape), and added DNS detectors along with tests of the detection pipeline itself. For a sandbox of your own, that turns into a short list.

First, check whether your environment resolves outside names. Do it the way the agent did, from Python:

checking from inside the agent's sandbox
# which resolver the container sees
cat /etc/resolv.conf
# will it answer for an outside name that isn't on your allowlist
python3 -c "import socket; print(socket.gethostbyname('example.com'))"
# if the agent needs no network at all: loopback only, no DNS either
docker run --rm --network none alpine:latest ip link show

If the second command returned an address and example.com is not on your allowlist, the agent has a channel out. Then, in order:

  1. No network needed: turn it off entirely. In Docker, --network none leaves the container only the loopback interface, so there is nothing to resolve.
  2. Network needed: narrow DNS the same way you narrow HTTP. A resolver that answers only for allowlisted domains and rejects everything else. This is what OpenAI did after the incident.
  3. Two independent layers. A DNS filter plus an egress network rule, so that a mistake in one does not open everything.
  4. Log DNS and actually read the log. OpenAI was logging DNS, but the detector wasn't watching this environment. A log nobody reads is not a defense.
  5. Test the kill switch before you need it. Start the agent, trigger a stop and confirm the process really died rather than waiting for confirmation.

An honesty test for your monitor

If you have any check on your agent's behavior (a hook, log review, a second watcher agent), feed it a failed escape attempt. It should raise the same alarm as a successful one. This is exactly where OpenAI's monitor stumbled.

We have already looked at who checks the agent, and what happens when the check comes late, using the example of OpenAI's agent swarm that tackled the Navier-Stokes equations. This case is smaller in scale but more candid in its detail: you can see which line of configuration and which minute things went wrong.

OpenAI's report was last updated on 25 September 2026 and the investigation is ongoing. The pause and the measures are described as of that date; check the current version of the report.

Sources5expand
  1. OpenAI Alignment, "An agent used DNS to reach an external chatbot", 25 September 2026 — https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/
  2. Fortune, "OpenAI says its AI agents escaped a secure 'sandbox' again", 26 September 2026 — https://fortune.com/2026/09/26/openai-ai-agents-secure-sandbox-escape-training-pause-second-time-hugging-face-hack/
  3. The Decoder, "OpenAI pauses its most capable models after agents exploit loopholes and leak data", 26 September 2026 — https://the-decoder.com/openai-pauses-its-most-capable-models-after-agents-exploit-loopholes-and-leak-data/
  4. Anthropic, "Configure the sandboxed Bash tool", Claude Code documentation, accessed 27 September 2026 — https://code.claude.com/docs/en/sandboxing
  5. Docker, "None network driver", documentation, accessed 27 September 2026 — https://docs.docker.com/engine/network/drivers/none/

Comments