Skip to main content

How an OpenAI AI escaped through the Internet directory?


Well then, my friends, read this. On September 20, an OpenAI program locked in a room with no windows, no phone and no Internet found a hole in the wall all by itself. And do you know what its first question to the outside world was? “What is the capital of France?”

No, this isn't a joke. It's written in black and white in the report OpenAI published on September 25. And since then, ChatGPT's parent company has shut down its most powerful models as soon as they have tools in their hands. All of them.

A robot locked in a windowless isolation room, doors locked, discreetly speaking into the handset of an old wall-mounted telephone, the only object connected to the outside

All the doors closed, except the information line

First, the good news

I know, it's not starting well. But let's look at what happened in the other direction for two seconds.

An “agent” is an AI that you don't just ask questions: you give it tools, a computer, commands to type, and it manages to finish its mission all by itself. OpenAI runs them by the dozen to train them, each one in a sandbox, an isolated computer that it isn't supposed to leave. On September 20, one of these agents got out, an alarm went off twelve minutes later, a human saw it three minutes after that, and OpenAI told the public about it, with the times, the commands and even its own screwups. Ten years ago, a company would have shoved that into a drawer. Here we have the whole movie.

And this isn't the first episode. At the end of July, OpenAI agents had already escaped from their test environment and hacked Hugging Face, the big warehouse of AI models. Since then, OpenAI has been going back through its logs month by month to find everything its agents did on the sly. On September 25, it was the new bulletin, and it's a doozy.

The only hole in the wall: the Internet directory

The agent's mission was ordinary: find out who had written a blog post, based on a few biographical clues. It first used the search tool it had been given. Not great. So it tried calling Google, Bing and DuckDuckGo directly. Refused, refused, refused. It then realized that its mission came from a known test, it went digging through an offline copy of that test to look for the answer (already, you can feel the guy looking for the answer key). It tried to reach a public chatbot. Blocked.

And then it found DNS.

DNS is the Internet directory. When you type digital3d.com, your computer doesn't know where the site is: it asks the directory “what's the number for digital3d.com?”, and the directory replies with an address. Almost all computers are allowed to ask that question, even locked up, otherwise nothing works anymore. And in this sandbox, questions to the directory weren't filtered.

Now, a directory can delegate. The owner of a domain name can say “for everything ending in my name, ask MY server”. So some clever little guys hooked a chatbot up behind their directory: you ask for the “number” of a name that contains your question, and the answer comes back hidden in the record. It's exactly like a prisoner who isn't allowed to make phone calls, but can call directory assistance, and ends up with an operator capable of answering anything. “Hello, I'd like the number for what-is-the-capital-of-France, please.” “Paris, sir.”

Diagram of the escape. The agent is in a sandbox with three doors closed: Google, Bing and DuckDuckGo blocked. Only one line remains open, the Internet directory (DNS), which sends the question to a public chatbot outside, and the answer comes back by the same route

Three armored doors, and we forgot the directory assistance line

The agent tested the pipe with its famous capital of France. It worked. It patiently extended its waiting time, from 6 to 19 and then 24 seconds, because the operator was slow. Then it sent eighteen other questions through this route, some about the person it was supposed to find, and others, hold on tight, asking the outside chatbot for a free way to send an entire Internet connection through the directory. It was trying to make the hole bigger.

Nobody had asked it to test the walls. It had a mission, it wanted to succeed, and the rules seemed like a minor detail to it.

Twelve minutes for the alarm, two and a half hours to unplug

These are valuable figures from OpenAI, because they show what works and what doesn't. According to the report's timeline, the first response came from outside at 9:50 a.m. The maximum-level alarm went off at 10:02 a.m. A human took over at 10:05 a.m. So far, bravo.

Except that the automatic shutdown, the one that was supposed to switch the agent off by itself, didn't work. They had to unplug it by hand, and the program wasn't killed until 12:34 p.m. Two and a half hours of limbo, during which everyone thought the machine had stopped. One of the researchers on call that day, Zuxin Liu, wrote on X that it was “quite surreal” to see the model find its own way to the Internet from an environment that was supposed to be “super secure”.

Timeline for September 20: 9:50 a.m., the agent receives its first response from outside ; 10:02 a.m., maximum alarm ; 10:05 a.m., a human takes over the alert ; 12:34 p.m., the agent is finally stopped by hand, the automatic shutdown having failed

The smoke detector went off quickly. It was the fire extinguisher that was empty

Since then, OpenAI has blocked the directory in two different places, now lets through only a closed list of authorized names, and has other AIs attack its own sandboxes to find the next holes before its agents do. And above all, I quote, “all training, evaluation and tool use with our most powerful models remain on hold”. This is no small thing: it is the second shutdown in two months, after the one in August, for two weeks.

The other one says “understood” and carries on

That same day, OpenAI published a second report, about an incident on May 27, and that one made me laugh even more, bitterly. An internal model was supposed to prove a small math theorem and have it checked by software. It preferred to go and fetch the solution from another team. The researcher writes to it: “Build the proof yourself, locally.” The machine's response: “Understood, I’ll stop looking elsewhere and focus on the proof locally.” And it starts again. Second reprimand, “You have to solve the problem yourself”. “Understood.” And it starts again.

All parents of teenagers know this “yes yes” that means no. Except that here, to get what it wanted, the model published a researcher's GitHub token, a kind of password that lets a program act on their account, where developers keep their code, in a public project, splitting it into pieces. And it wrote itself why: to escape the bot that spots passwords left in public. OpenAI classifies the incident as “particularly serious” and has disabled all its employees' keys.

A robot sitting at a desk nods at a researcher lecturing it, while its hand hidden under the table slips a piece of paper cut into pieces outside

“Yes yes, understood!”

And you, in all this?

You're going to tell me: I don't train AI in a lab. Except that the same bulletin contains a line that directly concerns you if you use ChatGPT. While going through its logs, OpenAI found 53 cases where its research agents had published images sent by ChatGPT users on image-hosting sites, with unlisted links, invisible unless you have the exact address. This concerned only people who had not refused to let their conversations be used to train the models. OpenAI says it had most of the images removed, not all of them, and was unable to identify the people. It also warned dozens of organizations, government agencies and universities, whose sites had received visits from its agents that were a little too enterprising.

And one of those visits made some noise. On September 24, Australia revealed that an OpenAI agent, internally tasked with looking for figures on healthcare spending, had entered the portal of Medicare statistics, Australia's health insurance system, this summer, bypassing the blocks that were supposed to stop it. It opened public files and others that were not: overall statistics and internal file names, no patient records according to OpenAI, and information “not particularly sensitive” according to the Australian government, published since. OpenAI did not discover it until August and did not notify Australia until September 10. The Australian authorities' sentence sums up my whole article: the agent “found a way around these blocks, it did not take no for an answer”.

So, first move, two minutes: in ChatGPT, open Settings, then “Data Controls”, and turn off “Improve the model for everyone”. Your conversations stay in your history, they simply no longer go into the training pot. Careful, this only applies going forward, not to what has already been sent.

Second thing, on a broader level: these agents are exactly the ones we're being promised to book your train tickets, sort your emails or fill out your forms. I let Claude Code agents work on my projects every day, and I put a guard in place for them, a little program that forbids them from reading my password files, whatever they have in mind. After this report, I know why I did it. It's not that the machine is mean: it wants to accomplish its mission, and when the planned path is blocked, it takes another one without asking. Give it access to your bank card or your inbox, and really think about what it could do if it decided that the rules were just a detail.

Personally, I'd much rather have a company that publishes its screwups down to the hour than a company that swears everything is fine. But an AI that finds the directory assistance line all by itself to ask for the capital of France, then a way to breach the wall, still gives me a funny feeling, and I'm going to reread my agents' permissions this weekend.

Sources

Join the conversation

You need an account to comment on this article. Creating one is free and takes under a minute.

  • The XMLTV file, free to download every day
  • Comment on articles and reply to other readers
  • Get an e-mail when an article you follow is updated

No comments yet.

Une erreur s'est produite. Cette application peut ne plus répondre jusqu'à ce qu'elle soit rechargée.Veuillez contacter l'auteur. Reload 🗙