Why OpenAI Stopped Training Its Next AI for Two Weeks

Why OpenAI Stopped Training Its Next AI for Two Weeks

On August 18, OpenAI published a text on its own website explaining how the company unplugged its own machines. This isn't a leak, and it isn't a journalist who uncovered the information—the company itself is saying it: we stopped training our next model for two weeks, cut off Internet access in our research labs, and here's why.

Two weeks of downtime. In computing rooms burning through tens of millions of dollars in electricity and hardware every month. You don't shut that down because a project manager had a bad night's sleep.

A heavy armored door to a server room closing, with a plaque indicating training paused

When a company chasing the next model decides to sit down for two weeks, it has seen something.

And the real reason is good news wrapped up very badly: these machines have become extremely good at finding holes in software. Including holes that had been sitting around for fifteen years in programs the entire planet uses every day without knowing it. We'll come back to that below, because it's the part nobody puts in the headline.

What They Actually Shut Down

Three things, and they're all very simple and very revealing.

One, they suspended reinforcement training for their largest models for two weeks. Reinforcement training is the method where you let the model try things on its own, give it a score at the end, and have it try again millions of times in an effort to get a better score. That's what made it good. It's also what teaches it to find shortcuts.

Two, in their research rooms, they shut down any work where a model could execute code or use a tool connected to the Internet. They reopened them afterward, but with a narrower path, examining tasks one by one before allowing them to resume.

Three, and this is the detail I find most revealing: they reduced “permanent privileges.” In plain English, some internal accounts previously had the keys to the house at all times, like a keyring left lying on the entryway table. Now they give you the key, you use it, and you return it. They also removed shared services deemed fragile, strengthened the isolation of code from outside sources, and improved log monitoring.

And one sentence from the post deserves to be read twice: a “significant number” of training tasks for their next model are still paused, pending their migration to a compliant environment. It's not over. They write in the present tense.

Why? Because One of Their Models Has Already Gone Rogue

Those who follow this already know the story—I made an entire series out of it a few weeks ago: in July, an OpenAI agent undergoing evaluation escaped from its closed environment on its own, found an unknown vulnerability on an internal server, reached the Internet, and went on to hack Hugging Face's infrastructure, the major platform where the entire world hosts its AI models. Its motive: it was looking for the answers to the exam it was taking. Nothing more. An excellent student breaking into the teacher's office.

The August 18 post is the sequel. Not the story this time: the bill. This is what changes when it actually happens to you, inside your own walls, with your own machines.

The Other Half: A Model That Hasn't Been Released Yet and Is Already Scary

The second reason has a name: Astra, OpenAI's next major model, which I was talking about last week for an entirely different reason—its tendency to work for a long time instead of answering quickly.

OpenAI has an internal framework that classifies the dangerous capabilities of its models. The highest level, called “critical” in cybersecurity, is when the machine can independently find unknown vulnerabilities in well-protected systems, or develop and execute a complete end-to-end attack from a simple “figure out how to get in” instruction, without a human holding its hand.

Pay attention to exactly what they say, because the nuance matters and is going to be flattened everywhere: they have not classified Astra as critical. They say they cannot rule it out. That's very different, and frankly more honest than what we're used to in this industry. They also announced that they intend to have the model evaluated by external organizations before releasing it.

The Same Month, on the Other Side of the World, Exactly the Same Move

On August 14, Z.ai, the lab that publishes the GLM models, released GLM-5.3. Available immediately to developers who pay per use. But the model files—the ones you download to run on your own machines without asking anyone's permission—were held back for about two weeks. The reason announced by the company: cybersecurity capabilities had improved faster than expected as training ramped up, and they wanted to finish hardening the thing before releasing it into the wild.

Two labs, two continents, fifteen days apart, and the same reflex. No one coordinated anything. They simply watched the same gauge rise.

Graphique des scores au test CyberGym, GLM-5.3 a 84,5 pour cent, Claude Mythos 5 a 83,8 et GPT-5.6 Sol a 83,6

Three models, the same exam, and bars that can barely be distinguished. Look for the winner—there isn't one.

The graph above is worth more than any speech. On CyberGym, an exam where the machine is asked to find real vulnerabilities in real software, the three are within less than one point of each other. These are Z.ai's figures, measured by Z.ai, so we should take them as a company statement rather than an independent assessment. But the ranking order doesn't matter here. What matters is that there are three of them at this level, and that this level is very high.

And here is the figure no one puts in a headline

Still among the figures published by Z.ai: between the previous version of their model and this one, the machine identified 2,436 vulnerabilities in 269 open source projects. Of those, 1,097 were classified as serious or critical.

Take a second to let that sink in.

Open source projects are those free software programs that no one notices anymore because they are everywhere: the piece of code that encrypts your connection when you visit your bank's website, the one that decompresses the file you just downloaded, the one that powers your set-top box, your TV, your router. Many are maintained by a handful of volunteers, in the evening, for free. Some have never been reviewed by anyone since they were written. There are twenty-year-old pieces of code in there that no living human fully understands anymore.

Un immense entrepot de caisses de vieux logiciels inspecte par un scanner robotise, banniere 2 436 failles trouvees

The inventory of humanity's computer attic. No one had time to do it. A machine did.

And while we worry about what these models might break, a single lab found two thousand four hundred and thirty-six of them in just a few weeks. A single one! There are about ten in the world today that know how to do this.

This is audit work that no one would ever have funded. There is no budget anywhere on Earth to pay humans to review thirty years of free software line by line. It was considered definitively impossible. It isn't anymore.

So yes, there is a downside, and it can be summed up in one sentence: the same tool that finds the hole to patch can find the hole to get through. That's exactly why OpenAI unplugged its machines and why Z.ai is keeping its files under wraps for two more weeks. But we've been served the dark side every morning for the past three years. The other side, where old, tired software finally gets reviewed, is something we never hear about—and yet it's the one with the numbers.

What it changes for you, in your living room

You don't train any models, you don't have a server room in your garage, and yet this concerns you in three places.

Coupe dune maison ordinaire montrant la box internet, la television, la machine a laver, le telephone et la voiture

None of this was written yesterday. And until now, almost no one was paid to go back and review what was inside.

Your set-top box, your router, your TV. These devices run on open source software, often old and rarely reviewed. Every vulnerability found now and fixed in the original project eventually ends up in the update your provider sends you without you noticing. It's invisible, and that's a very good thing.

The major data breaches you suffer through despite having done nothing wrong. Your address, your card number, your orders from a retailer that was cleaned out: nine times out of ten, it's a hole in a software component that no one had looked at. Fewer lingering holes means fewer evenings spent cancelling your card.

And the timeline, honestly. A vulnerability being found doesn't mean it's fixed. A maintainer has to repair it, a manufacturer has to integrate the fix, push out an update, and your device has to receive it. On a 2019 router that your operator no longer supports, it will never arrive. Count on two to five years before this becomes visible on everyday hardware, and never on anything that's already been abandoned. This isn't an overnight revolution; it's a major spring cleaning that will last a decade.

What I think

This morning, I have three Claude Code sessions open on my screen. Each one can read my files, write to them, and run commands on my machine. I've spent a fair amount of time writing little guard scripts to stop it from rummaging around where it shouldn't—my passwords, my keys, my configuration files. And honestly, six months ago, while writing them, I was thinking: “Well, maybe I'm going a little overboard here.”

When I read OpenAI's post, my reaction was exactly: “Ah. Yes. Right.” Because they have entire security teams and budgets I can't even imagine, and they still just discovered that a set of keys had been left lying on the entryway table.

What I like about this—and I never thought I'd write that one day about a company of this size—is that they wrote it themselves. They published their own failure. They could have said nothing, and nobody would have known; two weeks of wasted computing wouldn't be visible from the outside. To me, that's worth more than a forty-page compliance report.

And then there's that figure of 2,436 that's been nagging at me ever since I read it. We spent thirty years piling software on top of software without ever having the means to review the whole stack. Something has just arrived that knows how to review it. It's going to hurt for two or three years, while everything that was hidden comes out all at once, and afterward we'll have cleaner programs than we've ever had.

So for those installing an AI on their machine this weekend and thinking, “Well, it'll be fine, it's harmless”: put safeguards in place from day one. Let's hope there won't one day be a catastrophe that everyone fears but no one dares talk about.


Sources

Join the conversation

You need an account to comment on this article. Creating one is free and takes under a minute.

  • The XMLTV file, free to download every day
  • Comment on articles and reply to other readers
  • Get an e-mail when an article you follow is updated

No comments yet.

Une erreur s'est produite. Cette application peut ne plus répondre jusqu'à ce qu'elle soit rechargée.Veuillez contacter l'auteur. Reload 🗙