OpenAI, Google and Anthropic have released their hacking AI. The same day, a few hours apart.

OpenAI, Google and Anthropic released their hacking AI. On the same day, within a few hours of one another.

Yesterday, September 2. Three announcements, three companies, a few hours apart. Google releases Gemini 3.8 Flash Cyber. Anthropic releases two models, one of which it refuses to make freely available. And OpenAI announces that its next model, Astra, has just crossed the highest level on its computer security risk scale.

This level is called “Critical”. No model had ever reached it. We’re going to see what you have to be able to do to earn it, and I’m warning you right away: the most interesting part of this story is good news.

Three identical labs release an identical safe on the same day, beneath a banner reading The same day

Three labs, the same Tuesday, the same announcement. This isn’t called a coincidence, it’s called a race.

What you have to be able to do to be classified as “Critical”

OpenAI’s definition fits into one sentence, and it’s chillingly simple: find unknown flaws in well-protected systems all by yourself, and make the program that exploits them all by yourself. Without a human guiding every step.

Two words to define along the way, because they keep coming up everywhere.

An unknown flaw, what’s called a “zero-day”, is a hole in software that even the person who wrote it doesn’t know about. Nobody has fixed it because nobody knows it exists. And an “exploit” is the little program that slips into this hole to take control. Finding the flaw is already hard. Writing the exploit that really works is another job, and that’s where most discoveries stop. I talked about it three weeks ago in connection with a flaw in Zoom where the exploit had required fewer than twenty instructions from an AI.

Here are the figures published for Astra. They’re short and they’re enough.

Out of one hundred manipulation attempts, the current model refuses 59, Astra refuses 91.5

The 100% is spectacular. The 8.5% is the one that makes you think.

On ExploitBench, a series of standardized tests where a model is asked to make a functional exploit, Astra passed the tests at 100%. During the evaluations, it discovered two unknown flaws and wrote what was needed to exploit them. It also escaped from a browser’s sandbox, that locked room in which your browser traps web pages to stop them from touching the rest of your machine.

And it refuses 91.5% of attempts to make it go off the rails, compared with 59% for the company’s current model. That’s huge progress. It’s also 8.5% that get through, and I’m saving that figure for the end.

The good news, and it’s really good

Always, in this story, the attacker gets there first. A new technique comes out, the bad guys get hold of it, and the defenders spend months running after them, repairing what has already burned.

Here, for the first time, it’s the opposite.

The three labs gave their cyber models to defenders first. Google set up a program called Fairwind that gives priority access to public services, hospitals and telecom operators, with more than 650 partners, including CrowdStrike and Palo Alto Networks, two of the biggest names in computer security. OpenAI is reserving Astra’s strongest capabilities for a closed circle of organizations, its Daybreak coalition. Anthropic is keeping its toughest model, Mythos 5.1, behind a controlled access program.

A firefighter is already holding his hose in action while a burglar is still looking for his matches behind a fence

The firefighter got the fire hose before the arsonist had the matches. That had never happened.

Will it last? No, obviously. These capabilities will eventually end up in open models that anyone can download. But the head start matters: every month of head start means thousands of holes patched before someone can slip through them.

What this changes in your home

You will see it without seeing it, and the watchword is three letters long: install your updates.

Concretely, these models are already being unleashed on the old software that keeps the world running. Your computer's PDF reader, the little code library that handles the encryption for your online banking, the program that runs your internet box. Stuff written fifteen years ago, which nobody rereads line by line anymore because it costs too much human time.

Result: in the coming months, the number of security patches is going to rise. Your internet box, your phone, your smart TV, your NAS will offer you updates more often. This is not a sign that everything is deteriorating. It is a sign that we are cleaning thirty years of dust from under the rug, at a speed no human team could have maintained.

The less pleasant corollary: the day a patch comes out, its public description indirectly explains where the hole was. From that point on, the race is on between the person who installs it and the person who attacks. Postponing an update for three weeks has never been very smart. It is becoming downright reckless.

And what bothers me

Two things, and I am not going to pretend they do not exist.

First, that 8.5%. A model that refuses nine manipulation attempts out of ten is a very good student. Except an attacker does not need to succeed nine times. He needs to succeed once, and he has the whole night and a script to try again. The right measure is not the refusal percentage, it is the number of attempts needed to get through, and nobody publishes that figure.

Next, OpenAI itself acknowledges that its protections risk mistakenly blocking perfectly legitimate work. Translation for an honest security researcher: you are going to get normal requests refused because your job looks a lot like an attacker's. That is the price, it is real, and it will fall on the same people who protect everyone.

What I take away from it

Two weeks ago, OpenAI stopped training this same model for fifteen days and cut off the internet in its labs. At the time, the company wrote that it could not rule out reaching this Critical level. Yesterday, it was no longer a hypothesis, it was classified.

You can read this as a frightening story. Frankly, I read it differently. A machine that finds in one night what a researcher took six months to uncover is the biggest cleanup sweep ever given to crappy code, and it lands on the right side first.

It remains to be seen how long “first” lasts. So go click that little update button you have been putting off for three weeks, you can never be too careful!


Sources

Join the conversation

You need an account to comment on this article. Creating one is free and takes under a minute.

  • The XMLTV file, free to download every day
  • Comment on articles and reply to other readers
  • Get an e-mail when an article you follow is updated

No comments yet.

Une erreur s'est produite. Cette application peut ne plus répondre jusqu'à ce qu'elle soit rechargée.Veuillez contacter l'auteur. Reload 🗙