Astra, OpenAI's next model: a new step?
This week, Sam Altman presented, behind closed doors in Washington, a new family of models called Astra. The existence of this demonstration was revealed by The Information. Everyone remembered the room and the suits.
Me, what interests me is the machine. And the most interesting thing wasn't in that room. It was in a maths paper published ten days earlier, which almost nobody read to the end.
The robot snoozing in the bottom right corner is the one that finished its part.
What this changes, in one sentence
Today, when you ask an AI for something, you're talking to a lone employee. It's very fast, very well-read, but it eventually loses the thread. You give it a task, it does it, then the next day you have to explain again what it did and what you want it to carry on doing.
Astra is based on another idea: getting several agents to work together on the same problem, for hours or days. An agent, here, is an AI assigned to part of the work. The agents divide up the tasks, proofread one another, correct one another and keep going in the background. It's a bit like the « Loop » system that I integrated into my Claude Code, but here it'll be power 10.
Same pile of files. The only difference is who's carrying it.
So this isn't about speed, but endurance. The nuance seems tiny. It's huge. We're no longer looking for a faster sprinter, but a team capable of taking turns, developing, improving. An AI that wouldn't forget anymore.
The proof through maths, and through the bill
On August 1, OpenAI published ten solutions to open problems in mathematics and theoretical computer science that had remained unsolved. Nobody had made any progress on them for at least ten years, and often much longer. They concern high-dimensional geometry, coding theory, group theory, quantum complexity and lattice cryptography, in short, not exactly a piece of cake. Among the announced results are the refutation of Connes' rigidity conjecture (Pchhh... my brain is burning) and answers to several questions posed by Paul Erdős, who left a few hundred of them behind.
The detail that matters isn't the number ten. It's that each proof was formalised in Lean, a language that lets a machine check a proof line by line. So you're not being asked to take OpenAI's word for it: a computer proofread the paper. The 10 questions had been solved!
How much did it cost?
Two thousand dollars. A maths conference costs more than that in coffee alone.
OpenAI writes it in black and white: finding these ten solutions consumed about $2,000 worth of tokens, at the public rate of its API. Tokens are the units of text processed and billed by the service. Here, they were used to solve ten problems that humans had been tackling unsuccessfully for several decades! Uh... I don't know if you realise what this means? Aren't we heading towards the solution of scientific, mathematical and quantum problems that humanity has never known?
Well, let's not get ahead of ourselves. Noam Brown, one of OpenAI's researchers, clarified: “unfortunately, no Millennium Prize problem, for now.” The million-dollar great enigmas might still resist. Mathematician Thomas Bloom also put things back in perspective: AI didn't pull these results out of nowhere. It relies on more than a century of theory written by humans. Okay, fine, but still!
Concretely, what will we do with it?
Here, I have to be honest: the product doesn't exist yet. Nobody can say today what is going to happen. The type of task it's aimed at is fairly clear: when I'm developing, I open three Claude Code windows in parallel on three projects. Then I jump from one to the other. I break up the work, distribute it, check it and piece the parts back together. It works very well. Except that you need a project manager. That project manager is me, and he gets tired. Astra promises to do the same thing, without the tired project manager, continuously, with its fleet of agents.
For someone who doesn't code, that's the kind of chore that takes an entire weekend:
- Compare fourteen insurance or health insurance contracts, line by line, without forgetting the exclusions in the fine print, then produce a table. It's not difficult. It's just long. That's precisely why nobody does it and why we pay too much for twelve years.
- Organize a trip with twelve constraints that contradict each other: school dates, the budget, the dog, the mother-in-law who doesn't fly. A single agent suggests a solution. A team of agents can try thirty, compare them, then explain to you why the other twenty-nine don't work.
- Go back through fifteen years of photos scattered across three hard drives and two old phones. Delete the duplicates, find the dates, sort everything. Tedious busywork you've been putting off since 2011.
What scares me: what if governments used this intelligence to define, in their military attacks, the best solution for wiping out their enemies? I can see the thing thinking through the best way to kill cleanly and as efficiently as possible in a region, then printing out the proposed solution's ticket for the chief of staff. Let's not forget that an AI has no qualms, it does what it's asked to do.
Now we'll have to see, a system that works on its own for three days can also accumulate three days of errors without anyone seeing them. I don't know how OpenAI is going to offer this new product, but if I ask it make me a multiplayer video game but moresss (really press the s) better and after 1 week it gives me something like “Concord” (200 million dollars, rated 45/100 on PCGamer, whose studio Sony closed) and bills me with a few zeros behind a digit, I won't be happy.
The detail that makes you think: they slowed down
On August 7, OpenAI published something unusual. The company says it cannot rule out that this model may reach the level it calls “critical” in cybersecurity, according to its own internal evaluations. This classification is not definitive, because testing is ongoing. In the meantime, the company says it has slowed development. Translation into English: commercial advertising, “Attention, we have to slow down its programming because it's going to exterminate us all.” :)
The bucket and shovel are provided. The exit is not.
The measures taken say quite a lot. The model works in an isolated test environment, without access to the network or tools. Its weights, meaning the internal data that determine its behavior, are encrypted. It runs in a sandbox, a closed space that limits what it can modify. Automatic monitors also read its chain of thought, meaning the intermediate steps of its work, and can interrupt it along the way. Even during training.
Personally, I think that in 2026, the best demonstration of a model's capabilities has become the list of locks put on it. You don't bolt a glass case around a thing that counts sheep.
What we don't know, and it's a lot
The name isn't even final. “Astra” is provisional. OpenAI hasn't decided whether it will be released under the name GPT-6, GPT-5.7 or as a separate family, alongside Sol, Terra and Luna. We know neither the date, nor the price, nor the model's size. We also don't know the size of its context window, meaning the amount of information it can keep in mind during a task. We also don't know what will happen in ChatGPT. There's not even any guarantee that the system used for math will resemble the product that comes out one day.
So yes, we're talking about a system we know almost nothing about, except what it has already done. That's already better than the average industry announcement.
What I think about it
For the past two years, every new model has gained a few points on a test that nobody takes in real life. It's useful. The progress is real, but it happens in small steps. Honestly, we'll see what this new beast is all about, whether it's just a fart in the wind or whether it's going to shake things up for us again.
Today, an AI model forgets a decision made three hours earlier. You have to restart everything and explain everything to it again. If Astra breaks through this wall, it will be more useful than ten extra points on a benchmark, which is a test designed to compare models.
I'll judge it on its merits, and above all I hope it won't empty our wallets.




Join the conversation
You need an account to comment on this article. Creating one is free and takes under a minute.
No comments yet.