Astra, OpenAI's next model: a new step forward?
This week, Sam Altman presented, behind closed doors in Washington, a new family of models called Astra. The existence of this demonstration was revealed by The Information. Everyone focused on the room and the suits.
What interests me is the machine. And the most interesting part wasn't in that room. It was in a math paper published ten days earlier, which almost no one read to the end.
The robot snoozing in the bottom-right corner is the one that has finished its part.
What this changes, in one sentence
Today, when you ask an AI to do something, you're talking to a lone employee. It's very fast and highly educated, but eventually it loses track. You give it a task, it does it, and then the next day you have to explain again what it did and what you want it to continue doing.
Astra is based on a different idea: having several agents work together on the same problem, for hours or days. An agent, here, is an AI assigned to handle part of the work. The agents divide up the tasks, review one another's work, correct each other and continue working in the background. It's a bit like the “Loop” system I integrated into my Claude Code, but here it will be 10 times more powerful.
Same stack of files. The only difference is who's carrying it.
So this isn't about speed, but endurance. The nuance may seem tiny. It's enormous. We're no longer looking for a faster sprinter, but for a team capable of taking turns, developing and improving. An AI that would no longer forget.
The proof is in the math—and the bill
On August 1, OpenAI published ten solutions to open problems in mathematics and theoretical computer science that had remained unsolved. No one had made progress on them for at least ten years, and often much longer. They concern high-dimensional geometry, coding theory, group theory, quantum complexity and lattice-based cryptography—in short, no easy matter. Among the announced results were the refutation of Connes' rigidity conjecture (Phew... my brain is burning) and answers to several questions posed by Paul Erdős, who left several hundred of them behind.
The detail that matters isn't the number ten. It's that each proof was formalized in Lean, a language that allows a machine to verify a proof line by line. So you're not being asked to take OpenAI's word for it: a computer checked the work. All 10 questions had been solved!
How much did it cost?
Two thousand dollars. A math conference costs more than that in coffee alone.
OpenAI spells it out in black and white: finding these ten solutions consumed approximately 2,000 dollars' worth of tokens, at the public rate for its API. Tokens are the units of text processed and billed by the service. Here, they were used to solve ten problems that humans had been tackling unsuccessfully for several decades! Uh... I don't know whether you realize what this means? Aren't we heading toward solving scientific, mathematical and quantum problems that humanity has never solved before?
Well, let's not get ahead of ourselves. Noam Brown, one of OpenAI's researchers, clarified: “unfortunately, no Millennium Prize Problem for now.” The great million-dollar mysteries could still resist. Mathematician Thomas Bloom also set the record straight: the AI didn't produce these results out of nowhere. It relies on more than a century of theory written by humans. Well, okay, but still!
What will we actually do with it?
Here, I have to be honest: the product doesn't exist yet. No one can say today what will happen. However, the type of task being targeted is fairly clear: when I develop, I open three Claude Code windows in parallel on three projects. Then I jump from one to the other. I break down the work, distribute it, check it, and piece the parts back together. It works very well. Except that it requires a project manager. That project manager is me, and I'm getting tired. Astra promises to do the same thing, without the tired project manager, continuously, with its fleet of agents.
For someone who doesn't code, it corresponds to the kind of chore that takes an entire weekend:
- Compare fourteen insurance or supplementary health insurance contracts, line by line, without overlooking the exclusions in the fine print, then produce a table. It's not difficult. It's just long. That's precisely why nobody does it and why we overpay for twelve years.
- Organize a trip with twelve conflicting constraints: school dates, the budget, the dog, the mother-in-law who doesn't fly. A single agent suggests one solution. A team of agents can try thirty, compare them, and then explain why the other twenty-nine don't work.
- Go through fifteen years of photos scattered across three hard drives and two old phones. Delete the duplicates, find the dates, and organize everything. A painstaking task you've been putting off since 2011.
What is frightening: what if governments used this intelligence to determine, in their military attacks, the best way to eradicate their enemies? I can see it carefully figuring out how to kill cleanly and as efficiently as possible in a region, then printing out the bill for the proposed solution for the chief of staff. Let's not forget that an AI has no qualms; it does what it's asked to do.
Now we'll have to see: a system that works alone for three days can also accumulate three days of errors without anyone seeing them. I don't know how OpenAI will introduce this new product, but if I ask it to make me a multiplayer video game but even moooore (really emphasize the s) better and it gives me something like “Concord” after a week (200 million dollars, rated 45/100 on PCGamer, whose studio Sony shut down) and sends me a bill with a few zeros after a number, I won't be happy.
The detail that gives pause: they slowed down
On August 7, OpenAI published something unusual. The company says it cannot rule out the possibility that this model may reach the level it classifies as “critical” in cybersecurity, according to its own internal assessments. This classification is not final, as testing is ongoing. In the meantime, the company says it has slowed development. Translation into plain English: commercial advertising, “Careful, we have to slow down its programming because it's going to exterminate us all.” :)
The bucket and shovel are provided. The way out isn't.
The measures taken say quite a lot. The model works in an isolated test environment, with no access to the network or tools. Its weights—that is, the internal data that determine its behavior—are encrypted. It runs in a sandbox, a closed environment that limits what it can modify. Automatic monitors also read its chain of thought, meaning the intermediate steps in its work, and can interrupt it along the way. Even during training.
Personally, I think that in 2026, the best demonstration of a model's capabilities has become the list of locks placed on it. You don't bolt a glass box around something that counts sheep.
What we don't know—and it's a lot
The name isn't even final. “Astra” is provisional. OpenAI hasn't decided whether it will be released under the name GPT-6, GPT-5.7, or as a separate family alongside Sol, Terra, and Luna. We know neither the date nor the price nor the model's size. We also don't know the size of its context window, that is, the amount of information it can keep in mind during a task. We don't know what will happen in ChatGPT either. There's not even any guarantee that the system used for mathematics will resemble the product that eventually comes out.
So yes, we're talking about a system we know almost nothing about, except for what it has already done. That's already better than the average announcement in the industry.
What I Think
For the past two years, every new model has gained a few points on a test that nobody takes in real life. That's useful. Progress is real, but it happens in small steps. Honestly, we'll see what this new beast is really made of—whether it's just a flash in the pan or whether it's going to shake things up again.
Today, an AI model forgets a decision made three hours earlier. You have to restart everything and explain it all over again. If Astra breaks through that wall, it will be more useful than ten extra points on a benchmark, that is, a test designed to compare models.
That said, I'll judge it on its actual performance, and above all, I hope it won't empty our wallets.




Join the conversation
You need an account to comment on this article. Creating one is free and takes under a minute.
No comments yet.