GPT-6 Astra is out: is it really worth 2.5 times the price?
OpenAI dropped GPT-6 Astra Wednesday night, and since yesterday it has been rolling out in ChatGPT for Plus, Pro, Business and Enterprise subscribers, plus access for developers who plug it into their own programs.
The most widely cited independent ranking in the field, Artificial Analysis, gives it 61 out of 100. GPT-5.6 Sol, the OpenAI model it replaces, gets... 61 too. And Astra costs two and a half times more.
There, the article is over, they're taking us for fools, have a nice day everyone.
Except it isn't. This 61 doesn't say much anymore, and the interesting thing is elsewhere, in a column nobody looks at.
Meanwhile, for two days now, another table has been circulating everywhere, with 98% and 100% scores in it, and everyone is talking about the most advanced model of all time. Both images are true. We're going to have to untangle this.
Same exam grade, two and a half times the price. You're going to have to explain this to me.
First, what's this famous 61
Artificial Analysis is an independent company that gives all the models on the market the same tests, always under the same conditions, then averages the results. Math, scientific reasoning, code, general knowledge. It's a report card, with its overall average displayed prominently at the bottom of the page. And like all report cards, the average hides the essential part.
Because there comes a point when the test stops ranking the students. When the whole class hands in a dictation without a single mistake, the dictation no longer tells you who is the best, it just tells you the dictation is too easy. That's exactly where we are: the models are maxing out the tests one by one, and meanwhile they are improving at things the test doesn't measure.
By the way, and let's say it right away because hiding it won't make it any better: on this same ranking, the best one isn't OpenAI. Anthropic's Claude Fable 5.1 is at 66, Meta's Muse Spark 1.3 at 62, Astra at 61. The announced world champion comes in third with the independent referee.
And what is this 98.6% table worth?
It's authentic. It's the score table OpenAI itself published on its launch page, with its footnotes and its in-house tests.
The line everyone points at is the first one: ARC-AGI-3. Astra 98.6%, the model it replaces 7.8%, Anthropic's Claude Opus 5 30.2%. Seen this way, this is no longer a generational gap, it's an elephant..
ARC-AGI-3 isn't a general knowledge quiz. You drop the machine into a little game it has never seen before. Nobody tells it the rule, the goal, or what the buttons do. It has to feel its way around, figure out where it has landed, set itself a goal and reach it. It's deliberately made so that having swallowed the whole Internet is no use: the answer doesn't exist anywhere, you have to find it on the spot.
And there's a real achievement here. On this test, under the basic conditions, OpenAI's previous model used to top out at 7.8%. Astra, under exactly the same conditions, rises to 62.7%. Eight times better in a few months, on the test designed to withstand it. Nobody can play that figure down.
Wait. 62.7, when the table says 98.6?
Yes, and that's the whole point. Two paths, two plumbing setups around the same model. In the first, between two moves, the machine loses the thread of its reasoning: it only keeps the few notes it thought to write down. In the second, it keeps everything it had in mind and summarizes it as it goes. Same model, same questions, same cards. 62.7 on one side, 98.6 on the other.
The model hasn't changed between the last two bars. Its memory, though.
That's where I expected to find the trick, and there really isn't one. The two settings in question aren't a setup built for the occasion: they're ordinary options, open to all clients of the programming interface, and that's precisely why ARC Prize, the foundation that runs the test, accepts the result and verified it itself. What you can criticize OpenAI's leaderboard for is putting its best run up against the others' baseline run without saying so in the same entry.
Two things to remember, and the second is more useful than the first. First, ARC Prize writes in black and white, in its own report, that passing this test would not be proof of general intelligence and that it does not claim it is one. Next, if you tinker with these models: keeping the reasoning from one step to the next made the machine 3.66 times faster and made it consume 49% fewer tokens on the same tasks. It doesn't happen in the model, it happens in the way you call it.
Oh, and the detail that made me look up: on 96% of the levels, Astra needed fewer moves than a human to reach the end. Half as many on average.
So, revolution or not?
Two serious labs measured the same model in the same week and arrived at two opposite verdicts.
Epoch AI ranks it first, way out in front, at 169 points. Artificial Analysis puts it at 61, tied with its predecessor. Neither of them is cheating. They simply aren't weighting the same tests: Epoch loads up on math and puzzles, where Astra crushes everyone; Artificial Analysis loads up on code, where Claude Fable 5.1 is still ahead.
So the question “is it a revolution” has no answer until you finish the sentence with “for doing what”. For reasoning, exploring, figuring things out on its own in something unfamiliar, yes, the step up is huge. For writing code all day, no, nothing has moved this week.
François Chollet, the guy who invented this test and isn't known for his enthusiasm, has still moved his prediction forward. He was asked whether his 2030 horizon still held. His answer: “Earlier, because progress is arriving faster than I thought.”
My opinion, for what it's worth: this isn't a revolution, it's a stair step. But it's a high one, and it's sitting right where we've been told for years that these machines would never climb, that is, in a situation nobody described to them in advance.
It talks three times less to do the same work
Here is the column nobody looks at.
These machines charge by the token. A token is a piece of a word: “anticonstitutionally” is worth several, “the” is worth one. Everything you type gets split into tokens, everything the machine answers does too, and the bill counts both. Astra went from $4 to $10 per million input tokens, and from $20 to $50 per million output tokens. That's where the 2.5 times comes from.
Now for the second half of the calculation. Artificial Analysis has a test where the model doesn't answer a question but runs a real coding project all by itself, from start to finish, like a developer handed a ticket and left alone. On this test, Astra does the job while consuming about three times fewer tokens than its predecessor.
Two and a half times more expensive per token, three times fewer tokens. Do the math. The finished task costs you a tiny bit less than before, and Artificial Analysis notes that Astra reaches the level of Claude Fable 5 on that particular ranking, for less money.
Three columns, three different stories. Guess which one OpenAI put on its poster.
It's a taxi at €2.50 per kilometer that knows the city, against a €1 taxi that goes around the block three times before finding your street. The first one's meter scares you when you get in. It's the second one that ruins you.
And I'm going to tell you why this number speaks to me more than 61. When I work with Claude Code, I constantly have the percentage of credits I have left before the reset at the bottom of my screen. That percentage doesn't go down when the model is smart, it goes down when the model is talkative. A model that goes round in circles for twenty minutes, rereads the same file three times, starts its reasoning over from the beginning because it got lost, costs me a fortune and brings me nothing. So “it uses three times less”, frankly, that's the line in the press release that I read twice.
And when it doesn't know, it starts saying so
Another measure from Artificial Analysis, and this one is my favorite. They ask questions whose answers exist, they look at what the model does with the ones it can't handle, and they count how many times it makes up an answer instead of admitting that it doesn't know.
GPT-5.6 Sol made things up in 92% of cases. Astra drops to 51%.
The progress of the year isn't a more brilliant answer. It's a shrug.
It's huge, and it's the kind of progress that isn't visible on any poster because it doesn't make a good story. A machine that says “I don't know” doesn't make for a great demonstration at a conference. But if you've ever spent an evening checking a reference that a model had given you with magnificent confidence, and that simply didn't exist, you know exactly what this number is worth.
That said, we're talking about one time out of two. It's twice as good as before and it's still one time out of two. We're not there yet.
And for coding, does it beat Claude Code?
The question everyone is asking, and the honest answer is no. Not today.
On OpenAI's chart, Astra wins the code repair test, the one where they give it a real bug in a real project and see if the machine fixes it: 74.1% versus 67.4% for Claude Fable 5.1. Nice win.
Except that this test looks like an exam, not a day at work. On the one that looks like a day at work, where the model has to handle an entire project with its tools and over time, Claude Fable 5.1 in Claude Code is at 70 and Astra in Codex at 67. It's close, but it's in that direction.
The split is fairly clear in practice: Astra for driving a machine, chaining tools together, taking care of something from start to finish and for anything related to science ; Claude for staying in a big codebase for hours without losing its bearings. Me, I’m not moving my long projects. But I'm going to throw an old, really filthy project at it this week, just to see the look on its face!
And if you don't code a single line in your life?
Good question, because up to now I've been talking to you about tokens and rankings, and that doesn't fill a fridge.
The real change with this generation is the duration. The models before answered a question. This one is made to go off and work on its own for a long time, chaining steps together, using tools, clicking around in a browser by itself. Its working memory has grown to a little over one million tokens, which is roughly ten big novels that it keeps in its head from one end of the project to the other.
Before: a good sprinter that had to be restarted every three minutes. Now: a guy who leaves in the morning and comes back in the evening with the work done.
What does it look like in your life? The tax return you fill out by handing over your documents and monitoring, rather than typing in box after box. The health insurance file you hate, with its six forms and three attachments, entrusted to something that will click in your place. The eighty vacation photos to sort, rename and put away, which nobody ever does because it takes two hours and life is short.
When? Honestly, not tomorrow morning. This kind of task works in a demonstration and falls apart as soon as a website moves a button. Count on a few years before it becomes reliable enough to let it do its thing without watching. And remember that the first phone that could recognize your voice took ten years to become the assistant that sets your timer without getting it wrong. We're more or less at the same point in history here.
The bit they keep under lock and key
There's a part of Astra you won't get, and for once I think that's healthy.
OpenAI has placed this model at the highest level of its own cybersecurity risk scale. In plain English: it knows how to find holes in well-protected software and write the program that exploits them all by itself. The version you have in ChatGPT refuses this kind of request, and the full capabilities are only open to organizations selected with great care. I was talking about it the day before yesterday, when the three major labs released their hacking model on the same day, just a few hours apart.
What's new here is that the thing is now in the consumer product, with a locked door inside. It'll hold for as long as it holds.
So, do we switch or not?
If you pay for ChatGPT, you have nothing to decide, it arrives all by itself and it's better than what you had yesterday. Enjoy it.
If you're building something on top of it and paying by usage, don't look at the displayed token price, it's the trap of this launch. Measure what a finished task really costs you, on your own work. On text to classify or extract, where the model does two lines and shuts up, Astra is simply 2.5 times more expensive for nothing and you shouldn't touch it. On a long job where the machine manages on its own for an hour, it's the opposite.
And there's one thing I'm pretty happy about this morning. Twenty-five days ago, I wrote right here that Astra wouldn't answer faster but would work longer, and that this was the real shift. This week's figures say exactly that: the generalist exam score is going nowhere, endurance and efficiency are making a leap. We're changing metrics without anyone announcing it, and the posters keep selling IQ points.
So how much longer are we going to pass scores out of one hundred around while pretending that it means something? The only question I'm interested in now is: how much does the finished work cost, and is it fair?
Sources
- ARC Prize : OpenAI's GPT-6 Astra on ARC-AGI-3, the analysis by the foundation that runs the test, the two harnesses and their 62.7% and 98.6%, the comparison with humans and the explicit warning about general intelligence
- ARC Prize : GPT-6 Astra, ARC-AGI Results, September 2, 2026, the table of verified scores effort by effort and harness by harness
- OpenAI : How enabling two settings tripled our scores on the ARC-AGI-3 benchmark, the two settings in question, described by OpenAI itself
- The Decoder : Benchmarks disagree on GPT-6 Astra, Epoch AI's 169 points versus Artificial Analysis's 61, why the two diverge, and François Chollet's forecast brought forward
- Artificial Analysis : Benchmarking GPT-6 Astra, the independent measurement from September 3, 2026, the intelligence score, the coding task, token consumption and the rate of fabricated responses
- OpenAI : GPT-6 Astra, a new generation of intelligence, the original announcement, the pricing, the working memory size and the rollout
- Requesty : GPT-6 Astra scores 61 on the independent index, the same as Sol, at 2.5x the price, the comparison between the launch claims and the independent measurements
- Tech Startups : Top tech news today, September 4, 2026, the rollout schedule in ChatGPT and with hosting providers





Join the conversation
You need an account to comment on this article. Creating one is free and takes under a minute.
No comments yet.