Three announcements in two days, and only one question at the end: who's going to pay the bill. On September 21, xAI drops its Grok 4.7 to $2 per million input tokens. On the 22nd, Anthropic releases Opus 5.5. Ninety minutes later, OpenAI drops GPT-6 Sol and GPT-6 Luna, and Sol comes in at exactly the same input price as Grok. When three labs answer one another ninety minutes apart, it's no longer a product release, it's a reverse bidding war.
A quick reminder for those who've never paid an API bill, that doorway through which a program queries a model without going through a chat window. A token is a piece of a word. Count on three quarters of a word in French. A million tokens is therefore roughly 750,000 words, the equivalent of seven good doorstoppers. When a price drops from $4 to $2 per million, it's not a supermarket promotion, it's half the bill evaporating.
What OpenAI announces, and at what price
Two models, Sol and Luna, which replace GPT-5.6 Sol and GPT-5.6 Luna. Sol is the big one, Luna the fast and cheap one.
The prices, on OpenAI's model page, read on September 23: $2 in and $10 out for Sol, against $4 and $20 for its predecessor. For Luna, 10 cents and 50 cents, against 20 cents and $1.20. Cut in half, cleanly, announced as permanent and not as a launch promotion.
The price war didn't start this week, by the way. Last week, a Chinese model, Fugu Ultra v2, was already boasting that it could beat GPT-6 Astra without even trying, for half the price.
Right, there's a “but” that nobody reads. What you can send to a model all at once, what is called its context, goes up here to 1.05 million tokens. Except that beyond 272,000, the input goes back to full price and the output costs one and a half times as much. And when a program asks the same question a hundred times, the provider doesn't charge again for what it has already read: rereading costs 90% less. It's a price designed for programs that run in loops, not for Sunday conversation.
What it looks like on a bill
The same task, the same test, two generations apart
A price per million tokens doesn't mean anything to anyone. Artificial Analysis, an independent firm that runs the models through the same tests and publishes its measurements with their date, measures something else: the cost of a completed task. And on September 23, it looked like this.
Sol pushed to the max: $1.06 per task, against $1.99 for GPT-5.6 Sol. Luna: 7 cents, against 18 cents. Over a night of scripts chaining tasks together, there's no arguing with the difference.
Watch out for the detail that changes how you read the chart. Both models write MORE than before. 31,000 output tokens for Sol against 29,000, and 51,000 for Luna against 41,000. The drop comes from the price, not from a model that has become more economical. It thinks longer, it talks more, we just pay less for the word. This isn't a detail, it's half the story.
The scores that go up, and the two that go down
Sol moves forward, Luna moves backward, on the same tests
On the coding tests, Sol improves. Terminal-Bench 4.0, tasks to carry out in a terminal (install, fix, rerun until it works), goes from 37% to 43%. The test of questions about real code, SWE-Atlas-QnA, from 54% to 58%. The firm's coding index, which mixes it all together, gains two points.
Except that Luna goes the other way. Two points lost on the same index, 49% to 44% on the questions, 66% to 64% on bug fixing. The little one regressed while the big one moved forward. The firm's conclusion is honest and a little flat: both remain at GPT-5.6's level, with improvements here and declines there.
And there is one area where both clearly fall back. On GDPval, a test where professionals compare deliverables produced by the models, Sol loses around a hundred Elo points and Luna seventy-five. Elo is the unit used to rank chess players: the lower you go, the farther behind you are. The explanation is more interesting than the drop: less care in the presentation, deliverables that forget elements requested along the way. The model writes better and turns in a sloppier paper, which is exactly the kind of progress nobody wants.
This is the moment to bring back a question I asked when GPT-6 Astra came out: was the performance really worth two and a half times the price. OpenAI's answer comes from the opposite direction. It isn't the performance that goes up to justify the price, it's the price that goes down.
The hallucination rate trap
Two bars going down, and only one of them is good news
Here is the figure everyone picked up this week: the tendency to invent an answer collapses, from 92 to 60% for Sol and from 93 to 77% for Luna, on a test of specialized knowledge from the same firm. Presented like that, it's spectacular.
Look at the second bar in the drawing. The same report says that it only takes a shot at 83% of the questions now, compared with 99% before, and that the share of correct answers across the whole set falls from 59 to 54%. So it gives up on one question out of six, and it produces fewer correct answers overall than before.
It is better, clearly much better, and the firm says so too: on its overall knowledge score, Sol goes from 22 to 27 points. But that isn't the improvement the headline figure announces all by itself. A model that invents less because it tries less, that's called a dictionary. And this gap, for a program that chains tasks together with nobody behind it, means a task that stops halfway through.
Concretely, what does this change for you
You never call OpenAI's interface from your kitchen. Except that you use things every day that run on it, and that's where it matters.
Three examples that speak to you. The automatic translation in your browser, the one that turns a German page into French with one click: every translated page is a bill somewhere, and dividing that bill by two is what makes it possible to keep the feature free. Then, all those assistants that have started answering in your place (your bank's chat, meeting summaries, your phone's voice notes): they are billed by the same token, and a price divided by two means twice as many of these features that an editor can offer in the same subscription. And if you tinker with code, even just scripts to sort your photos, it's your bill that goes down.
What you shouldn't expect, on the other hand: a reduction in your subscription. That almost never happens. The token price goes down, the feature grows, and the line on your bank statement doesn't move. That's exactly what happened with mobile data: a gigabyte cost the price of a meal fifteen years ago, today it's buried in a fifteen-euro plan and nobody counts it anymore.
What I think about it
One last thing, and it's on my side of the screen. I run three Claude Code sessions in parallel on three projects, and sometimes I look at my quota counter before the reset the way you look at the fuel gauge. The price of a token, in that setting, isn't an abstract figure: it's what decides whether I can leave a machine working while I sleep.
For me, this is the best news of the week, and not for the reason you might think. I don't need this month's model to be smarter than last month's. I need it to do the same work for half the price, because that's what decides whether something can run all the time or only when I press the button.
The fact remains that Luna is falling back and Sol is turning in sloppier papers. Both are a year older and come with a lighter bill, with a little less care in them. It's a trade-off I'm happy to make for my scripts, not for everything.
In short, thanks for cutting it in half, we'll take it. And good luck to Grok and Opus, who will have been the price champions for ninety minutes.



Join the conversation
You need an account to comment on this article. Creating one is free and takes under a minute.
No comments yet.