A model 89 times cheaper nearly matches Opus’s score. The catch isn’t where you think.

A model 89 times cheaper nearly matches Opus's score. The catch isn't where you think.


On July 31, DeepSeek released V4-Flash-0731. The model costs $0.28 per million output tokens, compared with $25 for Claude Opus 5. In its results table, it also scores 82.7 on Terminal-Bench, compared with Opus 4.8's 85.0.

A two-point difference for a price 89 times lower. If the comparison holds up, everyone has a problem.

I checked. The figure is true. The comparison isn't. And the reason is more interesting than the model itself.

Two identical race cars, the one on the left surrounded by a full team of mechanics, the one on the right with a single man and a bicycle pump

Same engine, same track. Good luck comparing lap times.

First, the model

DeepSeek-V4-Flash-0731 isn't a new model. It's the same one as the trial version released a few weeks earlier: same architecture, same size, 284 billion parameters. Only the post-training was redone. Post-training is the training phase that comes after raw learning. It teaches the model to use tools and follow instructions, instead of simply completing text.

In other words, they didn't change the brain. They taught it how to work.

As for pricing, it costs $0.14 per million input tokens and $0.28 per million output tokens. A token is a piece of a word. One million tokens represents roughly 750,000 words. If the submitted text has already been processed before, the price drops to $0.0028, or 2% of the regular price.

Bar chart comparing the price of one million output tokens for five models, from $0.28 to $50

DeepSeek's bar is four pixels wide. It's not a display bug, it's the price.

The rest of the specifications hold up: a context window of one million tokens, meaning the amount of text the model can keep in view at once, or about 1,500 pages. It also produces 112.8 tokens per second and scores 52 on the intelligence index from Artificial Analysis. This organization conducts its own measurements instead of repeating the press release.

So far, so good. It's a good, inexpensive model, and its weights can be downloaded. Weights are the files that contain what the model has learned.

The figure that doesn't add up

In DeepSeek's table, Claude Opus 4.8 scores 85.0 on Terminal-Bench 2.1.

On the official Terminal-Bench leaderboard, the same Opus 4.8 scores 78.9.

A six-point difference. Same model, same test, same version number. Nobody lied. Yet the two figures don't measure exactly the same thing.

The difference comes from the harness.

What is a harness?

A language model can't do anything on its own. It reads text and produces text. That's it. To fix a bug in a project, someone has to provide it with the files, run the suggested commands, send back the error messages, then decide when the work is finished or when the model is going in circles.

That someone is a program. It's called a harness. Claude Code is one. Codex is another. Terminus 2 is as well. It's the harness used by the official leaderboard to test all models under the same conditions.

The harness decides everything that isn't directly handled by the model: the number of attempts before giving up, the available tools, how the problem is presented, or how it reacts to a command that takes three minutes. Change the harness and you change the score, without touching the model.

This isn't a hypothesis. It can be measured.

Graphique montrant trois modèles évalués chacun avec deux harnais différents, avec les marges d'erreur

Three models, two harnesses each, one evaluator. With Gemini, the two points overlap. That is information too.

Claude Fable 5 scores 83.8 with Claude Code and 80.4 with Terminus 2. Same model, same day, same evaluator, but a 3.4-point gap. Opus 4.7 scores 68.9 versus 66.1, a 2.8-point gap. Gemini 3.1 Pro comes in at 65.8 versus 65.6. This time, the difference lies entirely within the margin of error.

The latter case is the most instructive. The harness effect is not a constant that could be subtracted everywhere. It depends on the model. Some models benefit much more than others from their own tooling. That is not very surprising when the company that makes the model also designs the harness.

What DeepSeek says, and what it does not hide

Let us be fair: DeepSeek hid nothing. The method is written out in black and white.

Harnais   : DeepSeek Harness, mode minimal
Palier    : max
Réglages  : top_p = 0.95, temperature = 1.0
Publié    : non, annoncé comme à venir

DeepSeek even specifies that agent scores—that is, the results obtained when the model acts through tools—are extremely sensitive to the harness. They should therefore be treated as manufacturer figures until a third party has reproduced them. One would like everyone to state this just as clearly.

The problem is that this harness has not been published. No one can therefore repeat the measurement. And DeepSeek V4-Flash still does not appear in the official ranking, which has seventeen entries.

Now let us perform the calculation missing from the announcement. DeepSeek's harness gives Opus 4.8 a score of 85.0. The official ranking gives it 78.9. With Opus, this harness therefore produces a score approximately six points higher. I cannot prove that it has the same effect on DeepSeek, since I do not have access. But I find it hard to see why tooling that is so generous to the competitor would suddenly become harsh with the homegrown model. The announced score of 82.7 would then be worth something like 76 or 77 on the official ranking's scale. That is a good score. It is not Opus's score.

The only figure in the table I trust

There is one. It is also the most interesting.

In the same table, with the same harness and on the same day, V4-Flash scores 82.7. V4-Pro, its big brother, scores 72.1. The smaller model therefore beats the larger model from the same company by more than ten points.

This comparison is valid because everything else is identical. Here is the rule: comparing two models with the same harness makes sense. Comparing scores obtained with two different harnesses makes no sense at all.

This result tells us something more interesting than the race for rankings. A significantly smaller model, trained differently, performs better than a large model at agent work. Size no longer determines everything.

What does this change for you, concretely?

If you do not program, you may be wondering why I am telling you all this. There are three reasons. Two of them will affect your wallet.

Comparaison avant après, à gauche une pile de factures d'abonnement sur une table de cuisine, à droite la même table avec une seule pièce de monnaie

The goal is not to pay for one more subscription. It is to no longer need any.

The price of artificial intelligence is collapsing, and fast. High-end models still cost between 25 and 50 dollars per million tokens produced. This one costs 0.28 dollars, and its results are not ridiculous. At that price, a feature can stop being sold as a monthly subscription. It becomes a simple option in software you already use. Transcribing voice messages, sorting holiday photos, or generating automatic subtitles will soon no longer be billed separately, because their cost will become almost zero. The real relief is not paying for fewer subscriptions. It is no longer having to sign up for them.

A ranking does not measure the whole reality. It measures a result under specific conditions, which are almost never yours. When an advertisement claims that one assistant is better than another because it scores two points higher, you now know that those two points can be as thin as a hairline. That applies to AI. It was already true of smartphone speed tests and advertised car fuel consumption. Same mechanism, same century.

The precedent exists, and it is rather reassuring. In the 1990s, every PC manufacturer published its own performance tests. They were all flattering and all incomparable. That lasted until independent laboratories imposed their own protocols and in-house figures ceased to interest anyone. AI is at exactly that stage, a few years behind schedule. How long will it take? Two years, five years, or never if nobody funds neutral examiners.

What I think

I am not going to take a swipe at DeepSeek. The company published its settings and its warning, which puts it above the industry average. The problem is not DeepSeek, but the way the table is being used. When a company fills in its competitor's column itself, it is not a comparison. It is a sales brochure.

What really annoys me is the figure being repeated. Within a few days, I had read everywhere that this model was only two points behind Opus. Nobody went to consult the official ranking. It is public, free, and opens in thirty seconds. The figure spread on its own because it made a good headline.

The real information is not in this duel with Opus. A model with 284 billion parameters beats its own company's larger model without changing a single line of its architecture. All it took was retraining it. That is the news. It made no headlines.

Join the conversation

You need an account to comment on this article. Creating one is free and takes under a minute.

  • The XMLTV file, free to download every day
  • Comment on articles and reply to other readers
  • Get an e-mail when an article you follow is updated

No comments yet.

Une erreur s'est produite. Cette application peut ne plus répondre jusqu'à ce qu'elle soit rechargée.Veuillez contacter l'auteur. Reload 🗙