Fugu Ultra v2 beats GPT-6 Astra without using it, and costs half as much

There is a category of products starting to come out that is going to make people grind their teeth: models that aren't models. Since September 10, the Japanese company Sakana AI has been selling its own, and its argument fits in one sentence. When you ask it a question, it isn't the one that answers. It decides who answers.

The product is called Fugu Ultra v2, there is a second, cheaper version called Fugu Max, and the promise is the kind that makes you cough. In the tests Sakana has published, Fugu Ultra v2 comes ahead of GPT-6 Astra and Claude Fable 5.1 on several tests. Except that neither of them is part of its kitchen. It beats them without using them.

The garage switchboard operator

You call your garage about a strange noise in the car. The guy who picks up doesn't repair anything himself: he listens for three seconds, says “sounds like the gearbox, I'll put you through to Marcel,” and you end up with someone who knows gearboxes. If the noise is coming from the brakes, it'll be someone else. You dialed a single number, and you got the right person.

Fugu works exactly like this switchboard operator. Your request arrives at a single address, a program reads it, decides what kind of work it is, sends it to the model or models best placed for it, checks what comes back and gives you a single answer. Sometimes one specialist is enough, sometimes several are needed, working one after another or in parallel, and the switchboard operator pieces it all back together. From the outside, nothing changes: same address, same response format, same bill.

This isn't improvised. Sakana is a Tokyo lab that has been working on this subject for months, with two research papers published this year on how to get several models to work together. And the name is well chosen: fugu is that Japanese fish that kills you if the chef gets the wrong piece.

The figures, and the price

The figures now, keeping in mind that they come from Sakana and that nobody independent has reproduced them. On a test of reading charts and images, Fugu Ultra v2 scores 48.3 against 27.3 for Claude Opus 5 and 29.5 for Fable 5. On a coding test, it scores 74.3. Overall, it comes first or tied for first on five of its eight in-house tests, and is in the top two on seven of them.

The price is the other half of the story, and it's the half that concerns me directly. Fugu Ultra v2 costs $5 per million tokens sent and $30 per million tokens received. For comparison, GPT-6 Astra, which I talked about last week, is at 10 and 50. The Japanese model therefore costs half as much on input and 40% less on output, for announced results in the same ballpark.

Price of one million tokens on input and output: 2 and 6 dollars for Fugu Max, 5 and 30 for Fugu Ultra v2, 10 and 50 for GPT-6 Astra

The price of output, in dollars per million tokens. Three answers to the same question don't cost the same

And the cheaper one wins on their own tests

This is where it gets amusing. Fugu Max, the budget version, costs $2 on input and $6 on output, which is five times cheaper than the high-end version on output. And in Sakana's tests, it isn't the expensive version that wins: Max comes first on six of their tests, including Terminal Bench 2.1, the one that measures how well an agent gets through a real task on a real computer.

Sakana explains the difference through the composition of the kitchen. Max draws heavily from open models and specialized models, including Nvidia's Nemotron family, while Ultra v2 sticks to a more selective lineup. Translation: for a lot of work, a good assembly of average models does better than one very good model, for five times less. If this is confirmed anywhere other than on their own test suite, it's the most important news of the new season, and it doesn't concern the models. It concerns the bill.

The catch, and it's a big one

Now, what nobody says in the press releases. Sakana refuses to say which models are in the kitchen. It's written in black and white: the choice of providers and the routing plan are considered their trade secret. So when you use Fugu, you don't know who read your text.

Three consequences, in order of importance to you. One: you can't tell your client, your boss or your lawyer where their data went, because you don't know yourself. Two: the product can change models behind the same name, from one day to the next, without the name changing. Three: you also pay for the tokens the switchboard operator uses to think, and once again, not a single line on the bill tells you who did what.

There is one last detail that will interest European readers. The service is not available in the European Union, the United Kingdom, or Switzerland. Reason given: data protection issues have not been settled, and no date has been announced. For us in Belgium, the question is therefore settled for now: it is not an option, and not because someone decided to protect us, simply because the case is not ready.

Concretely, what does this change for you?

If you don't code, you already use this principle without knowing it. All the big assistants route things internally: a small model to rephrase your sentence, a big one for the difficult question. What changes here is not the method, it is that a company is making it into a product sold separately, and above all that it is talking about it in terms of price rather than power.

As far as your wallet is concerned, the timeline is worth stating honestly: none of this is going to cut your subscription bill tomorrow morning. Fugu is not even sold here. But the direction is clear, and it is a healthy one. For a year now, every new model has cost more than the previous one. Here, we are seeing machines that do the same thing by choosing the cheapest part, and that is the first good pricing news in a long time.

If you're a developer, you can already hack together the same thing yourself. I talked about it a few weeks ago in connection with OpenRouter, with that button that lets you switch AIs without rewriting a single line of code. Fugu is the same movement, except that the routing is done by a model trained for that purpose instead of by your own rules.

The front of an unbranded black box in a computer rack

What you see of the kitchen. You know what goes in, you know what comes out, and between the two there is no one left to ask

There is one question left that I am asking myself and that no one is answering for now. If the best model in the world is a switchboard operator, then the race for the biggest model may not be the only one that matters. And if an assembly of open models really beat the most expensive closed models at everyday work, that would be the best news of the year for everyone who pays for their tokens as they use them.

Join the conversation

You need an account to comment on this article. Creating one is free and takes under a minute.

  • The XMLTV file, free to download every day
  • Comment on articles and reply to other readers
  • Get an e-mail when an article you follow is updated

No comments yet.

Une erreur s'est produite. Cette application peut ne plus répondre jusqu'à ce qu'elle soit rechargée.Veuillez contacter l'auteur. Reload 🗙