OpenAI publishes the figures for its in house chip. It beats Nvidia, and you need to read the small print.

OpenAI publishes the figures for its in-house chip. It beats Nvidia, and you need to read the small print.

On Tuesday, OpenAI published the test results for its first ever chip, the one that's supposed to run its models in its own machines. It's called Jalapeño, it was designed with Broadcom and manufactured by TSMC, and the figure being highlighted is this: with the same amount of electricity, it does almost twice as much work as an Nvidia system.

That's a big figure. It's probably true. And it's measured by the one selling the chip, under the conditions it chose, which doesn't invalidate it but completely changes how you should read it.

Two identical electricity meters, two piles of tokens of different sizes

The meter on the left and the one on the right are running at the same speed. What's coming out underneath isn't.

The figure that matters isn't the one you think

For twenty years, we compared processors by speed. The fastest one won, period.

Today the question has changed, and it comes up like it does for a car: what matters now isn't top speed, it's the number of kilometres per litre. Except that here the litre is called the watt.

A little vocabulary note before we go any further, just one and it will be useful throughout the article. When an artificial intelligence answers you, it doesn't make sentences in one go: it produces pieces of words one after another, very quickly. These pieces are called tokens. The whole industry counts in tokens per second, just as a printing press counts in pages per hour.

The measurement published by OpenAI is therefore: how many tokens per kilowatt of electricity consumed. Pages per hour, per litre of fuel.

Why electricity, and not money

This is where I have to tell you about something I find fascinating, because it can't be seen anywhere and it decides everything.

A data centre today isn't limited by space. There's land available. Nor is it really limited by money, given the sums flying around. It's limited by the power that the local electricity grid is willing to deliver to it.

Server racks connected to a pylon by a narrow cable

The building can be as large as you like. The cable coming into it can't.

Once you've secured your fifty megawatts, you can't have any more for years: you'd have to build lines, substations, sometimes a power plant. Your number of megawatts is therefore fixed, and the only question left is: how much work can I manage to fit into it?

That's why performance per watt has become the real currency of this sector. Doubling the work per watt doesn't mean "going twice as fast". It means serving twice as many people without pulling a single extra cable.

The figures announced

On an open model that anyone can test, compared with an Nvidia system from the Blackwell generation, OpenAI announces about 1.9 times more work per kilowatt.

Tokens per kilowatt: 85,448 for Jalapeno, 44,960 for the Nvidia system being compared

Two bars, a single unit, and a caveat written at the bottom. It matters just as much as the bars.

Across all the scenarios tested, the company claims 1.5 to 1.9 times more compute per watt, and response times that are 1.7 to 3.6 times shorter. On paper, that's a real slap.

Now, the small print

And this is where it gets interesting, because what follows applies to all tests, not just this one.

A magnifying glass over the last line of a results table

The most instructive line is always the smallest one.

First question: who is holding the stopwatch? Here, it's OpenAI testing OpenAI's chip. It's not cheating, the figures will be checked by others, but whoever chooses the tests always chooses the ones they're good at.

Second question: compared to what? The competitor chosen is Nvidia's Blackwell generation. Except Jalapeño will arrive facing the next generation, Rubin, which is already being shipped. Comparing next year's car to last year's model, it happens a lot and is rarely said out loud.

Third question: are the weapons equal? No, and that's the strongest point of the criticism. Jalapeño carries memory from a more recent generation than the one in the Nvidia system it's being compared with, and for this kind of work, memory accounts for a good part of the result. Some of the measured lead is therefore not in the chip, it's in the memory sticks around it.

Fourth question: does it really exist? Not quite yet. We're talking about development samples. Deployment is supposed to begin in OpenAI's own machines by the end of the year, serious production is expected in 2027, and full capacity not before 2028. A lot can happen before then, including nothing.

A detail that looks like nothing and says everything: OpenAI specified that it was continuing to buy from Nvidia. A lot. When you announce that you've done better than your supplier while confirming that you're still placing orders with it, you know very well where you stand.

What does it change for you?

Nothing this year, and I prefer to say that right away.

In the medium term, two things, and they're concrete.

The first is the price. What you pay to use artificial intelligence, whether it's a monthly subscription or a pay-per-use bill, is tied to two costs: the hardware and the electricity it guzzles. A machine that does the same work with half as much power, that's downward pressure on what it costs. Not a guarantee, pressure.

The second is dependence. Today, almost everything that runs AI in the world comes out of the same company, and when a supplier is alone, it sets the prices and the deadlines all by itself. Every player that manages to run its models on its own chip takes a little of that power away. It's the same battle as the one I was telling you about eight days ago with a small company aiming for exactly the same place, except that here it's no longer an ambitious startup, it's Nvidia's biggest customer making its own hardware.

And there's a knock-on effect that concerns you more directly than you think. Data centers and your house draw from the same grid. Every time one of these machines does the same work with half as much power, those are megawatts that aren't being demanded from someone's grid. It won't lower your bill, but it counts in the queue.

What I take away from it

Personally, what I like about this isn't the score. It's the shift in the question.

For years, the race was about “who has the biggest model”, then about “who has the most graphics cards”. Now it's about how much work we can get out of a megawatt, and that's a physical constraint, the only one marketing can't get around. You can't convince an electrical transformer with a good presentation.

And then, let's be honest for two seconds: a chip called Jalapeño, that deserved an article!


Sources

Join the conversation

You need an account to comment on this article. Creating one is free and takes under a minute.

  • The XMLTV file, free to download every day
  • Comment on articles and reply to other readers
  • Get an e-mail when an article you follow is updated

No comments yet.

Une erreur s'est produite. Cette application peut ne plus répondre jusqu'à ce qu'elle soit rechargée.Veuillez contacter l'auteur. Reload 🗙