GLM-5.3-FlashX: 5 times faster, 2.5 times more expensive

For one week in August, an unknown model showed up on two services used by developers, free of charge, with no manufacturer name, under the pseudonym Ox Alpha. It had to be tested without knowing who we were dealing with. It became the most-used model of the week on those two platforms. And after a few days, people guessed which family it came from just by looking at how it split text into pieces, before Z.ai confirmed it and published the weights.

It was GLM-5.3-Flash, and this little game wasn't a whim: the company wanted feedback from developers who wouldn't have any preconceived opinion about the brand. Nicely done. And on Friday, the same company put another coin in the machine, with a version that goes five times faster and costs two and a half times more.

What was announced on Friday

The newcomer is called GLM-5.3-FlashX. Two figures in the press release, and they go in opposite directions. Speed first: up to 200 tokens per second, compared with around 45 for the Flash version. Five times faster. Price next: the rate goes up to $0.37 per million input tokens, where Flash is at $0.15. For output, $1.25 versus $0.50. In short, the same 2.5 factor applied to the whole rate table.

A token, to put it in context, is a piece of a word. Count on roughly three quarters of an ordinary English word. So one million tokens is about 750,000 words, or two big novels, for a few dozen cents.

An unlabeled box sitting on the shelf. Everyone tested it for a week without knowing who had made it

An unlabeled box sitting on the shelf. Everyone tested it for a week without knowing who had made it

What it costs, side by side

The price of one million tokens from the same manufacturer, on September 18, 2026

The input price of one million tokens, from the same manufacturer, on September 18, 2026

Z.ai's complete table is instructive, because it shows that the company didn't just raise a price: it built another tier. At the bottom, there's GLM-5.3-Flash at $0.15 for input. In the middle, the new FlashX at $0.37. At the top, GLM-5.3, the full model, at $1.40 for input and $4.40 for output, which is almost four times the price of FlashX.

What you need to look at isn't the raw figure, it's what the machine spends to do the work. On Artificial Analysis's index, which measures a mix of reasoning, coding and knowledge, GLM-5.3-Flash comes out at 57 points, which puts it third among all models, and the task costs $0.09 there. Claude Opus 4.8, pushed to its maximum reasoning level, gets roughly the same score but costs $2.03 per task. That's the real gap: about twenty times, not a factor of 40 as you sometimes read. And FlashX, at 2.5 times the price of Flash, is still a long way behind.

What's running underneath

The technical sheet is surprising for a model in this range. 320 billion parameters in total, but only 18 billion activated for each word it produces. That's the mixture-of-experts principle: the machine keeps dozens of specialists on hand and wakes up only eight of the 288 available to process the piece of text currently in progress. Hence a bill that stays low despite a displayed size that would be frightening.

It's also the first in the family to see, in the literal sense: text, images and video enter the same model, up to one million input tokens at once. This is a long way from the hack where you stick an image reader next to a text model.

And there's a detail the Americans noticed before we did: throughout the free testing week, the model ran exclusively on chips made in China. Not a single Nvidia accelerator in there. For a model of this quality served free of charge to thousands of developers at the same time, it's as much a demonstration as it is a press release.

What this changes on your bill

If you never use a model through the programming interface, this news doesn't concern you directly, and I prefer to say that. What's coming down is volume pricing, not subscriptions.

For everyone else, the effect is simple and you see it at the end of the month. An agent that sorts your emails, summarizes your meetings or spends the night reviewing code consumes tokens continuously, and the price of the token decides whether you let it run or shut it off. Dividing the bill by twenty completely changes what we dare to launch. My own use looks like this: I run Claude Code all day, and sometimes I switch to GLM when the month's quota runs out too quickly.

The agent working while you sleep. It's not free, it has just become cheap enough not to think about anymore

The agent working while you sleep. It's not free, it has just become cheap enough not to think about anymore

Still, watch out for a trap that this kind of announcement sets. A fast model is great, but speed also comes at the cost of rereading. At 200 tokens per second, the machine produces faster than you can read and much faster than you can check. With code, that means errors arrive in batches, and rereading becomes the real bottleneck.

What's still unclear, and that's no small thing

Three things the spec sheet doesn't say. The knowledge cutoff date hasn't been published, nor has the source of the training data, which is the norm for almost everyone but is worth saying every time. Then, the weights can be downloaded for free under the MIT license, but the full GLM-5.3 version adds a clause: a company with more than ten billion dollars in revenue has to undergo a security review before it can use it. This isn't the end of free software, it's a condition set at the top end of the market.

Finally, there's an amusing twist in the figures. The FlashX is five times faster than the Flash, but the full GLM-5.3 model remains faster than the FlashX in actual throughput, 78 tokens per second versus 45. The speed advertised at peak and the speed you get on your request aren't the same thing. That's true of all models, it's just more visible when you stack three versions of the same one.

The real question isn't the price

For two years, the promise of Chinese models came down to one sentence: just as good, much cheaper. It was true, it still is, and now they're adding a level above that and charging almost three times the price of the one below.

You have to read it like this: they arrived through price, and now they want to be paid for something else. Z.ai has been publicly traded since this year, is raising billions, and no longer needs to slash prices to exist. That's exactly what OpenAI did before it, and it's a sign that a category is growing.

Still, the free Flash, the one that served as the showcase, is still at $0.15 per input. How much longer, if nobody comes along to compete with it from below?

Join the conversation

You need an account to comment on this article. Creating one is free and takes under a minute.

  • The XMLTV file, free to download every day
  • Comment on articles and reply to other readers
  • Get an e-mail when an article you follow is updated

No comments yet.

Une erreur s'est produite. Cette application peut ne plus répondre jusqu'à ce qu'elle soit rechargée.Veuillez contacter l'auteur. Reload 🗙