Grok 4.7 for $2, MiMo v2.6 to download: which AI model should you choose?

Monday afternoon, 3:39 p.m. UTC time. A repository appears on Hugging Face with 309 billion parameters signed by Xiaomi, under the MIT license, the one that allows pretty much everything, including making money from it. Eighteen seconds later, the Pro version was online right next to it. That same afternoon, SpaceXAI was releasing Grok 4.7 and announcing 2 dollars per million input tokens, 6 dollars for output.

Two models on the same day, that's nothing exceptional, three come out every week. What's interesting is what the two announcements tell us together. One is rented by the minute, the other is downloaded. One needs an air-conditioned room, the other fits in a folder on your hard drive.

Grok 4.7: half the price, and it doesn't win everywhere

The official page announces the most powerful model for “code and knowledge work”, twice as fast and half the price of comparable competition, offered at the same price as Grok 4.6. That's the sentence every lab writes on release day, and it's worth as much as the charts that come with it.

As it happens, there is one. Four tests, three models. I redrew it with the figures exactly as they are, keeping the test versions announced by each one:

Grok 4.7, GPT-5.6 Sol max and Fable 5.1 max compared across four tests

Two out of four tests for Grok. And the only one that really matters to me, repairing code, ends in a tie

Translation, because test names don't mean anything to anyone. CursorBench means writing code in an editor when you're given the existing project. DeepSWE means repairing real damaged code in a repository. Terminal-Bench means managing on your own in a command line, with no interface and no safety net. Harvey Legal means a US law exam, the one lawyers take.

Bottom line: Grok 4.7 crushes it on law, 19.6% against 2.5% and 6.7% for the other two, which is no small slap in the face. It also blows up the competition on pure engineering. And it gets left behind on the terminal, where Fable 5.1 max hits 57.9% while it tops out at 38%. On code repair, the three are within two and a half points of one another, which amounts to saying they're tied.

The price, now, because that's the real announcement of the day:

Price per million input and output tokens for the three models

The orange bar is what you pay for what the machine writes. With Grok, it's five times shorter than with Fable

Compared with real work, that changes a bill. I run three coding assistant sessions constantly, and the “model” line in my budget has become a line I look at every month, like other people look at their electricity bill. At 2 dollars per million tokens read, we're no longer having the same conversation as at 10. If you want the details of all the prices, both subscriptions and APIs, I went through the whole thing in this comparison per million tokens a month and a half ago, and the orders of magnitude haven't changed.

MiMo v2.6: 309 billion parameters, and 15 doing the work

At Xiaomi, the same day, it was a different story. The model is called MiMo v2.6 and it's enormous: 309 billion parameters in total, spread across 48 layers and 256 experts. Except that out of those 256 experts, only 8 wake up to process a given word, which means that around 15 billion parameters are effectively running at once. That's the mixture-of-experts principle: a big brain, but one that only turns on the useful areas. Result, a model the size of a semi-trailer truck that consumes like a van.

The rest of the technical specifications, checked in the configuration file published with the weights: one million context tokens, meaning enough to swallow an entire large codebase and keep talking, and an input that accepts text, images, video and sound. The license is MIT, with no research clause and no revenue ceiling, which is rare for a model this size. The weights weigh 172.9 GB, split into 65 chunks to download.

Eight graphics cards mounted in parallel inside an open chassis

Eight cards, and that's the minimum announced. The electricity bill is not included in the model's price

What you'd need to run it at home

172.9 GB won't fit in your graphics card. It's not even a memory problem, it's a machine problem: the deployment documentation talks about a minimum of eight cards in parallel, sixteen plus two in the other recommended configuration, across several servers connected to each other. In other words, a rack, networking, a room that gets rid of the heat.

It's the same cliff as the one I described for another 750-billion-parameter model released this summer : beyond a certain size, "free" means "free for you if you already have a datacenter".

Except Xiaomi thought about the rest of the world. Alongside the big model, the same family exists in a reduced 9-billion-parameter version, and that one has already been converted by the community to run on ordinary hardware : a GGUF version for PCs, a 4-bit MLX version for Apple Silicon Macs. Copies have been circulating since the night it was released. It's the real gift of the day, and it isn't on anyone's cover.

Concretely, what does this change for you?

Three things, in order of arrival.

Tonight, nothing. You have neither the cards nor the rack, and the small 9-billion-parameter model won't be of much use for anything your current assistant doesn't already do. If you like tinkering, though, you can download it and fine-tune it on your own data without asking anyone's permission and without it costing you a cent more than a hard drive. That was impossible with the big models a year ago.

This year, the price. Grok 4.7 at $2 per million tokens read puts pressure on everyone, and it's your subscription that eventually benefits from it. When an enterprise model costs five times less than its equivalent from six months ago and works better, the only question left is who will be the first to match it.

In two or three years, the floor below. A 309-billion-parameter model under an MIT license means that any company, any administration, any hospital can install it at home and run it without a single line of its data going elsewhere. Two years ago, this kind of capability was reserved for five companies in the world. Today, it's a 172 GB file you download overnight. Does that mean everything is going to become free? No, because someone still has to pay for the electricity. But it means that the day your employer wants its own AI, the answer will no longer be "impossible".

What you shouldn't believe

Xiaomi's figures, first of all. They are self-reported, based on tests whose versions aren't the ones in Grok's table, and on a model they trained themselves. Comparing MiMo's 87.6% on Terminal Bench 2.1 with Fable 5.1's 57.9% on Terminal Bench 4.0 would be a calculation error, and nobody did it because nobody released both in the same test. The two can't be measured side by side, and that's exactly the kind of territory where marketing has a field day.

The repository's success, next, which makes me smile. As I write this, the Hugging Face page shows 151 people who have hearted the model and zero downloads. Zero. Hugging Face statistics are known to lag, but admit that for a model announced as a revolution, it's a nice indicator of what's really happening: people applaud, then go back to their ten-dollar subscription.

And Grok's table itself, which is published by SpaceXAI on its own site. It wins where it wins, it loses where it loses, that's not a surprise, but the twenty-point gap on the terminal rarely appears in the announcements. It's worth noting.

Two models on the same day, one you rent and one you get for free. If you had the choice, which one would you take? I keep the one that runs on my machine, even when it's stupider. The day my connection goes down, at least it's still there.

Join the conversation

You need an account to comment on this article. Creating one is free and takes under a minute.

  • The XMLTV file, free to download every day
  • Comment on articles and reply to other readers
  • Get an e-mail when an article you follow is updated

No comments yet.

Une erreur s'est produite. Cette application peut ne plus répondre jusqu'à ce qu'elle soit rechargée.Veuillez contacter l'auteur. Reload 🗙