ByteDance is building a 10,000 billion parameter model. It's the least interesting figure in the story.

ByteDance is building a 10,000 billion parameter model. That's the least interesting figure in the whole affair.


On August 7, the Financial Times announced that ByteDance, TikTok's parent company, is training an artificial intelligence model that could reach 10,000 billion parameters. The information comes from three people familiar with the project. ByteDance refuses to comment and no release date has been announced.

The figure went around the world in two days. Yet it says almost nothing. The model I was telling you about yesterday helps explain why.

A gigantic industrial engine suspended from a crane, being lowered toward the hood of a small city car that clearly cannot accommodate it

It fits, promise. You just have to remove the seats. And the car.

What we really know

A parameter is an internal setting that the model learns during its training. A model can contain thousands of billions of them. It is also one of the few figures a company can announce before it has even finished its model. Handy for making a headline without having proved anything yet.

The model is in pretraining. This is the first phase, when it gulps down data in bulk. It lasts three to six months. The model then has to be fine-tuned to teach it to answer correctly. Even its final number of parameters has not yet been set.

Bar chart of the number of parameters in billions, from 284 for DeepSeek V4-Flash to 10,000 for ByteDance, with the estimates shown as hollow bars

Solid bars represent published figures. Hollow bars, estimates. Guess which ones make the best headlines.

To put the project into perspective, Kimi K3, so far the largest Chinese model, runs at around 2,800 billion parameters. The industry estimates Anthropic's Mythos 5 at around 8,000 billion, but nobody has confirmed it. ByteDance would therefore be aiming for the top of the global rankings.

Why this figure says almost nothing

Yesterday, I was telling you the story of DeepSeek V4-Flash-0731. Look at the first bar in the chart: 284 billion parameters. That's thirty-five times fewer than ByteDance's project. Yet this small model beats its own company's big model by more than ten points in tests where it has to carry out tasks. Its architecture has not changed. Only its fine-tuning phase was redone.

The number of parameters is to a model what engine displacement is to a car. It gives you an idea of the engine's size. It says nothing about fuel consumption, road handling or the thing's ability to start in the morning.

What really matters is the model's architecture, the quality of the data and above all its fine-tuning. Three elements impossible to sum up in a single number. Strangely, press releases are much less interested in them.

The detail nobody put in the headline, even though it's worth more than the figure

Zhang Yiming, the founder of ByteDance, told his teams that he refused distillation for this model. Distillation consists of learning by copying the answers of an already high-performing competing model. You ask it millions of questions, collect its answers, then train your own model with that data. It's fast, effective and pretty much everyone does it.

Refusing this method means accepting several months of delay behind its Chinese rivals. It's a costly and consequential industrial choice. Personally, I find that a hundred times more interesting than a parameter counter. I haven't seen it anywhere in a big headline, though.

Doubao, the world's second assistant you've never heard of

That's the real lesson of this story. And it has nothing to do with the training currently under way.

ByteDance already has an assistant called Doubao. In April 2026, it had 336 million monthly active users. It was number one in China and number two in the world, behind ChatGPT and ahead of all the others.

A huge stadium filled with tens of thousands of people with a banner reading 336 million users, and in the foreground a person on their couch who has no idea

Second assistant in the world. Ask around you at lunch today, you’re going to have fun.

Ask ten people around you to name an AI assistant. You’ll hear ChatGPT ten times. Maybe Claude or Gemini once. Doubao, zero. A product used by more people than the population of the United States remains perfectly invisible over here.

It’s not a question of quality, but of the market. A good part of what’s happening in AI today simply never reaches us.

Concretely, what does it change for you

If you don’t code and you don’t have TikTok, you’re probably wondering what you’re doing here. There are three answers. The first is the most useful.

Comparison, on the left an old camera displaying 20 megapixels with a blurry photo, on the right a modern phone displaying 12 megapixels with a sharp photo

We’ve already experienced exactly this. The numbers were accurate. The conclusion, much less so.

You’ve already learned to be wary of this kind of number. Remember megapixels. For ten years, every camera displayed a higher number than the one next to it. So we bought the one that had the most. Then we understood that the sensor size, the optics and the software made the photo, not that number. Today, a 12-megapixel phone crushes a 20-megapixel compact camera from 2009. For parameters, we’re exactly at the same stage. When you read the next headline announcing 10,000 billion, you’ll know what to do with it.

This race is bringing your bill down. Not theirs, yours. When four giants fight to offer the same feature, the price collapses. At DeepSeek, we went down to $0.28 per million generated words, compared with $25 for high-end models. Result, transcribing your voice messages, sorting your photos or automatic subtitles stop being monthly subscriptions. They become simple options in software you already own.

You may use this model without knowing it. A model isn’t only sold directly. It can also be rented to other applications. The engine that answers in your photo-editing app or in your car assistant has no reason to have a familiar name. Let’s be honest about the timeline: this model doesn’t exist yet. It may come out in six months, a year from now or never. No one has announced a date.

What I think of it

The most striking thing isn’t the race. It’s our blind spot. We keep commenting on three American companies while an assistant used by 336 million people develops without getting a line from us.

As for the number itself, I’m going to be direct: today, 10,000 billion parameters are not information. It’s an intention. The model doesn’t exist yet. It can’t do anything and no one has tested it. We’re commenting on a quote, not a house.

The real question is whether Zhang Yiming’s gamble will hold. Refusing to copy his competitors when everyone is doing it is either costly pride, or the only way to eventually overtake them. Seen from here, the two look very much alike.

Join the conversation

You need an account to comment on this article. Creating one is free and takes under a minute.

  • The XMLTV file, free to download every day
  • Comment on articles and reply to other readers
  • Get an e-mail when an article you follow is updated

No comments yet.

Une erreur s'est produite. Cette application peut ne plus répondre jusqu'à ce qu'elle soit rechargée.Veuillez contacter l'auteur. Reload 🗙