ByteDance is building a 10,000-billion-parameter model. That's the least interesting figure in the story.
On August 7, the Financial Times reported that ByteDance, TikTok's parent company, is training an artificial intelligence model that could reach 10,000 billion parameters. The information comes from three people familiar with the project. ByteDance declined to comment, and no release date has been announced.
The figure went around the world in two days. Yet it means almost nothing. The model I told you about yesterday helps explain why.
It fits, I promise. You just have to remove the seats. And the car.
What we really know
A parameter is an internal setting that the model learns during training. A model can contain thousands of billions of them. It's also one of the few figures a company can announce before it has even finished building its model. Handy for creating a headline without having proved anything yet.
The model is in pretraining. That's the first phase, when it gulps down data in bulk. It lasts three to six months. The model then needs to be fine-tuned to teach it how to respond correctly. Even its final number of parameters has not yet been determined.
Solid bars represent published figures. Hollow bars represent estimates. Guess which ones make the best headlines.
To put the project into perspective, Kimi K3, until now the largest Chinese model, has around 2,800 billion parameters. The industry estimates Anthropic's Mythos 5 at around 8,000 billion, but nobody has confirmed it. ByteDance would therefore be aiming for the top of the global rankings.
Why this figure tells us almost nothing
Yesterday, I told you the story of DeepSeek V4-Flash-0731. Look at the first bar in the chart: 284 billion parameters. That's thirty-five times fewer than ByteDance's project. Yet this small model beats its own company's larger model by more than ten points in tests where it has to complete tasks. Its architecture hasn't changed. Only its fine-tuning phase was redone.
The number of parameters is to a model what engine displacement is to a car. It gives you an idea of the engine's size. It tells you nothing about fuel consumption, road handling, or whether the thing can start in the morning.
What really matters is the model's architecture, the quality of the data, and above all its fine-tuning. Three elements that cannot be summed up in a single number. Strangely, they attract much less attention in press releases.
The detail nobody put in the headline, even though it's worth more than the figure
Zhang Yiming, ByteDance's founder, told his teams that he refused to use distillation for this model. Distillation consists of learning by copying the answers of an already high-performing competing model. You ask it millions of questions, collect its answers, and then train your own model using that data. It's fast, effective, and pretty much everyone does it.
Refusing this method means accepting a lead of several months for its Chinese rivals. It's a costly and consequential industrial choice. Personally, I find that a hundred times more interesting than a parameter counter. Yet I haven't seen it anywhere in a major headline.
Doubao, the world's second assistant you've never heard of
That's the real lesson of this story. And it has nothing to do with the training currently underway.
ByteDance already owns an assistant called Doubao. In April 2026, it had 336 million monthly active users. It was number one in China and number two in the world, behind ChatGPT and ahead of all the others.
Second assistant in the world. Ask around at lunchtime today, you'll have fun.
Ask ten people around you to name an AI assistant. You'll hear ChatGPT ten times. Maybe Claude or Gemini once. Doubao, zero. A product used by more people than the population of the United States remains completely invisible here.
It's not a question of quality, but of market. A significant part of what's happening in AI today simply never reaches us.
Concretely, what does this change for you?
If you don't code and you don't have TikTok, you're probably wondering what you're doing here. There are three answers. The first is the most useful.
We've already experienced exactly this. The figures were accurate. The conclusion, much less so.
You've already learned to be wary of this kind of figure. Remember megapixels. For ten years, every camera displayed a higher number than the one next to it. So we bought the one with the most. Then we understood that the sensor size, optics, and software made the photo—not that number. Today, a 12-megapixel phone outperforms a 20-megapixel compact camera from 2009 by far. With parameters, we're at exactly the same stage. When you read the next headline announcing 10 trillion, you'll know what to make of it.
This race brings down your bill. Not theirs—yours. When four giants compete to offer the same feature, the price collapses. With DeepSeek, we've gone from 25 dollars for high-end models to 0.28 dollars per million generated words. As a result, transcribing your voice messages, sorting your photos, or generating automatic subtitles stops being a monthly subscription. They become simple options in software you already own.
You may use this model without knowing it. A model isn't sold only directly. It can also be rented to other applications. The engine that responds in your photo-editing app or in your car assistant has no reason to bear a familiar name. Let's be honest about the timeline: this model doesn't exist yet. It may be released in six months, in a year, or never. No one has announced a date.
What I think
The most striking thing isn't the race. It's our blind spot. We endlessly comment on three American companies while an assistant used by 336 million people develops without getting a single line written about it here.
As for the figure itself, I'll be direct: today, 10 trillion parameters is not information. It's an intention. The model doesn't exist yet. It can't do anything, and no one has tested it. We're commenting on a quote, not a house.
The real question is whether Zhang Yiming's gamble will hold. Refusing to copy his competitors when everyone else does is either costly pride or the only way to eventually overtake them. From here, the two look very much alike.



Join the conversation
You need an account to comment on this article. Creating one is free and takes under a minute.
No comments yet.