Parametric scale of domestic large models is continuing to rise beyond 10 trillion. Qwen3.8 Max and Kimi K3 have already reached 2.4 trillion and 2.8 trillion parameters, respectively, and the next step could be 5 trillion. This time, ByteDance may take the lead. LatePost reported that ByteDance is discussing the training of a large model with more than 5 trillion parameters, which would become the largest known model in China. However, the plan is still in its early stages, and it may not ultimately be released.

What is the scale of 5 trillion parameters? It is considered to be at the level of GPT-5.6 or Opus 5. Generally, the larger the number of parameters, the stronger the performance of a large model. OpenAI and Anthropic have maintained their leading positions by having higher parameter counts—reports suggest that both companies are training or have completed preparations for 10 trillion parameter large models. However, extremely high parameter counts also mean extremely high computing power investments. Based on H100, training a 5 trillion parameter model would require 100,000 H100s running for 347 days, or one million graphics cards running for 35 days. A million graphics cards have entered the GW-level computing power scale, with extremely high barriers. ByteDance will not be able to invest such a huge amount of computing power all at once. A more reasonable approach would be to train with 200,000 to 300,000 graphics cards for 3 to 4 months, then add preparation and post-training phases, which would take at least half a year, or even a year, before it could be completed.

In terms of computing power reserves, OpenAI currently has the most abundant resources. SemiAnalysis estimates that it has about 2 million graphics cards; Anthropic has about 620,000; DeepSeek only has around 60,000, many of which are relatively outdated H20 and H800. With its financial strength, ByteDance can easily feed a 5 trillion parameter large model. With it, Dou Bao's capabilities could reach the level of GPT-5.6 or Opus 5 without doubt, truly maximizing its intelligence, no longer needing to be mocked by netizens. The only remaining uncertainty is commercial returns—can this astronomical investment bring equivalent returns in practical applications?