On August 5, 2026, The Information reported that Zhang Yiming, the founder of ByteDance, made it clear at the Seed staff meeting held a month ago that ByteDance would not use model distillation to accelerate its ability to improve its big language model.
According to the analysis of the report, ByteDance refused to distill the model of the American cutting-edge AI laboratory out of fear that the distillation would lead to retaliation in Washington and endanger the company’s high-profile social media asset TikTok.
And what we have learned is that TikTok’s prospects account for almost zero weight in the considerations of Seed’s distillation of cutting-edge models in the United States.
At the Seed plenary session, Zhang Yiming explained the reason for prohibiting distillation of other models, saying, “In order to achieve long-term goals, we should be willing to sacrifice some short-term benefits.
The ‘long-term goal’ here can easily be interpreted as TikTok’s future destiny. However, among the priorities of Zhang Yiming, who has removed the daily management of ByteDance, the current importance of TikTok is far less than the pursuit of real intelligence of the model. Therefore, this sentence should be interpreted as: In order to pursue the infinite intelligence of the model, distillation should be prohibited at this stage.
We can accept temporary backwardness, but don’t distill it, “insiders told us. Zhang Yiming has made such statements multiple times within Seed.
In the past six months, Anthropic, a cutting-edge AI laboratory in the United States, has publicly accused Chinese laboratories such as Tongyi Qianwen, DeepSeek, Darkface of the Moon, and MiniMax of using large-scale fake accounts and proxy networks to obtain Claude’s output and improve their model capabilities. Model distillation is transitioning from a common training method to an increasingly sensitive term in the competition between Chinese and American artificial intelligence.
But as early as 2023, when “distillation” is far from becoming the key word of Sino US AI competition, ByteDance has made clear the opposition to distillation.
At that time, ByteDance’s large model team was still in the early exploration stage, and some engineers had used the data generated by GPT API for experimental research on a smaller model. In April 2023, after ByteDance introduced the GPT API call specification check, the model team immediately put forward a clear requirement: data generated by GPT should not be added to the training dataset of ByteDance model. Afterwards, another round of internal inspection was conducted to sample and test the similarity between the model output and GPT, in order to prevent data annotators from using GPT without authorization.
ByteDance’s vigilance towards distillation training data clearly predates today’s AI competitive narrative around distillation.
And this internal regulation did not end the debate. As model competition intensifies, researchers will still come up with distillation solutions when specific projects encounter capability or computing power gaps.
As far as we know, there have been at least three more significant debates within Seed regarding the “distillation” route since 2025.
The first debate occurred in January 2025, when Anthropic had not yet publicly accused Chinese laboratories.
After the release of DeepSeeker R1, it demonstrated powerful reasoning ability with limited training costs, shaking Silicon Valley and putting enormous pressure on other leading laboratories in China. Some researchers within Seed are particularly anxious about this. In their view, Seed’s talent density, computing power investment, and real research ability are not inferior to DeepSeek, but the reasoning ability demonstrated by the model at that time did not reach the level of R1.
At that time, there was already speculation in the industry that DeepSeek had used generated data from cutting-edge models in the United States during the training process, but this judgment has not been confirmed to this day, except for Anthropic’s unilateral accusation.
But within Seed, it poses a very realistic question: since peers can quickly improve performance by distilling the American SOTA model, why can’t Seed moderately adopt the same approach?
Some researchers suggest introducing the generation results of the US head closed source model to improve the inference and comprehensive performance of the Seed language model. This does not require giving up one’s own pre training, nor does it mean that the entire model is built on someone else. It is more like a quick acting tonic: retaining the original training system while using the output of a stronger ‘teacher model’ to compensate for the exposed ability gap.
But Seed ultimately rejected the proposal internally.
One year later, the debate over whether to “distill” occurred for the second time during the large-scale deployment of Nvidia Blackwell series GPUs for commercial use in the United States.
This time, Seed faced more specific pressure.
NVIDIA’s Blackwell series GPUs will enter the top laboratories in the United States on a large scale by the end of 2025; Due to the targeted export control of high-end GPUs by the United States, Chinese laboratories are unable to obtain the most advanced performance “B cards”.
An unknown piece of information is that Seedance 2.0, which established the global SOTA status of the ByteDance video model, was trained in a computing environment built on a large number of Nvidia H20 stacks. H20 is a “compliant castrated version” specifically designed for the Chinese market, with a comprehensive performance for model training that is only 1/50 of B200.
In any case, the existing computing power gap between American and Chinese laboratories will be sharply widened by a new GPU generation at the end of 2025 and the beginning of 2026.
Seed noticed an unsettling fact internally: from GPT-5.5 to Claude 4.8, top models in the United States have made intelligent leaps in reasoning, programming, and science, all closely related to the large-scale deployment of Blackwell series GPUs. Especially the intelligence demonstrated by Claude Fable 5 left a deep impression on Seed researchers.
The motion of ‘distillation if it really doesn’t work’ has reappeared and has received support from many researchers. This has become a gradually spreading emotion internally, especially in the case where the self-developed Seedance 2.0 is already one of the best performing video generation models in the world, while the Seed series language models are still relatively backward.
Researchers who support distillation believe that since the huge gap in computing power cannot be eliminated in the short term, data should be exchanged to concentrate extremely scarce and inherently insufficient training resources on proven effective directions.
This proposal still has not received support from the top management of the company, but there is a growing trend within Seed to become a consensus.
At this point, the external pressure of Seed’s internal “distillation” initiative is still only catching up with the leading SOTA model in the United States. A few months later, the success of Kimi K3 truly sparked this discussion.
K3’s performance in programming, tool calling, in-depth research, and complex tasks has entered the same competitive range as the top closed source models in the United States. For Seed, this pressure is even more direct than R1.
A laboratory from China, with a scale and computing power resources far smaller than ByteSeed, has already made its open weight language model enter the world’s top camp; Despite investing more and having more talent in Seed, we still haven’t come up with a language model that is equally positioned.
The proposal of distillation was loudly raised for the third time and received more responses from people.
The proposer believes that as long as the generated results of the top-level SOTA model are systematically used, the performance of the Seed language model can be easily improved in the short term. It does not need to replicate the full capabilities of a particular model, but only needs to have the strong model generate enough inference, code, tool calls, and complex task data to quickly fill in the parts that most affect the ranking and user perception.
This time, a seemingly more secure compromise solution has emerged internally: since distilling the closed source model of the US laboratory may violate the terms of service, face public accusations, or be involved in geopolitical disputes. So, Seed can at least distill the best performing open weight model.
The authorization of open weight models is usually more lenient, and the models can be deployed on their own servers without the need for bulk account registration or bypassing geographical and access restrictions. It can also generate high-quality training data with much lower legal and commercial risks.
This time, they waited for Zhang Yiming’s direct statement.
Zhang Yiming said: Closed source models cannot distill, and open weight models cannot distill either.
He also stated that Seed can accept temporary lag, but does not accept relying on distillation competitors to eliminate this lag.
And as far as we know, this is also one of the backgrounds for Seed’s convening of the all staff meeting.
Zhang Yiming hopes for a unified understanding within Seed and a clearer technical judgment: the key is not short-term gains or losses, but rather the fundamental work and research that truly affect the long term, whether they are truly valued and done well.
The direct statement from the founder of ByteDance is that after the launch of K3, Seed’s internal attitude to whether other models can be distilled has shown unprecedented vacillation and position loosening.
After this plenary session, Seed has introduced a new policy internally: explicitly prohibiting the distillation of open weight models such as K3, and tracing and investigating suspected distillation behavior through API technology testing and other means.
At this point, the internal debate and decision on distillation within Byte has become a consensus, and it has also transcended the significance of legal, commercial terms, and geopolitical AI competition.
Model distillation itself has no original sin. It is a very mature training technique: allowing a ‘student model’ to learn the probabilities, answers, reasoning trajectories, or tool calling processes generated by the ‘teacher model’, reproducing some of the abilities of the teacher model at a lower cost. All head laboratories conduct self distillation and use their own models to generate synthetic data, which is then sent back to the training system.
Seed rejects the source of this ability, and top-level performance models – whether open source or closed source, cannot become its teachers.
When a laboratory finds itself lagging behind, it can study competitors’ papers, analyze their technical reports, understand their evaluation results, and re-examine its data, architecture, and training methods.
It can also turn competitors directly into “teachers”, mass produce answers, and then make its own models move forward along the path that the other party has already taken.
The latter path is usually faster.
It can enable a model to complete its reasoning, programming, and agent capabilities within a few months, narrow the benchmark gap, and also enable the product to enter the first camp faster. But it cannot tell the laboratory exactly how these abilities are generated; It cannot be guaranteed where the distillation side should go next when the intelligent boundary of the current SOTA model reaches its end.
This is the reason why the distillation motion has been repeatedly put forward and rejected in ByteDance.
Distillation is not a controversial technology, but a shortcut that is effective enough, enticing enough, and capable of changing the long-term research culture in laboratories.
Inside the byte, Seed wanted to use it for the first time to catch up with the inference shock brought by R1; The second time I want to use it to bridge the computing power gap between H20 and Blackwell; The third time I want to use it to narrow the gap with Kimi K3 and the US head closed source model.
Three proposals were made when Seed did have a reason to choose shortcuts, but they were all rejected.
Behind it is a route debate that has lasted for over a year. The prospects of TikTok, Anthropic’s accusations, and the increasing policy risks in the United States certainly make it easier to explain the rejection of distillation, but they are not enough to constitute principles. The technical principle of ByteDance non distillation SOTA model originated before these risks and external pressures.
Zhang Yiming’s biggest concern is probably whether Seed, as an AI laboratory, can make its own achievements in pursuing cutting-edge intelligence and reach the ceiling of global model competitiveness.
At least one thing is clear: a model can enter the forefront through distillation, but a laboratory cannot become a cutting-edge laboratory through distillation.
This means that Seed may fall behind at some point, and it also means that it has to pay more computing power, time, and failure costs for independent exploration. But now, they are all acceptable.