A technique called model distillation, long used to shrink large AI systems into cheaper ones, has become a new flashpoint between Washington and Beijing. U.S. firms accuse Chinese rivals of using it to extract capabilities from proprietary models without consent, while Chinese researchers have also used the same method on U.S. models in public research projects.
Washington and leading U.S. AI firms accuse Chinese rivals of using model distillation to pull capabilities out of proprietary models without permission, opening a new front in the U.S.-China contest over AI dominance. The technique trains a smaller "student" model on the outputs of a larger "teacher" model, producing a system that runs on less powerful hardware.
How the technique works
The teacher model generates examples, including answers and computer code, and the student model trains on that output. The result is not a copy of the teacher: it does not inherit the teacher's weights, architecture or full range of abilities, but instead learns select behaviors that let it handle specific tasks more efficiently. Because a distilled model can run on cheaper hardware, it appeals to companies and governments that want to deploy AI in devices, factories, vehicles and private networks without the cost of a frontier system.
Reasoning traces raise the stakes
Newer AI systems have made it possible to transfer not just final answers but the reasoning steps behind them, known as "reasoning traces". ETH Zurich machine-learning security researcher Florian Tramèr compared learning from full worked solutions to learning from answers alone, saying it is far easier for a student to progress with the steps shown. As reasoning traces have grown more valuable, access to a model's outputs has become correspondingly more sensitive, since those traces can expose the methods advanced systems use to solve hard problems.
Accusations run largely one way
Distillation itself is a standard, widely used training method: Stanford's Alpaca project and Microsoft's Orca research both relied on outputs from more advanced models to improve smaller ones, and Chinese researchers have likewise used U.S. model outputs in public projects, including efforts to build Chinese-language instruction models. The dispute centers on access rather than the method: open-weight models let outsiders inspect and modify their parameters, while closed models such as OpenAI's ChatGPT and Anthropic's Claude stay under company control.
Anthropic has accused Chinese entities including DeepSeek, Moonshot and MiniMax of running large-scale campaigns to obtain software-engineering and advanced-reasoning capabilities from Claude, and OpenAI has said it detected similar attempts by Chinese actors against its own models. No Chinese company has yet made the reverse accusation against a U.S. rival.
Source: Reuters
Trading involves risk.