Physical Address
304 North Cardinal St.
Dorchester Center, MA 02124
Physical Address
304 North Cardinal St.
Dorchester Center, MA 02124

Moonshot, the Chinese company behind Kimi K3, the largest open-weight LLM available, built its model by copying Anthropic’s Fable LLM while using unauthorized chips for export to China, White House science adviser Michael Kratsios said.
Kratsios: “Large-scale secret industrial distillation aimed at stealing American technology and undermining American research is unacceptable.” booksamid reports of discussions about Ban on Chinese models with open weight That shook the artificial intelligence sector. Moonshot did not respond to questions about its training process, and Kratsios did not share more details about the sources of his claims.
Kratsios’ tweet echoed comments from Treasury Secretary Scott Bessent that “we are finding watermarks of our large American language models on many Chinese models, and this is unacceptable.” It’s not clear what these watermarks consist of, and the Treasury Department did not respond to an inquiry.
However, experts doubt that distillation – the process of querying the LLM to determine its inner workings and copying its capabilities – is responsible for the advanced capabilities displayed by the Kimi K3.
“I don’t think you get a model this powerful that fast on the heels of Fable that does strict distillation,” Braden Hancock, a researcher at the Laude Institute and co-founder of Snorkel AI, told TechCrunch. “There’s not even time, honestly, is there? Fable has only been available to the public since July 1st. You can’t extract that much data, train a model, and release it in two weeks.”
“I’ve been seeing the trickle-down effect becoming less and less over time as the Chinese models get closer to the limits and the training system transitions to reinforcement learning,” Nathan Lambert, an AI researcher at the Allen Institute for AI, said in an article. Podcast Released yesterday. “(If) that were the case, everyone would have easily been able to catch up to GLM or K3 by using their distillation data. But we haven’t done that, or won’t see that, through supervised fine-tuning alone.”
The distillation procedure requires the laboratory to systematically query the target model in order to generate data that can be used in post-training. Sometimes this explicitly involves asking the model to explain his train of thought to understand how the problems are solved. Other times, prompts and responses from the model are used to train a new model in a process called supervised fine-tuning, or SFT.
It is this process of fine-tuning that can lead to a model ostensibly created by a third party claiming to be Claude. Fine-tuning, in Lambert’s view, is where the model “acquires its morality.”
But Lambert says the benefits of SFT become less important as models become more complex. Extracting myth-like abilities will likely require reinforcement learning techniques. In many cases, this means having a proxy for the larger model evaluate the responses of the smaller model, adjusting based on the score.
More advanced technologies also require more significant infrastructure. Large reinforcement learning operations can require tens of millions of agents. Using a frontier lab’s API to do this “would be very expensive, and would likely be a time bottleneck because these models are very slow, and to be frank, they might not even give you a performance lift.”
It seems likely that previous boundary models may have contributed to the emergence of Kemi; Anthropic Publicly accused Moonshot, DeepSeek, and MiniMax systematically distilled their models earlier this year. Anthropic said it detected millions of exchanges between its models and users it identified at those companies through IP addresses and other metadata. These queries were “distinct from normal patterns of use, and reflect deliberate capacity extraction rather than legitimate use.” Anthropic did not respond to TechCrunch’s inquiries about Fable’s distillation.
However, distillation is seen as popular among AI companies, and not just in China. Elon Musk to attest Earlier this year, his company SpaceXAI distilled OpenAI models for Grok development, and that the practice was common in the industry. For example, the line between distillation and developing synthetic datasets can be somewhat blurry.
“In general, Americans understand the technical expertise of these Chinese teams,” Hancock said. “One of the founders of Moonshot was a PhD student at Carnegie Mellon University. These are forensic researchers and engineers doing serious work. … If the American models stop, I think China’s progress will slow, but it will continue. They are not just stopping here.”
It’s also difficult to separate the distillation from the second part of Kratsios’ comment – that Moonshot acquired advanced Nvidia chips and Grace Blackwell 300s, and also had access to GB300-equipped servers in Thailand. These chips are banned from being exported to China, but there is a black market, according to Sam Bresnick, a research fellow at the Georgetown Center for Security and Emerging Technology. In May, the founder of Supermicro, an American server builder, was indicted on charges of smuggling advanced chips into China.
“I’m a supporter of ‘know your customer’ laws for data centers around the world,” Bresnik said. “If you’re allowing a company to do massive training on your modern hardware, there needs to be a reporting mechanism on who that company is and what it does.”
President Joe Biden’s Department of Commerce Suggested Federal know-your-customer rules for data centers in 2024, but no further progress appears to have been made under Donald Trump. However, exporters are shipping advanced chips abroad It is supposed to guarantee They are used only for approved purposes.
When you make a purchase through the links in our articles, We may earn a small commission. This does not affect our editorial independence.