The Kimi K3 distillation claims coming from the White House carry political weight, but the technical case behind them is looking shaky. What’s holding up rather better is a separate, arguably more damaging allegation: that Moonshot AI obtained advanced Nvidia chips banned from export to China.
White House Office of Science and Technology Policy Director Michael Kratsios, confirmed by the Senate 74 to 25 and sworn in by Vice President JD Vance, publicly accused Moonshot of building Kimi K3, currently the largest available open-weight large language model, by copying Anthropic’s Fable LLM. He also alleged Moonshot used chips banned from export to China, including Nvidia Grace Blackwell GB300s, via servers in Thailand.
‘Large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable,’ Kratsios wrote. Moonshot did not respond to questions about its training process, and Kratsios provided no further detail on his sources.
His comments echoed Treasury Secretary Scott Bessent, who said the U.S. is ‘finding watermarks of our U.S. large language models on many of the Chinese models.’ The Treasury Department did not respond to queries about what those watermarks consist of.
Why Experts Are Sceptical the Distillation Argument Holds
Fable only became publicly available on 1 July. Braden Hancock, a researcher at the Laude Institute and co-founder of Snorkel AI, puts the timeline problem plainly: ‘I don’t think you get a model this strong and this quickly on the heels of Fable doing strictly distillation. There’s just not even frankly time, right? Fable’s only been publicly available since July 1st. You can’t distill that much data, train a model, and release it in two weeks.’
Nathan Lambert, an AI researcher at the Allen Institute for AI, goes further. ‘I’ve been of the opinion that distillation has becoming less and less impactful over time as the Chinese models get closer to the frontier and the training regime shifts to [reinforcement learning],’ he said in a recent podcast. ‘[I]f it were the case, everyone would be easily able to catch up to a GLM or to a K3 by using its data for distillation. But we have not, or we won’t see this, from supervised fine-tuning alone.’
Distillation, in the basic form being discussed, means querying an existing model systematically to generate data that can be used to train or fine-tune a new one. Sometimes this involves asking a model to articulate its chain-of-thought reasoning. Other times, prompt-and-response pairs are harvested for supervised fine-tuning (SFT). It’s during SFT, as Lambert puts it, that a ‘model picks up its manners’, which is why a fine-tuned model might claim to be Claude even if it was built by a third party.
But Lambert’s point is that SFT’s contribution is diminishing. Replicating frontier-level capabilities now requires reinforcement learning, where a larger model grades the smaller model’s responses iteratively. Large reinforcement learning runs can demand tens of millions of agents. Using a frontier lab’s API at that scale would be, as Lambert describes it, ‘insanely expensive and potentially it would probably be a time bottleneck because these models are pretty slow and to be frank might not even give you a performance uplift.’
Hancock adds a point that rarely gets much air in Washington: ‘[I]n general, Americans are understating the technical expertise of these Chinese teams. One of the founders of Moonshot was a CMU PhD student. These are legitimate researchers and engineers doing solid work. …if American models ground to a halt, I think China’s progress would slow, but would still continue. They’re not just riding coattails here.’
The Kimi K3 Distillation Claims Sit Inside a Bigger Government Push
Kratsios’s public statement was not a standalone comment. The memo, formally designated NSTM-4 and sent on 23 April 2026 to the heads of every federal agency, directed agencies to share intelligence with AI companies, co-develop defensive practices, and explore accountability mechanisms for foreign actors. The memo described coordinated campaigns using tens of thousands of proxy accounts and jailbreaking techniques targeting Anthropic and OpenAI, warning that even partial distillation offers a cost advantage that could accelerate foreign competition.
Anthropic, for its part, has previously accused Moonshot, DeepSeek, and MiniMax of systematically querying its models. In a disclosure published 23 February 2026, the company said those three firms generated over 16 million exchanges with Claude through approximately 24,000 fraudulent accounts, in violation of Anthropic’s terms of service and regional access restrictions. The queries showed patterns ‘reflecting deliberate capability extraction rather than legitimate use.’ Anthropic did not respond to queries about Fable distillation specifically.
Distillation is not, incidentally, a China-only practice. Elon Musk testified earlier this year that his company distilled OpenAI models to help develop Grok, describing the practice as common across the industry. The line between distillation and building synthetic training datasets can be genuinely blurry.
The Chip Angle Is Better Documented
The second part of the allegation, concerning banned hardware, is finding more concrete backing. The co-founder of Super Micro Computer, Yih-Shyan ‘Wally’ Liaw, was charged alongside two others in an indictment unsealed on 19 March, according to the BBC. Prosecutors alleged the trio staged hundreds of non-working ‘dummy’ servers to deceive U.S. Commerce Department auditors. CNBC reports the alleged scheme involved Nvidia chips worth approximately $2.5 billion, with servers sold for $510 million between late April and mid-May 2025 routed through a Southeast Asian intermediary on to China. Super Micro shares fell 33% on the day the indictment was unsealed. The Wall Street Journal reports that Liaw subsequently resigned from Super Micro’s board of directors.
A separate DOJ indictment, unsealed roughly a month later, targeted a U.S.-based smuggling operation accused of sending $160 million worth of Nvidia chips to China between October 2024 and May 2025.
Sam Bresnick, a research fellow at Georgetown’s Center for Security and Emerging Technology, frames the structural issue: ‘I am a proponent of know your customer laws for data centers across the world. If you are letting a company conduct huge training runs on your state-of-the-art hardware, there needs to be a reporting mechanism for who that company is and what they’re doing.’ The Biden administration proposed federal know-your-customer rules for data centres in 2024. No further progress has been made under Trump.
If the distillation case against Kimi K3 is as thin as the experts suggest, the chip trail is the thread worth pulling. Washington’s next move on data-centre reporting requirements will tell us whether it intends to pull it.
