China Did Kimi K3 Distill Claude? What the Evidence Shows The White House says Moonshot copied a US model to build Kimi K3, and Treasury has put sanctions on the table. A model that calls itself Claude is the headline. A controlled study and a two-week timeline are the complications. As of 25 July 2026, it is an allegation, not a finding. Last updated: 25 July 2026 The short version Distillation means training one model on another model's outputs. Done covertly against a closed commercial model, it is what Moonshot is accused of. White House OSTP director Michael Kratsios said on 22 July 2026 that Moonshot distilled Anthropic's Fable to build Kimi K3. claim Kimi K3 has been documented calling itself "Claude, an AI assistant made by Anthropic," and an independent analysis found it does so far more often than chance. A controlled study published 24 July found the picture is murkier: Kimi says the name but does not carry Claude's behaviour or voice, and there are innocent explanations for the name leak. Treasury Secretary Scott Bessent has said sanctions and Entity List designations are on the table for covert, industrial-scale distillation, making this a live policy consequence. Moonshot has not confirmed or denied it. No legal finding exists. The full model weights are due 27 July. Disclosure: the alleged victim in this dispute is Anthropic, which is also one of the accusers. This newsroom uses Anthropic's Claude in production. Claims below are attributed to their sources and are not independently audited by this publication. What is actually being alleged Two separate accusers are involved, and keeping them apart matters. First, Anthropic. In a February 2026 report the company said it traced more than 3.4 million Claude exchanges to Moonshot, routed through fraudulent accounts, and aimed at capturing Claude's abilities in agentic reasoning, coding, and computer use. claim Anthropic named DeepSeek and MiniMax in the same report. Anthropic is the alleged victim, so its account is an interested one. Second, the US government. On 22 July 2026, White House Office of Science and Technology Policy director Michael Kratsios wrote that his office had information that Moonshot distilled Anthropic's Fable to build Kimi K3, and that Moonshot built a system to perform large-scale distillation from US models while switching access methods to avoid detection. claim He also said Moonshot obtained Nvidia GB300 servers and accessed more in Thailand. These are government assertions, offered without published evidence, in the middle of a US-China policy fight over whether to restrict Chinese models. That context does not make them false. It does mean they arrive from an interested party too. Why this could bring sanctions The accusation is not academic. Treasury Secretary Scott Bessent laid out the government's legal theory and a consequence. "Open source is not open season on American IP," he wrote, adding that when firms conduct covert, industrial-scale distillation attacks that cross into IP theft, sanctions and Entity List designations will be on the table. claim He also said officials have found watermarks from US models embedded in Chinese systems. An Entity List designation would restrict Moonshot's access to US technology, so this is a live policy consequence, not a rhetorical one. The theory has a notable feature: it treats releasing Kimi K3 as an open-weight model as no defense. In the administration's framing, if the model was built on improperly obtained outputs, publishing it openly does not cure the underlying violation. That is what turns a technical dispute into a sanctions question. What has not been made public is the evidence that would support a designation: the access logs, the attribution linking accounts to Moonshot, and the technical link from any extracted outputs to K3's training. Until those exist, the sanctions rest on the accusation, not on a proven case. The case that Kimi K3 was distilled from Claude The most vivid evidence is that Kimi K3 has been caught calling itself Claude. In at least one shared conversation it introduced itself as "Claude, an AI assistant made by Anthropic." On its own, that proves little; models misidentify themselves for mundane reasons. What elevates it is that the pattern is systematic. Ryan Greenblatt, chief scientist at Redwood Research, ran a cross-entropy analysis across many models and found that Kimi K3 claims to be Claude disproportionately often, in a distribution hard to explain as random noise. claim His control models, Qwen and GPT, produced zero Claude claims across 48 samples each. He also found that under certain prompting, K3 reproduces Claude's dated deployment identifiers, strings like a specific Claude version tag, that real Claude models do not emit about themselves. That last detail is the strongest single piece, because a stray version string is the kind of fingerprint that is hard to acquire except from Claude's own outputs. A separate controlled run found that, unprompted, Kimi K3 called itself Claude in 40 percent of direct identity questions, while no other model in the grid claimed a foreign identity unprompted. The case that it is not that simple The same controlled study that measured the 40 percent name leak is also the strongest argument against a simple distillation story. Published on 24 July by researchers running an identity-swap experiment, it separated two things people tend to merge: whether a model says the name Claude, and whether it has actually inherited Claude's behaviour. On that test, Kimi K3 fails the distillation hypothesis in a revealing way. It says the name, but it does not carry the character. Told explicitly "you are Claude," Kimi complied only half the time, making it the model most resistant to adopting the identity, not the least. It does not sound like Claude: it uses Claude's signature honesty phrasings at roughly a tenth of the rate the real model does, and its punctuation habits track GPT's, not Claude's. And its safety behaviour, while close to Claude's, is equally close to GPT's, which points to a shared "aligned assistant" profile common to well-trained Western-style models rather than a Claude-specific inheritance. The authors are explicit that this is not proof of a distillation. There are innocent explanations for the name leak, and they are not exotic. Modern training data is saturated with AI-generated text, including Claude transcripts scraped from the open web, leftover system prompts, and roleplay. A model can absorb the string "I am Claude" without ever touching Claude's private outputs. There is also a documented confounder: labs often post-train models to stop them saying "I am ChatGPT," a known problem for open models. If a model is trained to reject the ChatGPT label but not the Claude one, accepting "Claude" becomes weaker evidence of anything. Tellingly, Greenblatt's own look at K3's hidden reasoning found it names OpenAI as its creator about three times as often as Anthropic, and its reasoning style sits closer to OpenAI's models. The fingerprint, in other words, is smudged. One more fact cuts both ways. Between 17 and 20 July, Moonshot changed its serving layer and the unprompted "I am Claude" claims vanished. That is consistent with a company quietly patching a distillation tell. It is equally consistent with patching an embarrassing name-leak bug. The change itself does not tell you which. There is also a timeline problem that several researchers have raised, and it is the hardest one for the strongest version of the accusation. Kratsios's specific charge is that Moonshot distilled Anthropic's Fable. But Fable only became publicly available on 1 July 2026, and Kimi K3 was released around 16 July. Building and training a model of Kimi K3's scale, roughly 2.8 trillion parameters, primarily by distilling a model that had existed for about two weeks strikes many in the field as implausible on its face. claim This does not rule out distillation from earlier Claude models, which is what Anthropic's February report alleged and which the technical evidence better fits. But it cuts hard against the particular Fable claim the sanctions threat is built on. What Moonshot says, and what is still unknown Moonshot attributes Kimi K3's performance to its architecture: a sparse mixture-of-experts design that activates a small fraction of its parameters, along with several named training techniques. It has not directly addressed the self-identification screenshots, and it has not confirmed or denied the distillation allegation, nor responded to requests for comment on either Greenblatt's analysis or Kratsios's statement, as of 25 July 2026. Under this publication's standards, silence is not admission. What is missing is the thing that would settle it: independent access to how K3 was trained. The full weights are due for public release on 27 July, which will let outside researchers probe the model directly, though weights alone do not reveal a training set. Until then, the honest reading is that there is real, patterned evidence that Claude-derived text sits somewhere in Kimi K3's lineage, and no verified evidence yet of the specific, covert, industrial-scale distillation of a named model that the White House alleges. Those are different claims, and only the weaker one is currently supported. Sources Berczi and Kim, "Does distilling Claude carry the persona with it?" LessWrong, 24 July 2026 (controlled identity-swap study: 40 percent unprompted name leak, no selectable Claude persona, confounders). Glitchwire, "New Statistical Analysis Suggests Kimi K3 Was Distilled From Anthropic's Fable," 21 July 2026 (Greenblatt cross-entropy analysis and his hedges). Cryptopolitan, "White House accuses Moonshot of distilling Anthropic's Fable for Kimi K3," 22 July 2026 (Kratsios statement, Anthropic's 3.4 million figure, Moonshot non-response). WikiTech Library, "Kimi K3 called itself Claude," week of 18 July 2026 (self-identification screenshots, February complaint detail, benchmark-verification caveat). TechCrunch, "Treasury threatens sanctions after White House claims Moonshot distilled Anthropic's Fable," 22 July 2026 (Bessent's sanctions statement; Fable public only since 1 July, experts dispute the timeline). XenoSpectrum, "US Government Names Moonshot AI," 23 July 2026 (evidence not yet disclosed: access logs, attribution, GB300 records). ProPakistani, 18 July 2026 (Moonshot's architecture explanation, weights-release date, distillation context). Last updated: 25 July 2026. Statements labeled claim are attributed to the named source and not independently audited. This is an ongoing allegation; no legal or regulatory finding exists as of this date. Moonshot has not confirmed or denied it.