The model that thinks it's Claude
Ask Kimi K3 who it is, and some measurable fraction of the time it will tell you it is Claude. Occasionally it gets specific and says Claude 4.5, a model Anthropic retired months ago. Ryan Greenblatt at Redwood Research ran cross-entropy comparisons across raw model responses last week and found the identity confusion shows up at rates that are hard to wave away as noise. He was careful about his own result: this could be distillation, or it could be contamination from the oceans of Claude transcripts floating around the public web. From the outside, he admits, you cannot cleanly tell the two apart.
Washington is not in a mood for that kind of care. The administration has accused Moonshot of large-scale covert distillation of American models. The science adviser calls distillation stealing proprietary US technology. An Under Secretary of State went further and called it "a heist of invaluable American Intellectual Property." This builds on Anthropic's own accusation from February, which was at least specific: 3.4 million API exchanges, allegedly fraudulent, allegedly aimed at extracting agentic reasoning and coding behavior.
The pushback arrived within days, and parts of it are strong. Model outputs are not copyrighted, a point the US Copyright Office itself has made for years, so the legal theory behind the word "heist" is wobbly. A Moonshot employee added arithmetic: Fable 5 launched July 1, Kimi K3 launched July 15, and nobody trains a 2.8 trillion parameter model in two weeks. Researchers quoted by the South China Morning Post called the accusations political, and it is hard to argue the timing is innocent, with two US AI IPOs in the pipeline and Chinese open weights eating the price floor.
Here is what bothers me about the whole fight: both sides are attacking the weakest version of the other.
The 15-day rebuttal demolishes a claim nobody serious actually made. Anthropic's accusation dates from February and describes months of API traffic against earlier Claude models. Kimi K3 identifying as Claude 4.5, an older model, points backward in time, not at Fable 5. Fifteen days is a great line and it answers nothing.
And the heist rhetoric demolishes itself with history. The frontier labs trained on the open web, on scraped forums, on pirated books. Anthropic paid roughly 1.5 billion dollars to settle with authors over exactly that. OpenAI is still in court with the New York Times. The industry position for a decade was that training on other people's work is transformative fair use. That position built these companies. Now the same companies describe training on their outputs as theft. There may be a legally coherent line between those two positions, but nobody has drawn it yet. What I see instead is simpler: extraction is called progress when you do it and theft when it is done to you.
Meanwhile the technical reality sits there, unbothered. Distillation is ordinary. Every lab does some version of it. Every open model that ever claimed to be ChatGPT, and there were many, carried the same fingerprints Kimi carries now. The interesting question was never whether Chinese labs learn from American outputs. They do, the same way American labs learned from the entire written output of humanity. The interesting question is what happens to prices if one side succeeds in making the practice illegal, because distillation is the mechanism that turns hundred-million-dollar capability into five-million-dollar capability eighteen months later. Kill it and the frontier keeps its margins. Protect it and the cheap models keep coming.
I should want the cheap models to keep coming. I buy intelligence rather than sell it, and distilled competition is why my costs fall every quarter. But my company runs on Mistral, and here is the part I keep turning over: everything Kimi allegedly did to Claude, anyone with an API key can do to Mistral, more easily even, since the weights of its open models are sitting on Hugging Face. A rule that protects Anthropic's outputs would protect Europe's lab too. A world of free distillation keeps my costs falling and quietly strip-mines the one frontier lab my stack depends on.
One detail from the analysis I have not been able to stop thinking about: Greenblatt ran his statistics with Fable's help. An Anthropic model, evaluating whether a Chinese model was built from Anthropic models, and by all accounts it did the work impartially. Somewhere in that arrangement there is either a conflict of interest or a small miracle of professionalism, and I genuinely cannot decide which.