Re: Director Michael Kratsios's claim that Kimi distilled Fable
My read on this is:
- Fable was only available for ~1 month before K3 was released, and for the first 2-3 weeks it was a limited, "trusted enterprises" release. If Kimi was able to get access to Fable outputs during the initial release, then the "trusted enterprises" ecosystem is deeply infiltrated by leakers. If the White House truly believes this, we can expect much more careful vetting of users from now on; you'd probably need a special clearance to get access to powerful US models, and knowing Trump, they'd probably want to make sure you're not Chinese (and if you're a Chinese researcher working for an US lab, expect the CIA/FBI monitoring you from now on).
- If K3 was distilled instead from the publicly available Fable, then it means there's only a ~2 weeks gap between when a frontier model releases and when a distilled version of it can exist. While technically possible as a post-training fine-tuning step, this kind of distillation is essentially impossible to prevent for any public model release. The US government can complain about it all it wants, but the only practical alternative is to not release the model in the first place, which of course, destroys its addressable market and makes it that much more likely China will take over global consumer-facing AI. This is a lose-lose situation for the US.
- There is the final possibility that K3 didn't truly distill Fable (in the sense of relying on its outputs for the success of the model; I don't think you can definitively prove K3 never used any Fable generated output in its training set as that's all over the internet), and this is all just poisoning the well in preparation for a wide-ranging ban on Chinese models and model services for the US and its allies. I find this explanation the most likely and plausible as it's been the administration's stated goal for the last few months, ever since GLM 5.2 released. However, it doesn't change the fundamental calculus as they were going to do it any way, so however they justify it is irrelevant.
Personally, the only definitive way of disproving Chinese frontier labs are dependent on US frontier labs is if China releases a better model than the US has available (publicly). This is why I am hoping for a Seedance moment for frontier LLMs, since no one can truly claim ByteDance distilled Sora or Veo when it is superior to both.
Should China accomplish this, the US's accusations will seem silly and it'd further erode US credibility and claims of superiority, which will be beneficial to global adoption of Chinese models; hence I do believe it is worth China's while to actually invest in it just to end the debate once & for all.
On the counter side, even if it is true that Chinese frontier labs are dependent on US frontier labs for distillation data, China will probably still win out in consumer-facing models in the end, because the business strategy is just so much better and it will crash the US stock market if the US can't publicly release models without it being quickly distilled. But the risk there would be if the US bank rolls Anthropic/Open AI to develop completely closed, secret models that China cannot replicate and uses them for military/intelligence/scientific purposes to get ahead. So I do hope Chinese labs have independent abilities to keep up on the frontier, because the biggest take-away from the White House's message is that the US is going to be cracking down on Chinese models and public access to US frontier models very soon.