Artificial Intelligence thread

Topazchen

Junior Member
Registered Member
Would be funny if this ends up saving Microsoft’s flagging stocks.

Please, Log in or Register to view URLs content!

Who am I kidding, it most likely will!!!
It's kind of ironic and funny that Microsoft shut down its AI lab in Shanghai because they didn't want Chinese researchers getting their hands on cutting edge AI tech... Lmaoooo
 

9dashline

Major
Registered Member
I never said the model is bad. In fact is quite good. I'm hypothesizing that they are trying to recreate a Mythos type level model that when you remove the guardrails becomes a terminator type cyberweapon, that can hack or even do serious damage (if have how) on its own. OpenAI are looking to meet Pete Hegseth and Palantir psychopathic demands of unrestricted Warfare, something that Dario, to his credit, refuse to do. I'm hypothesizing because of that the model is behaving more erratically because is not fully polished.
GPT 5.6 has "accidentally" hard wiped a lot of peoples harddrives... no joke
 

Eventine

Senior Member
Registered Member
Re: Director Michael Kratsios's claim that Kimi distilled Fable

My read on this is:
  • Fable was only available for ~1 month before K3 was released, and for the first 2-3 weeks it was a limited, "trusted enterprises" release. If Kimi was able to get access to Fable outputs during the initial release, then the "trusted enterprises" ecosystem is deeply infiltrated by leakers. If the White House truly believes this, we can expect much more careful vetting of users from now on; you'd probably need a special clearance to get access to powerful US models, and knowing Trump, they'd probably want to make sure you're not Chinese (and if you're a Chinese researcher working for an US lab, expect the CIA/FBI monitoring you from now on).
  • If K3 was distilled instead from the publicly available Fable, then it means there's only a ~2 weeks gap between when a frontier model releases and when a distilled version of it can exist. While technically possible as a post-training fine-tuning step, this kind of distillation is essentially impossible to prevent for any public model release. The US government can complain about it all it wants, but the only practical alternative is to not release the model in the first place, which of course, destroys its addressable market and makes it that much more likely China will take over global consumer-facing AI. This is a lose-lose situation for the US.
  • There is the final possibility that K3 didn't truly distill Fable (in the sense of relying on its outputs for the success of the model; I don't think you can definitively prove K3 never used any Fable generated output in its training set as that's all over the internet), and this is all just poisoning the well in preparation for a wide-ranging ban on Chinese models and model services for the US and its allies. I find this explanation the most likely and plausible as it's been the administration's stated goal for the last few months, ever since GLM 5.2 released. However, it doesn't change the fundamental calculus as they were going to do it any way, so however they justify it is irrelevant.
Personally, the only definitive way of disproving Chinese frontier labs are dependent on US frontier labs is if China releases a better model than the US has available (publicly). This is why I am hoping for a Seedance moment for frontier LLMs, since no one can truly claim ByteDance distilled Sora or Veo when it is superior to both.

Should China accomplish this, the US's accusations will seem silly and it'd further erode US credibility and claims of superiority, which will be beneficial to global adoption of Chinese models; hence I do believe it is worth China's while to actually invest in it just to end the debate once & for all.

On the counter side, even if it is true that Chinese frontier labs are dependent on US frontier labs for distillation data, China will probably still win out in consumer-facing models in the end, because the business strategy is just so much better and it will crash the US stock market if the US can't publicly release models without it being quickly distilled. But the risk there would be if the US bank rolls Anthropic/Open AI to develop completely closed, secret models that China cannot replicate and uses them for military/intelligence/scientific purposes to get ahead. So I do hope Chinese labs have independent abilities to keep up on the frontier, because the biggest take-away from the White House's message is that the US is going to be cracking down on Chinese models and public access to US frontier models very soon.
 

SanWenYu

Major
Registered Member
Re: Director Michael Kratsios's claim that Kimi distilled Fable

My read on this is:
  • Fable was only available for ~1 month before K3 was released, and for the first 2-3 weeks it was a limited, "trusted enterprises" release. If Kimi was able to get access to Fable outputs during the initial release, then the "trusted enterprises" ecosystem is deeply infiltrated by leakers. If the White House truly believes this, we can expect much more careful vetting of users from now on; you'd probably need a special clearance to get access to powerful US models, and knowing Trump, they'd probably want to make sure you're not Chinese (and if you're a Chinese researcher working for an US lab, expect the CIA/FBI monitoring you from now on).
  • If K3 was distilled instead from the publicly available Fable, then it means there's only a ~2 weeks gap between when a frontier model releases and when a distilled version of it can exist. While technically possible as a post-training fine-tuning step, this kind of distillation is essentially impossible to prevent for any public model release. The US government can complain about it all it wants, but the only practical alternative is to not release the model in the first place, which of course, destroys its addressable market and makes it that much more likely China will take over global consumer-facing AI. This is a lose-lose situation for the US.
  • There is the final possibility that K3 didn't truly distill Fable (in the sense of relying on its outputs for the success of the model; I don't think you can definitively prove K3 never used any Fable generated output in its training set as that's all over the internet), and this is all just poisoning the well in preparation for a wide-ranging ban on Chinese models and model services for the US and its allies. I find this explanation the most likely and plausible as it's been the administration's stated goal for the last few months, ever since GLM 5.2 released. However, it doesn't change the fundamental calculus as they were going to do it any way, so however they justify it is irrelevant.
Personally, the only definitive way of disproving Chinese frontier labs are dependent on US frontier labs is if China releases a better model than the US has available (publicly). This is why I am hoping for a Seedance moment for frontier LLMs, since no one can truly claim ByteDance distilled Sora or Veo when it is superior to both.

Should China accomplish this, the US's accusations will seem silly and it'd further erode US credibility and claims of superiority, which will be beneficial to global adoption of Chinese models; hence I do believe it is worth China's while to actually invest in it just to end the debate once & for all.

On the counter side, even if it is true that Chinese frontier labs are dependent on US frontier labs for distillation data, China will probably still win out in consumer-facing models in the end, because the business strategy is just so much better and it will crash the US stock market if the US can't publicly release models without it being quickly distilled. But the risk there would be if the US bank rolls Anthropic/Open AI to develop completely closed, secret models that China cannot replicate and uses them for military/intelligence/scientific purposes to get ahead. So I do hope Chinese labs have independent abilities to keep up on the frontier, because the biggest take-away from the White House's message is that the US is going to be cracking down on Chinese models and public access to US frontier models very soon.
Not that I believe what OpenAI, Anthropic and the US politicians are saying, but Google search's AI mode says that same-size distillation can result in student model outperforming the teacher model with better accuracy.
 

bsdnf

Senior Member
Registered Member
Not that I believe what OpenAI, Anthropic and the US politicians are saying, but Google search's AI mode says that same-size distillation can result in student model outperforming the teacher model with better accuracy.
The premise is that the teacher model provides a complete CoT, while OpenAI and Anthorpic have hiding it, where users only see a summary.
 

tphuang

General
Staff member
Super Moderator
VIP Professional
Registered Member
Re: Director Michael Kratsios's claim that Kimi distilled Fable

My read on this is:
  • Fable was only available for ~1 month before K3 was released, and for the first 2-3 weeks it was a limited, "trusted enterprises" release. If Kimi was able to get access to Fable outputs during the initial release, then the "trusted enterprises" ecosystem is deeply infiltrated by leakers. If the White House truly believes this, we can expect much more careful vetting of users from now on; you'd probably need a special clearance to get access to powerful US models, and knowing Trump, they'd probably want to make sure you're not Chinese (and if you're a Chinese researcher working for an US lab, expect the CIA/FBI monitoring you from now on).
  • If K3 was distilled instead from the publicly available Fable, then it means there's only a ~2 weeks gap between when a frontier model releases and when a distilled version of it can exist. While technically possible as a post-training fine-tuning step, this kind of distillation is essentially impossible to prevent for any public model release. The US government can complain about it all it wants, but the only practical alternative is to not release the model in the first place, which of course, destroys its addressable market and makes it that much more likely China will take over global consumer-facing AI. This is a lose-lose situation for the US.
  • There is the final possibility that K3 didn't truly distill Fable (in the sense of relying on its outputs for the success of the model; I don't think you can definitively prove K3 never used any Fable generated output in its training set as that's all over the internet), and this is all just poisoning the well in preparation for a wide-ranging ban on Chinese models and model services for the US and its allies. I find this explanation the most likely and plausible as it's been the administration's stated goal for the last few months, ever since GLM 5.2 released. However, it doesn't change the fundamental calculus as they were going to do it any way, so however they justify it is irrelevant.
Personally, the only definitive way of disproving Chinese frontier labs are dependent on US frontier labs is if China releases a better model than the US has available (publicly). This is why I am hoping for a Seedance moment for frontier LLMs, since no one can truly claim ByteDance distilled Sora or Veo when it is superior to both.

Should China accomplish this, the US's accusations will seem silly and it'd further erode US credibility and claims of superiority, which will be beneficial to global adoption of Chinese models; hence I do believe it is worth China's while to actually invest in it just to end the debate once & for all.

On the counter side, even if it is true that Chinese frontier labs are dependent on US frontier labs for distillation data, China will probably still win out in consumer-facing models in the end, because the business strategy is just so much better and it will crash the US stock market if the US can't publicly release models without it being quickly distilled. But the risk there would be if the US bank rolls Anthropic/Open AI to develop completely closed, secret models that China cannot replicate and uses them for military/intelligence/scientific purposes to get ahead. So I do hope Chinese labs have independent abilities to keep up on the frontier, because the biggest take-away from the White House's message is that the US is going to be cracking down on Chinese models and public access to US frontier models very soon.

why would you give any credence to this nonsense?

distillation on a small scale is very common in the AI industry. Elon even said this recently. Most of the so called distillation from Chinese labs of Claude models were done against Opus months ago. The Chinese models are now good enough and getting enough hard problem through these API calls that they generate enough synthetic data to no longer need distillation.

Think about it this way:
if you distill by passing a problem to Fable and then get output/solution. You can get a just as good solution/output from running it against GLM-5.2 or Kimi-3. So, do you actually need to run Closed source models from distillation? It seems you don't.

Now, you probably do want to do api calls against Fable and GPT-5.6 models for benchmark purposes. But it's hard for me to see justification that distillation of Fable at this point can still lead to significant improvement to model development.

This is why I keep talking about RSI. If you achieve self improvement, then you don't need to distill off other people's models.
 
Top