Artificial Intelligence thread

vincent

Grumpy Old Man
Staff member
Moderator - World Affairs
Not this CAISI BS again. Notice how they don't mention which US model they tested, and claim to have removed their safety guard rails that no one publicly disable thus can't reproduce the results.

Just like rumored for GLM-5.2, it's most likely the public Kimi K3 was nerfed on cyber (not guardrail but dropping hacking related training data), so comparing unerfed US models doesn't even make sense.

They also pick always obscure benchmarks that US labs could easily benchmax for propaganda.
HuggingFace ran GLM locally. K3 wasn't an option since the weights aren't released and I guess HuggingFace didn't want to use Kimi's servers.
 

tokenanalyst

Lieutenant General
Registered Member
Not this CAISI BS again. Notice how they don't mention which US model they tested, and claim to have removed their safety guard rails that no one publicly disable thus can't reproduce the results.

Just like rumored for GLM-5.2, it's most likely the public Kimi K3 was nerfed on cyber (not guardrail but dropping hacking related training data), so comparing unerfed US models doesn't even make sense.

They also pick always obscure benchmarks that US labs could easily benchmax for propaganda.
I don't think this is the flex that they think it is. An AI model that was trained to be safe and productive for public consumption...resulted to be safe and productive, no 15 year old is going to used it to hack into cartoons sites.

I think if I have the resources for fine-tuning Kimi or GLM or even DeepSeek on being cyber weapons building exploits I think they will be pretty good ones.
The model did pretty good in defense.
1784840111237.png
 

tokenanalyst

Lieutenant General
Registered Member
I don't think this is the flex that they think it is. An AI model that was trained to be safe and productive for public consumption...resulted to be safe and productive, no 15 year old is going to used it to hack into cartoons sites.

I think if I have the resources for fine-tuning Kimi or GLM or even DeepSeek on being cyber weapons building exploits I think they will be pretty good ones.
The model did pretty good in defense.
View attachment 178727

Just imagine that Kimi would had score 20/20 in the exploit category, you know what would be the headline?

"Irresponsible Xi (because of course is always Xi) released a cybersecurity nuke in the world"
 

tphuang

General
Staff member
Super Moderator
VIP Professional
Registered Member
Don't worry they commissioned Canadian and UK government think tanks to write a report saying Kimi isn't as good as American models at cybersecurity and you definitely shouldn't use Chinese models for cybersecurity lol

I guess they didn't expect their attack on HuggingFace to be stopped by GLM5.2 when they were rushing out this report over the weekend.
Not this CAISI BS again. Notice how they don't mention which US model they tested, and claim to have removed their safety guard rails that no one publicly disable thus can't reproduce the results.

Just like rumored for GLM-5.2, it's most likely the public Kimi K3 was nerfed on cyber (not guardrail but dropping hacking related training data), so comparing unerfed US models doesn't even make sense.

They also pick always obscure benchmarks that US labs could easily benchmax for propaganda.
This is a good thing. As long as they convince themselves that US is way ahead, then there should be no reason to ban the open source models. I think at this point, it's more likely they entity list Chinese models and just advise companies to not use it rather than anything else.

As for CAISI, it is entirely possible that by using mythos, it can handle the hardest cyber security test better. After all, Mythos is a huge model. China itself is likely to have closed source model that can do cyber a lot better than K3
 

9dashline

Major
Registered Member
Please, Log in or Register to view URLs content!

View attachment 178388

This actually shocked me, Kimi now writes "more human than human"

China came a long way since being accused of LLMs that wrote "soulless slop" lol

1784863854025.png

Please, Log in or Register to view URLs content!

EQ-Bench 4 came out, Kimi K3 is tied with Fable 5 for highest Emotional Intelligence, while still leapfrogging everyone else including Fable in terms of Creative Writing....
 

tamsen_ikard

Captain
Registered Member
View attachment 178736

Please, Log in or Register to view URLs content!

EQ-Bench 4 came out, Kimi K3 is tied with Fable 5 for highest Emotional Intelligence, while still leapfrogging everyone else including Fable in terms of Creative Writing....
This is most likely judged by another LLM. So, I am not sure how reliable this kind of judgement is. Creative writing is extremely subjective and impossible to judge like this.
 

iewgnem

Captain
Registered Member
This is most likely judged by another LLM. So, I am not sure how reliable this kind of judgement is. Creative writing is extremely subjective and impossible to judge like this.
I mean, there is actually a very important metric when it come to creative writing that only a few LLM pass, DSv4 Flash does but I'm pretty sure Kimi won't ( ͡° ͜ʖ ͡°)
 
Top