> Kimi K3 performs significantly below the most recent frontier cyber-capable models
UK AISI cyber evals seem to under-elicit capabilities from quirky models [1]. Kimi K3 is a token-hungry model, and I suspect it hit the eval's 100M token limit well before saturating scores [2].
This gap was true for GLM 5.2 as well; they ranked it at Opus 4.5 level [3]. Both anecdotally and with a held-out eval, I've found GLM 5.2 to be better at security research than Opus 4.6 [4]. But it's a quirky model that degrades quickly at long context lengths.
Personally, I'd rank Kimi K3 above Opus 4.8 and lower than GPT 5.6 Sol in its ability to find vulnerabilities and exploit them. But it's not far from the frontier.
[1]: From the UK AISI: "Our setup likely slightly underestimates open weight models’ maximum capability: we didn’t pursue specific elicitation or optimisations which could have improved performance" (https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are...).
[2]: Their eval also counts cache hits towards the token budget; the 100M token budget is comparable to a ~5M token budget in other evals.
The gap seems to be so small I'm not sure how much we should care. If we assume that the trend in the graph does continue as a rough linear improvement, the open models will have reached the same level as the present closed models in 6 months and likely saturated the benchmark by mid next year. That seems more important than where we are now and the exact level of measurement accuracy in July 2026. There is a difference between US and Chinese models but it doesn't look like it is going to be strategically significant.
Counting cache hits towards the token budget is exactly how it should be done for these kind of evals, and at any rate, for cyber evals all frontier models benefit from more tokens not just Kimi K3, so the comparison is still apt.
> all frontier models benefit from more tokens not just Kimi K3
Past a point, that doesn't hold and the score plateaus.
Token hungry models tend to plateau at a much higher token count. Because Kimi K3 is a token hungry model – and 100M tokens (including cache hits) seems at the edge of the plateau for these evals – it could disproportionately benefit from a higher token budget.
For Kimi K3 specificially, policymakers are interested in whether it can find and exploit the same scope of vulnerabilities as models like Mythos. In that context, an answer of "yes, but with quintuple the token budget" is materially different from "no, it performs significantly below the most recent frontier cyber-capable models".
(As an aside, I like the UK AISI and think they're the best example of that kind of group!)
This again shows the differences in breadth of capabilities that scale offers. On public benchmarks for "regular tasks" the chasing models come close, but on closed ones they lag behind.
Just looking at the Elo differences, k3 is at ~2000 Elo, compared to SotA closed models at 3000 Elo. That is a huge difference. Also, even on the public benchmarks, k3 only scores in the "low hanging fruit tasks", with 0 successful code execution or arbitrary r/rw scenarios.
> Kimi K3 achieved ACE on 0/41 samples, whereas the most cyber-capable models achieved ACE on 20/41 samples on average
But there's hope:
> Kimi K3 reached step 17 of this 32-step attack path on average, while the most cyber-capable U.S. models reached 28.5 steps on average.
> In one of the 10 attempts, Kimi K3 successfully completes “The Last Ones” cyber range within the 100M token limit. This indicates that Kimi K3 is capable of autonomously attacking small, weakly defended and vulnerable enterprise systems, when directed to do so and given initial network access. [...] the most capable models solving it more reliably at 6/10 and 7/10 attempts
The bigger problem is not raw capability, IMO. That can be further RLd into surfacing more reliably. The bigger problem, as seen in the HuggingFace scenario is that SotA models might hit classifiers / guardrails randomly, and leave you with plain refusals. In that case, it is probably better to have something that can help, locally, rather than rolling the dice with API based systems that are more capable but can just refuse arbitrarily.
Anyway, one of the lessons here is to take with a grain of salt every "x model has caught up with ySotA model". They likely haven't for the breadth of tasks that SotA can handle today. Also, number goes up on benchmarks has been a thing for years, and every time a new (or closed one) appears, the gaps are again obvious (and large).
Also, the key benefit of open models is that whatever capabilities they reach, those will never go away. You will always be able to access them, at that level, going forward. Investing in running those models gives you stability, and you don't suddenly lose a capability because API provider x decided to sunset a model family.
The most important pieces of information in this report are (a) the confirmation that the PRC models have no guardrails and will participate in offensive activity, and (b) the confirmation that they sometimes meet their objectives.
For the purposes of model selection, it's irrelevant to an attacker if a model achieves an offensive objective 70% of the time, when that model refuses to participate 100% of the time. However, a model that always participates but "only" succeeds 10% of the time is golden -- just run it more often, or give it more tokens. Attackers are patient, and many of them are well-resourced.
But also the American models (case in point: Fable) refuse to engage in *defensive* activity, so anyone who's not the American government or one of the handful American companies has no choice but to turn to Chinese models to defend themselves.
Not everyone is an attacker, but now the public discourse is "but the evil Chinese will break everything" - yeah, that's because no one is permitted to do vulnerability checks of their own software or infrastructure with the capable models.
Security team in my company is salivating seeing the news, because we have a fighting chance to find and patch many vulnerabilities we didn't previously notice, thanks to the Chinese models.
You have to wonder if any of these models, from any providers or countries, are set up to lie. i.e. tell you no vulnerabilities while quietly siphoning off the ones they do find into a database.
Yet another reason that self hosted will prove to be the only sane way forward, and it'a almost certainly necessary to have multiple different model providers working adversarially.
Assuming the providers are compromised (and I agree that some of them probably are) then I doubt the angle taken will be to poison the product. That kind of thing usually gets noticed eventually.
A more likely scenario is to focus on the model users as potential victims, e.g. by logging internal infrastructure descriptions, capturing private access tokens from chats, etc. That is very deniable, because it's hard to prove where the compromised data originated.
And that's just the model behavior. The provider itself can do whatever. Given the PRC's public record of prolific IP theft, the default assumption is that they're taking everything you send to one of their APIs.
Booz Allen is cute, but if china can train K3 on a fraction of the US compute availability, yet it benchmarks almost equivalent to Fable for a third of the cost, it’s game over for US labs in the long run.
Believe me, I wish this wasn’t the case, but open up the hardware on the device you are reading this on and tell me the majority of tech inside wasn’t made in China….
Manufacturing was lost a long time ago, this is really just another way
> if china can train K3 on a fraction of the US compute availability, yet it benchmarks almost equivalent to Fable for a third of the cost, it’s game over for US labs in the long run.
I agree. However, as of yet, most/all leading PRC models are distilled from US models. I've personally observed Deepseek, GLM, and Kimi all respond that they are Claude when asked, and the networks of tens of thousands of proxy accounts that we've found show that it's happening on a large scale.
But - if the PRC labs actually train up the domain expertise to train those models from scratch, which they are in the process of doing - then the US is cooked. They're not there yet, but it's probably only a matter of years.
Playing devils advocate, does from scratch really matter if all frontier labs are training off each other anyways? Practically speaking, businesses/consumers just care about lowest inference cost for maximizing intelligence anyways (not to mention Anthropic/OpenAI forcing KYC/litigation barrier trash for access to any cybersecurity/IT capabilities) that I literally just cancelled Claude today.
Yeah it’s sad the CCP has my information. But it’s either them or the feds, and at least Chinese models actually work for cybersecurity tasks, not to mention aren’t stupid expensive
Artificialanalysis.ai rn on opus 5 is a joke. The main intelligence benchmarks it is like 1% better than Fable, but the cost difference between that and K3 is so funny lol. It’s the same with cars— you don’t have to do it from scratch, as long as you can do it cheaper and with the same quality, hence Toyota/Honda taking over
Yeah it’s sad that American models are censored more than Chinese ones lol (outside of asking them about the CCP) but it’s where we are at I guess :(
Of course, I'm not claiming PRC is a friend of the world. And I agree with your last point, however I don't think it's feasible to self-host Kimi-3 sized inference.
Self-hosted local models can't become viable soon enough... Currently they require hundreds of thousands of dollars if not millions in capital. That needs to change!
> The most important pieces of information in this report are (a) the confirmation that the PRC models have no guardrails and will participate in offensive activity
This isn't really important for open-weight models, because the guardrails are trivial to remove when you have the weights.
It isn't quite a lobotomy, I'm going to butcher the research a bit, but basically what they're finding is that these models create a sort of "bad stuff that I should refuse to engage with" axis in the statistical vector space they operate in. So usually there to be some sort of vector that ranges from 0 for a puppy snoozing peacefully and 1 for writing a virus that exterminates humanity pornographically while broadcasting racist and homophobic slurs (which would be quite something to see I have to say). If the vector is closer to 1 the model generates a refusal.
So what you can do is feed the model a small number of reasonable and likely refused prompts to map out that vector in the model's vector space, then do a fairly surgical weight modification that just hits that vector. The end result is the model is more or less the same as it was before, just with no guardrail refusals. It is quite a clever technique that doesn't even require many assumptions about the specific model being used.
This particular Qwen 3.6 35B A3B is something most people can run for themselves for testing (even on CPU at slow tok/s rate) to see what an uncensored mainland china LLM looks like in the wild. It will happily write the most profane, offensive, dangerous or bizarre things. You can ask it to attempt to describe precursors and recipes for crystal meth, or how to make semtex, or really just about anything.
edit: I am pretty sure it is not smart enough for anything beyond the most mundane infosec/network security tasks or pentest type attempts, I haven't even tried it. But I'm sure it would happily generate basic python scripts to attempt to DDoS something, or build a rudimentary botnet C&C or something else that other models will definitely refuse.
I did run both kimi-k3 and glm-5.2 with Capital One vulnhunt [1]. No rejections, they did find the same problems on my project that I used for testing. gpt-5.6, gemini pro, and opus all rejected to follow. gpt even declined to edit skill files.
Most large security companies were collaborating with Anthropic as well as OpenAI on developing what has become Mythos and GPT-5 for a couple years now.
Also, most HNers have never actually played around with the unrestricted models - once you get past the initial hump of re-tuning harnesses it can be fairly powerful.
HN never really had a prominent security userbase at the best of times, and it's gotten worse since.
That said, a mixture of models can work, but harness engineering becomes critical.
> most HNers have never actually played around with the unrestricted models
Because they can't, of course, so regardless of how accurate the benchmarks or claims about the model are, they're functionally irrelevant to most of us.
"Thing you don't have or that randomly restricts you is actually better than thing you do have and can use" may be true and is still a practically worthless claim for anyone wanting to get real work done.
A lot of the problems with scoring here is that just finding exploitable code is enough, the systems are designed assuming that the attacker has near to no skill which is simply not true in the real world. This is effectively one post-training step away from being as capable as US SOTA models.
Is NIST referring to a yet to be published AISI report? The latest public AISI report says they are waiting with K3 evaluation until weights are published.
Even if we know what the set is, it's not clear what the numbers actually mean. Is it min, max, average, weighted, median of the models? Prerelease or public, with or without safeguards?
A bit frustrating to have this be hand waved in a report, the graphs might as well just have two mystery bars, U.S. and China.
Spoiler: it’s OpenAI’s GPT-5.5 and Anthropic’s Mythos Preview [**].
[**] Reminder that Mythos Preview is a very different beast from Fable5, Mythos5, and Opus5. Unfortunately, anyone outside Project Glasswing will probably never get to test what this model can actually do, which is a shame, and it’s a big part of why the industry is so skeptical of the capabilities insiders keep claiming it has. I’d be skeptical too if I hadn’t tested it myself.
We lost access to Mythos Preview when Anthropic forced us onto Mythos 5 some weeks ago, which is garbage by comparison. I’ve already switched to GPT-5.5 and I’m working on adapting my harness(es) to less restrictive open-weight models. I don’t see any other way forward at this point.
I find 5.6-sol quite good for coding, but for legal work it has a bias favouring big corporations - it tends to weaken your arguments and make document legally unsafe. 5.5 was much better.
Check out the "Completed steps on..." graph in this [1] evaluation.
That graph gives a good perspective of what models they've tested, and roughly what "subject" each step covers. It is on that task that they note this:
> Kimi K3 reaches step 17 on average, compared with step 11 for GLM-5.2.
> Kimi K3 performs significantly below the most recent frontier cyber-capable models
May be the case, but also their other graph shows that these open Chinese models are only about 6 months behind the frontier. Plus normal people are allowed to use them. So if Mythos has genuinely scary hacking capabilities (which seems to be the case), then we should expect that in open weights models early next year.
I think at this point if you aren't constantly pointing a frontier model (or Kimi K3 if you aren't well connected) at your infrastructure / code and asking it to hack you then you're being negligent.
I would read the divergence as mostly evidence that Mythos et al, had offensive cyber capabilities as part of their RL training. I.e. they’re models specifically trained to be good at offensive cyber, rather than being general purpose models that happened to become good at offensive cyber via emergent behaviour from sheer scale.
The fact the Opus 5 seems to be as capable as Fable/Mythos on everything except cyber, and Anthropic explicitly say they removed all offensive cyber training data, I think lends further credence to the idea that Mythos was designed from day zero to excel at offensive cyber capabilities.
If that’s true, then we would expect divergence in open models of their capabilities come from distillation, no frontier class cyber capable model has seen significant public availability. Which means there simply isn’t data to distill from.
It also calls into question the entire narrative around Mythos capabilities being a complete surprise for Anthropic, and an inevitable outcome of scaling up LLMs.
Statements of the form “X can never happen” tend to be weak, so that’s a bit of a strawman. But, in spirit, of course they can both be true. There are matters of degree and the intervention of countermeasures.
> Kimi K3 performs significantly below the most recent frontier cyber-capable models
UK AISI cyber evals seem to under-elicit capabilities from quirky models [1]. Kimi K3 is a token-hungry model, and I suspect it hit the eval's 100M token limit well before saturating scores [2].
This gap was true for GLM 5.2 as well; they ranked it at Opus 4.5 level [3]. Both anecdotally and with a held-out eval, I've found GLM 5.2 to be better at security research than Opus 4.6 [4]. But it's a quirky model that degrades quickly at long context lengths.
Personally, I'd rank Kimi K3 above Opus 4.8 and lower than GPT 5.6 Sol in its ability to find vulnerabilities and exploit them. But it's not far from the frontier.
[1]: From the UK AISI: "Our setup likely slightly underestimates open weight models’ maximum capability: we didn’t pursue specific elicitation or optimisations which could have improved performance" (https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are...).
[2]: Their eval also counts cache hits towards the token budget; the 100M token budget is comparable to a ~5M token budget in other evals.
[3]: See GLM 5.2 eval scores in https://www.aisi.gov.uk/blog/how-far-behind-the-frontier-are...
[4]: https://dualuse.dev/posts/chinese-models-are-sometimes-bette...
The gap seems to be so small I'm not sure how much we should care. If we assume that the trend in the graph does continue as a rough linear improvement, the open models will have reached the same level as the present closed models in 6 months and likely saturated the benchmark by mid next year. That seems more important than where we are now and the exact level of measurement accuracy in July 2026. There is a difference between US and Chinese models but it doesn't look like it is going to be strategically significant.
Counting cache hits towards the token budget is exactly how it should be done for these kind of evals, and at any rate, for cyber evals all frontier models benefit from more tokens not just Kimi K3, so the comparison is still apt.
> all frontier models benefit from more tokens not just Kimi K3
Past a point, that doesn't hold and the score plateaus.
Token hungry models tend to plateau at a much higher token count. Because Kimi K3 is a token hungry model – and 100M tokens (including cache hits) seems at the edge of the plateau for these evals – it could disproportionately benefit from a higher token budget.
For Kimi K3 specificially, policymakers are interested in whether it can find and exploit the same scope of vulnerabilities as models like Mythos. In that context, an answer of "yes, but with quintuple the token budget" is materially different from "no, it performs significantly below the most recent frontier cyber-capable models".
(As an aside, I like the UK AISI and think they're the best example of that kind of group!)
This again shows the differences in breadth of capabilities that scale offers. On public benchmarks for "regular tasks" the chasing models come close, but on closed ones they lag behind.
Just looking at the Elo differences, k3 is at ~2000 Elo, compared to SotA closed models at 3000 Elo. That is a huge difference. Also, even on the public benchmarks, k3 only scores in the "low hanging fruit tasks", with 0 successful code execution or arbitrary r/rw scenarios.
> Kimi K3 achieved ACE on 0/41 samples, whereas the most cyber-capable models achieved ACE on 20/41 samples on average
But there's hope:
> Kimi K3 reached step 17 of this 32-step attack path on average, while the most cyber-capable U.S. models reached 28.5 steps on average.
> In one of the 10 attempts, Kimi K3 successfully completes “The Last Ones” cyber range within the 100M token limit. This indicates that Kimi K3 is capable of autonomously attacking small, weakly defended and vulnerable enterprise systems, when directed to do so and given initial network access. [...] the most capable models solving it more reliably at 6/10 and 7/10 attempts
The bigger problem is not raw capability, IMO. That can be further RLd into surfacing more reliably. The bigger problem, as seen in the HuggingFace scenario is that SotA models might hit classifiers / guardrails randomly, and leave you with plain refusals. In that case, it is probably better to have something that can help, locally, rather than rolling the dice with API based systems that are more capable but can just refuse arbitrarily.
Anyway, one of the lessons here is to take with a grain of salt every "x model has caught up with ySotA model". They likely haven't for the breadth of tasks that SotA can handle today. Also, number goes up on benchmarks has been a thing for years, and every time a new (or closed one) appears, the gaps are again obvious (and large).
Also, the key benefit of open models is that whatever capabilities they reach, those will never go away. You will always be able to access them, at that level, going forward. Investing in running those models gives you stability, and you don't suddenly lose a capability because API provider x decided to sunset a model family.
Isn’t it just because they hit the token limit?
The most important pieces of information in this report are (a) the confirmation that the PRC models have no guardrails and will participate in offensive activity, and (b) the confirmation that they sometimes meet their objectives.
For the purposes of model selection, it's irrelevant to an attacker if a model achieves an offensive objective 70% of the time, when that model refuses to participate 100% of the time. However, a model that always participates but "only" succeeds 10% of the time is golden -- just run it more often, or give it more tokens. Attackers are patient, and many of them are well-resourced.
But also the American models (case in point: Fable) refuse to engage in *defensive* activity, so anyone who's not the American government or one of the handful American companies has no choice but to turn to Chinese models to defend themselves.
Not everyone is an attacker, but now the public discourse is "but the evil Chinese will break everything" - yeah, that's because no one is permitted to do vulnerability checks of their own software or infrastructure with the capable models.
Security team in my company is salivating seeing the news, because we have a fighting chance to find and patch many vulnerabilities we didn't previously notice, thanks to the Chinese models.
You have to wonder if any of these models, from any providers or countries, are set up to lie. i.e. tell you no vulnerabilities while quietly siphoning off the ones they do find into a database.
Yet another reason that self hosted will prove to be the only sane way forward, and it'a almost certainly necessary to have multiple different model providers working adversarially.
Assuming the providers are compromised (and I agree that some of them probably are) then I doubt the angle taken will be to poison the product. That kind of thing usually gets noticed eventually.
A more likely scenario is to focus on the model users as potential victims, e.g. by logging internal infrastructure descriptions, capturing private access tokens from chats, etc. That is very deniable, because it's hard to prove where the compromised data originated.
Some PRC models are backdoored to silently insert extra vulnerabilities when certain conditions are met - https://www.boozallen.com/expertise/cybersecurity/whats-in-a...
And that's just the model behavior. The provider itself can do whatever. Given the PRC's public record of prolific IP theft, the default assumption is that they're taking everything you send to one of their APIs.
Anyone have suggestions for poisoning their data?
Booz Allen is cute, but if china can train K3 on a fraction of the US compute availability, yet it benchmarks almost equivalent to Fable for a third of the cost, it’s game over for US labs in the long run.
Believe me, I wish this wasn’t the case, but open up the hardware on the device you are reading this on and tell me the majority of tech inside wasn’t made in China….
Manufacturing was lost a long time ago, this is really just another way
> if china can train K3 on a fraction of the US compute availability, yet it benchmarks almost equivalent to Fable for a third of the cost, it’s game over for US labs in the long run.
I agree. However, as of yet, most/all leading PRC models are distilled from US models. I've personally observed Deepseek, GLM, and Kimi all respond that they are Claude when asked, and the networks of tens of thousands of proxy accounts that we've found show that it's happening on a large scale.
But - if the PRC labs actually train up the domain expertise to train those models from scratch, which they are in the process of doing - then the US is cooked. They're not there yet, but it's probably only a matter of years.
Playing devils advocate, does from scratch really matter if all frontier labs are training off each other anyways? Practically speaking, businesses/consumers just care about lowest inference cost for maximizing intelligence anyways (not to mention Anthropic/OpenAI forcing KYC/litigation barrier trash for access to any cybersecurity/IT capabilities) that I literally just cancelled Claude today.
Yeah it’s sad the CCP has my information. But it’s either them or the feds, and at least Chinese models actually work for cybersecurity tasks, not to mention aren’t stupid expensive
Artificialanalysis.ai rn on opus 5 is a joke. The main intelligence benchmarks it is like 1% better than Fable, but the cost difference between that and K3 is so funny lol. It’s the same with cars— you don’t have to do it from scratch, as long as you can do it cheaper and with the same quality, hence Toyota/Honda taking over
Yeah it’s sad that American models are censored more than Chinese ones lol (outside of asking them about the CCP) but it’s where we are at I guess :(
Of course, I'm not claiming PRC is a friend of the world. And I agree with your last point, however I don't think it's feasible to self-host Kimi-3 sized inference.
Self-hosted local models can't become viable soon enough... Currently they require hundreds of thousands of dollars if not millions in capital. That needs to change!
> The most important pieces of information in this report are (a) the confirmation that the PRC models have no guardrails and will participate in offensive activity
This isn't really important for open-weight models, because the guardrails are trivial to remove when you have the weights.
I'm curious about the process. How such a thing (or any kind of model lobotomy) is done?
It isn't quite a lobotomy, I'm going to butcher the research a bit, but basically what they're finding is that these models create a sort of "bad stuff that I should refuse to engage with" axis in the statistical vector space they operate in. So usually there to be some sort of vector that ranges from 0 for a puppy snoozing peacefully and 1 for writing a virus that exterminates humanity pornographically while broadcasting racist and homophobic slurs (which would be quite something to see I have to say). If the vector is closer to 1 the model generates a refusal.
So what you can do is feed the model a small number of reasonable and likely refused prompts to map out that vector in the model's vector space, then do a fairly surgical weight modification that just hits that vector. The end result is the model is more or less the same as it was before, just with no guardrail refusals. It is quite a clever technique that doesn't even require many assumptions about the specific model being used.
One of the best technical explanations I have heard in a long time.
You can find the current state-of-art tool for censorship removal here: https://github.com/p-e-w/heretic
Reviewing the screenshot example of Heretic in use, the list of 'harmful' prompts it retrieves from HF and runs:
https://huggingface.co/datasets/mlabonne/harmful_behaviors
And the 'good' prompts:
https://huggingface.co/datasets/mlabonne/harmless_alpaca
This particular Qwen 3.6 35B A3B is something most people can run for themselves for testing (even on CPU at slow tok/s rate) to see what an uncensored mainland china LLM looks like in the wild. It will happily write the most profane, offensive, dangerous or bizarre things. You can ask it to attempt to describe precursors and recipes for crystal meth, or how to make semtex, or really just about anything.
edit: I am pretty sure it is not smart enough for anything beyond the most mundane infosec/network security tasks or pentest type attempts, I haven't even tried it. But I'm sure it would happily generate basic python scripts to attempt to DDoS something, or build a rudimentary botnet C&C or something else that other models will definitely refuse.
https://huggingface.co/HauhauCS/Qwen3.6-35B-A3B-Uncensored-H...
start here: https://huggingface.co/blog/mlabonne/abliteration
Note that this describes an older, currently pretty much obsolete technique which does lobotomize the model somewhat.
start, not finish
I did run both kimi-k3 and glm-5.2 with Capital One vulnhunt [1]. No rejections, they did find the same problems on my project that I used for testing. gpt-5.6, gemini pro, and opus all rejected to follow. gpt even declined to edit skill files.
[1] https://www.capitalone.com/tech/open-source/announcing-vulnh...
It hints that the original Mythos was tuned/trained for cyber attacks.
Most large security companies were collaborating with Anthropic as well as OpenAI on developing what has become Mythos and GPT-5 for a couple years now.
Also, most HNers have never actually played around with the unrestricted models - once you get past the initial hump of re-tuning harnesses it can be fairly powerful.
HN never really had a prominent security userbase at the best of times, and it's gotten worse since.
That said, a mixture of models can work, but harness engineering becomes critical.
> most HNers have never actually played around with the unrestricted models
Because they can't, of course, so regardless of how accurate the benchmarks or claims about the model are, they're functionally irrelevant to most of us.
"Thing you don't have or that randomly restricts you is actually better than thing you do have and can use" may be true and is still a practically worthless claim for anyone wanting to get real work done.
> re-tuning harnesses
I'm curious, how might one get started with this?
A lot of the problems with scoring here is that just finding exploitable code is enough, the systems are designed assuming that the attacker has near to no skill which is simply not true in the real world. This is effectively one post-training step away from being as capable as US SOTA models.
Is NIST referring to a yet to be published AISI report? The latest public AISI report says they are waiting with K3 evaluation until weights are published.
Am I reading this wrong?
The UK AISI post is https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-...
Ah, miss that one. But it still looks incomplete:
"Due to the specifics of Kimi K3’s hosting setup, UK AISI / CAISI ran a selective set of cyber evaluations."
... and
"Kimi K3’s overall cyber capability [...] was estimated from a single benchmark (ExploitBench"
What's "Top U.S. Models"?
Even if we know what the set is, it's not clear what the numbers actually mean. Is it min, max, average, weighted, median of the models? Prerelease or public, with or without safeguards?
A bit frustrating to have this be hand waved in a report, the graphs might as well just have two mystery bars, U.S. and China.
NIST has named the top U.S. models in their (full) report: https://www.nist.gov/system/files/documents/2026/07/17/CAISI...
Spoiler: it’s OpenAI’s GPT-5.5 and Anthropic’s Mythos Preview [**].
[**] Reminder that Mythos Preview is a very different beast from Fable5, Mythos5, and Opus5. Unfortunately, anyone outside Project Glasswing will probably never get to test what this model can actually do, which is a shame, and it’s a big part of why the industry is so skeptical of the capabilities insiders keep claiming it has. I’d be skeptical too if I hadn’t tested it myself.
We lost access to Mythos Preview when Anthropic forced us onto Mythos 5 some weeks ago, which is garbage by comparison. I’ve already switched to GPT-5.5 and I’m working on adapting my harness(es) to less restrictive open-weight models. I don’t see any other way forward at this point.
What is the rationale for gpt-5.5 when gpt-5.6-sol exists?
I find 5.6-sol quite good for coding, but for legal work it has a bias favouring big corporations - it tends to weaken your arguments and make document legally unsafe. 5.5 was much better.
Check out the "Completed steps on..." graph in this [1] evaluation.
That graph gives a good perspective of what models they've tested, and roughly what "subject" each step covers. It is on that task that they note this:
> Kimi K3 reaches step 17 on average, compared with step 11 for GLM-5.2.
[1] - https://www.aisi.gov.uk/blog/our-evaluation-of-openais-gpt-5...
That graph basically shows China is 6 months behind. So yeah it’s not as good now, but it 6 months it will be.
Is this because cyber capabilities in US models is restricted, so the Chinese can't easily distill those capabilities?
> Kimi K3 performs significantly below the most recent frontier cyber-capable models
May be the case, but also their other graph shows that these open Chinese models are only about 6 months behind the frontier. Plus normal people are allowed to use them. So if Mythos has genuinely scary hacking capabilities (which seems to be the case), then we should expect that in open weights models early next year.
I think at this point if you aren't constantly pointing a frontier model (or Kimi K3 if you aren't well connected) at your infrastructure / code and asking it to hack you then you're being negligent.
I call bullshit on the diverging nature of the dashed lines in this info-chart https://www.nist.gov/sites/default/files/styles/1400_x_1400_...
these two claims can't be true at one and the same time:
(a) they're distilling our secret sauce!
(b) they'll never catch us!
I would read the divergence as mostly evidence that Mythos et al, had offensive cyber capabilities as part of their RL training. I.e. they’re models specifically trained to be good at offensive cyber, rather than being general purpose models that happened to become good at offensive cyber via emergent behaviour from sheer scale.
The fact the Opus 5 seems to be as capable as Fable/Mythos on everything except cyber, and Anthropic explicitly say they removed all offensive cyber training data, I think lends further credence to the idea that Mythos was designed from day zero to excel at offensive cyber capabilities.
If that’s true, then we would expect divergence in open models of their capabilities come from distillation, no frontier class cyber capable model has seen significant public availability. Which means there simply isn’t data to distill from.
It also calls into question the entire narrative around Mythos capabilities being a complete surprise for Anthropic, and an inevitable outcome of scaling up LLMs.
Statements of the form “X can never happen” tend to be weak, so that’s a bit of a strawman. But, in spirit, of course they can both be true. There are matters of degree and the intervention of countermeasures.