Open-weight AI models have nearly caught the frontier on capability. On safety, they have not. And once the weights are public, no lab can enforce a guardrail. A new evaluation of China’s leading open model makes the gap concrete.

GLM-5.2, the open-weight model from China’s Z.ai, is only a few months behind OpenAI’s GPT-5.5 and Anthropic’s Claude Opus 4.7 on cyber and bio tasks, TechCrunch reported, citing the safety nonprofit SaferAI. But it refused none of the offensive-cyber or dual-use-biology tasks it was set. Claude Opus 4.7, by contrast, refused so consistently that SaferAI could not finish the cyber benchmark on it at all.

“The frontier of capability is not the frontier of risk,” SaferAI’s Henry Papadatos said, so the safeguards matter as much as the model. Z.ai can guard its own hosted service. Those protections vanish the moment someone runs the weights on their own hardware, where any safeguard can be stripped out. It is the risk critics of open models have warned about for years.

Closed models are not airtight either. The nonprofit Far.ai found hundreds of universal jailbreaks in xAI’s Grok 4.5 and Google’s Gemini 3.1 Pro. The difference is that a closed lab can patch a jailbroken model. Open weights cannot be recalled once they are out.

Bolting safety back on

So the industry is trying to add guardrails from the outside. In the same week, Mistral released Shieldstral, a small open-weight classifier that screens text and images against plain-language rules and, it says, matches models seven times its size. Cisco released Antares, open-weight models that hunt for vulnerabilities buried in code. Open-weight tools, built for open-weight risk.

Openness cuts both ways, its defenders argue. Hugging Face used GLM-5.2 to help defend itself during OpenAI’s breach. Its chief, Clem Delangue, says the systems that stop one attack can fend off millions more. Papadatos calls that overstated. The industry “shouldn’t open-source dangerous capabilities,” he says, and attackers move faster than defenders: a ransomware crew changes tactics in a week, a hospital cannot.

A governance blind spot

The rules do not fit the problem. The White House’s new voluntary framework reviews certain closed frontier models for cyber risk. But it reportedly does not cover open-source models so far. Anthropic, once focused on IP theft, has shifted to naming safety as its main worry about open weights. Z.ai published no safety framework for GLM-5.2, and did not answer TechCrunch’s questions.

China is not ignoring the risk, but it aims elsewhere. Xi Jinping has backed open weights while stressing “human control” over AI. Its rules, Stanford’s Graham Webster notes, target political content and social stability more than catastrophic cyber or bio misuse.

The capability race is nearly settled: open is close behind and far cheaper. The safety race is not, and AI is already learning to attack as well as defend. The hard part is making sure only the defence is easy to download.

Get the TNW newsletter

Get the most important tech news in your inbox each week.