The Cyber Risk Debate Needs More Nuance
Discussions about the cybersecurity risks of open models often assume one of two things: that banning open models would stop bad actors, or that China does not care about AI saf...
By AI Engineering Team
Discussions about the cybersecurity risks of open models often assume one of two things: that banning open models would stop bad actors, or that China does not care about AI safety. Either assumption could lead to policies that reduce American AI competitiveness while increasing long-term cyber risk. A better approach requires acknowledging uncertainty, trade-offs, and uncomfortable realities.
Several positions are especially visible in the Western AI ecosystem:
- Open-weight models pose an unacceptable risk to society, a view associated with some frontier-lab leaders and members of the U.S. national security community.
- Open-weight models are necessary for defense, and banning them could make the world less safe, a position held by AI-risk moderates to varying degrees. This includes the author, Hugging Face’s perspective after the OpenAI incident, and commentators such as Joshua Saxe.
- Chinese companies continue releasing open-weight models with strong cyber capabilities, based on risk assessments shaped by their society and government.
These positions are not equally prominent. Risk-focused arguments have received disproportionate attention, including some work that makes cross-party engagement more difficult. There has also been little substantive analysis of how Chinese companies assess risk. Instead, discussion often relies on claims that Chinese labs simply do not care about safety.
The assumptions behind opposition to open weights
Anthropic’s report on the risks of GLM-5.3 as an offensive cyber tool is a recent example of this debate. The technical research is largely reasonable, but the report does not address broader questions. What would happen if open models were banned because of cyber risks? Why do Chinese companies consider these models safe enough to release?
Those questions are necessary for understanding the wider AI ecosystem. Policy interventions change the shape of risk, but they rarely produce universally better outcomes. Every intervention involves trade-offs.
This narrow focus also resembles arguments based on classified briefings. Some people who support open models publicly say that information they have seen about a Chinese open-weight model shows an overwhelming threat that requires immediate action.
That view conflicts with the public evidence available so far. Documented cyber attacks have primarily involved closed models. The observed attacks involving OpenAI models, along with the broader evidence collected by FelonyBench, are among the clearest available indicators of how AI-related cyber risks are currently developing.
Several explanations remain possible, and determining which is correct could take years:
- Open model weights and closed-model APIs with safeguards may both be relatively easy to misuse. Instead of "Open Dangerous, Closed Safe," the more accurate description may be "Open Unsafe, Closed Unsafe."
- There may be fewer bad actors willing to use AI for highly visible cyber attacks against critical U.S. infrastructure or powerful but unprepared institutions. Closed models may offer stronger capabilities and, in theory, stronger safeguards, but the capability advantage could produce more net harm if safeguards in both systems are porous.
- Open-weight models may become an important tool for cyber defense. This remains an area requiring more analysis, but it is reasonable to argue that open models could help prevent harm over the next several years. Sensitive government agencies, for example, may need models that can run on air-gapped networks, where open weights are currently the only practical option.
The reality may be difficult to measure even when the logical trade-offs are clear. Closed models could ultimately be safer because they can be surrounded by more layers of protection. In the near term, however, widely accessible model APIs could cause more harm than expected. Control and deployment practices matter as much as the tools themselves, and concentrating such an important capability among a small number of companies could create its own risk.
This leads to a broader policy implication. If the newest open-weight models should be banned to slow the spread of cyber capabilities, then public-facing APIs for frontier closed models may also need to be restricted. Closed models currently have stronger safeguards than open-weight models, but those safeguards are not perfect. Their cyber capabilities may improve faster than their guardrails, creating a situation in which banning open models while allowing closed models to advance widens the offensive-defense gap.
Building the mitigations needed for defenders to run powerful models on private infrastructure would also take significant time if access to strong open-weight models were removed.
Understanding China’s AI risk posture
China does care about AI safety, but it approaches the issue through its own culture, institutions, and power structures. Chinese society is generally techno-optimistic, while the government places strong emphasis on political stability. Both characteristics affect how emerging risks such as cybersecurity are evaluated.
China has its own AI risk framework, although it has not received the same public attention as the political relationship between the White House and U.S. frontier labs. Chinese companies are understood to register major model releases with the government. These registrations include evaluations that initially focused on controlling access to information prohibited by the government. It is unclear how extensively this framework now covers risks such as cyber operations or biological threats. China’s government is decentralized, and information from AI labs must pass through existing institutions before reaching senior leadership.
At the same time, China appears to have a stronger emphasis on social risk. Leading AI researchers may face restrictions on leaving the country, and the industry is becoming less open to foreign investment. The personal consequences of causing serious domestic harm may therefore be greater than in the United States, where the most severe outcome might be the collapse of a company.
Political oversight also appears to occur earlier in the development process. Western AI companies often operate according to a principle close to: the best way to stop a bad actor with AI is to ensure that good actors have AI first. Chinese companies tend to be more pragmatic about building useful tools, often integrating them with existing businesses.
Chinese AI labs also appear to consider the international consequences of their work. At the same time, they are strongly incentivized, and sometimes publicly encouraged by the government, to compete. Comprehensive safety evaluations for a frontier model such as Kimi K3 could cost tens of millions of dollars in compute. Companies may prefer to direct that compute toward training instead.
The more useful policy question is therefore how much compute a lab should be required to spend on safety testing before releasing a model. There is a reasonable case that Chinese labs should spend more or provide greater transparency. It is less clear that the required minimum should match the spending of Anthropic or OpenAI.
Neither company’s institutional commitment to understanding risk necessarily places safety above business value and economic success. Incidents involving Hugging Face and OpenAI have raised concerns about whether frontier labs monitor their models closely enough in a highly competitive environment. The Hacktron hacking of OpenAI provides another example: a company seeking to act as a trusted cybersecurity partner can face difficulties if its own systems experience extensive leaks.
Individual researchers at these labs may work diligently to improve safety, but institutional reliability is a separate question.
Would Claude Mythos have been acceptable as an open-weight model?
When Claude Mythos was announced, it was presented as a potential new class of cyber weapon. The implication was that an accidental release into the wrong hands could cause widespread social destabilization.
That prediction now appears questionable.
By the available measures, GLM-5.3 is the model that crosses the relevant capability threshold. There is little public evidence that the situation has changed substantially, even more than a month after the model weights were released.
If Claude Mythos had been accidentally released as an open-weight model, the consequences might have been serious but manageable. Cybersecurity incidents would likely have increased, and the leak would have been harmful to society. However, the result might have looked more like an acceleration of existing trends than a fundamental change in risk.
Proponents of the latest wave of concern about open weights have made a testable prediction: current open-weight models will cause unprecedented harm by severely disrupting cyber infrastructure.
Based on the evidence available so far, that prediction appears unlikely to hold. After years of debate about the risks of open weights, the continuing release of capable models may finally provide enough evidence to evaluate these claims more directly.