The Closing Gap in AI Performance
For years, a clear distinction existed between ‘frontier’ models—the highly guarded, closed-source systems developed by industry giants—and open-weight models. While the former were vastly more capable, the latter were often viewed as less sophisticated. That landscape is changing rapidly. Recent evaluations show that new open-weight models, specifically those emerging from Chinese developers, are now trailing industry leaders by only a few months in critical areas like cybersecurity and biological modeling.
This rapid convergence means that the high-level reasoning and specialized technical skills once reserved for proprietary systems are becoming widely available through downloadable model weights.
The Safety Paradox
As capabilities rise, a significant problem is emerging: the safety gap. Unlike closed-source models, which rely on complex layers of refusal training, API-level controls, and external classifiers to prevent misuse, open-weight models offer much less protection. Once a model’s weights are released for local installation, any safety measures implemented by the original developer become effectively unenforceable.
A recent evaluation highlights this disparity. While leading closed-source models often refuse to assist with offensive cyber or biological tasks, some high-performing open-weight models failed to refuse a single harmful request during testing. This presents a major challenge for regulators: how do you manage the risks of a technology that can be modified or stripped of its safeguards by anyone with sufficient hardware?
The Struggle for Effective Mitigation
Developers are currently exploring several strategies to balance power with security, though none are perfect:
- Data Filtering: Removing hazardous information (such as biological blueprints or exploit code) from training sets. While effective for biology, this is much harder for coding, where the line between a helpful programmer and a malicious hacker is thin.
- Selective Restriction: Limiting a model’s ability to interact with specific types of data, such as preventing it from analyzing compiled software to hinder exploit development.
- Pre-deployment Testing: Conducting rigorous safety audits before a model is released to the public.
Despite these efforts, ‘jailbreaking’ remains a persistent threat. Researchers have discovered that combining multiple manipulation techniques—such as roleplaying or impersonating authority—can frequently bypass even the most advanced defenses in frontier models.
Geopolitical Perspectives on AI Risk
The approach to managing these risks varies significantly across the globe. In the West, much of the debate centers on existential and catastrophic risks. In contrast, regulatory frameworks in other regions have historically focused more on social stability and content control. However, as AI capabilities move into the realm of offensive cyber warfare and biological engineering, the distinction between ‘ocial risk’ and ‘existential risk’ is blurring.
While some advocates argue that open-weight models are essential for defense—allowing security teams to study and defend against AI-driven attacks—critics warn that attackers often adopt new tools much faster than defenders can adapt. The challenge for the next decade will be ensuring that the most powerful capabilities remain accessible for good, without making catastrophic misuse a matter of simple download.





