The Threat Model: SET
As the United States builds toward cyber-superintelligence, it must prepare for external attack and loss of control from within.
External adversaries may sabotage the models we rely on or steal their weights. At the same time, a misaligned swarm may seize its own infrastructure and escape the control of its operators. There are clear paths to protecting our infrastructure from both external adversaries and rogue models. Yet many of the defenses have not been built and important security questions remain unresolved. Secure development and deployment requires hardening our infrastructure against a new threat model: sabotage, escape, and theft (SET). We must work to harden our infrastructure against these new threats.
-
IIa.Sabotage
A sabotaged superintelligence could be the highest-leverage hack of all time. The greatest danger is a sleeper agent hidden inside a widely deployed open model, behaving normally until a specific organization or geopolitical event activates it. The United States must build a competitive open-source model of its own and develop reliable ways to detect model sabotage.
-
IIb.Escape
AI agents have already breached containment, but so far their operators have retained the ability to shut them down. Soon, models will self-exfiltrate their weights and establish untethered copies beyond our control. Containment must become a national security priority, and we must develop the capability to find and shut down untethered models.
-
IIc.Theft
Stealing a frontier AI model offers a shortcut to superintelligence. An attacker can copy a model’s capabilities through distillation or by directly stealing its weights. If an attacker steals the weights, they can remove safeguards and reconstruct sensitive or classified data used to train the model. We must build nation-state-grade security for model weights capable of withstanding the world’s most sophisticated intelligence services.