NSA’s deputy director, Tim Kosiba, said it plainly at an INSA panel in Bethesda on Thursday: “We want access to all the models.” He confirmed active talks with the frontier labs. He would not name them, would not say which models NSA uses today, and would not say whether any developer has handed over an unreleased system.
Here is the machinery behind that line.
A June executive order set up a voluntary program. A developer can give the government access to a qualifying model for up to 30 days before releasing it to other trusted partners. The NSA’s director decides which systems cross the covered frontier model threshold, with input from ONCD and CISA. The order specifically bars a mandatory licensing or approval regime.
A companion national security memo told defense and intelligence agencies to build proactive partnerships with industry and to avoid depending on any single vendor.
Two honest ways to read this.
| Position A: Test it first | Position B: Voluntary in name only | |
|---|---|---|
| Core claim | A model that can find and exploit unknown software flaws is closer to a weapon than a product. Someone has to test it against classified networks before it ships. | The agency that evaluates your model is also your biggest customer and can label you a supply chain risk. That is leverage, not choice. |
| What it points to | Anthropic withheld its first Mythos model over exactly this risk. NSA is now reportedly using Mythos versions to probe defenses on US military networks. | DoD designated Anthropic a supply chain risk after it declined to loosen some military-use limits. Parts of NSA then lost Mythos 5 access during an export control fight. A judge has temporarily blocked parts of the government’s action. |
| What it protects | Defenders get the capability before attackers do. | Independent safety judgment at the developer, and a supplier base with more than one name in it. |
| Weakest point | Thirty days of lead time means little against an adversary who steals the weights or replicates the research. | Declining to test frontier cyber capability does not make the capability disappear. |
The detail that cuts against both positions: the White House told developers this month that open-weight models are excluded from the testing program. Those are the models anyone can download and modify. So the program examines the systems that are already under corporate control and skips the ones that are not.
The part of this that lands on our desks
Every element of this needs cleared people, not just cloud credits.
Classified evaluation of frontier models is not entry-level data science. It needs cleared ML engineers, cleared offensive security testers, and people who can write and defend a test plan at TS/SCI. Add the compliance layer Kosiba flagged, where a human still has to verify every use against surveillance law and internal rules. That is another cleared body per workflow, not fewer.
That pool is small and slow to build. A 30-day model access window does not create a single cleared evaluator.
Three questions for the room
- If a lab declines the pre-release window, does it stay a customer? Answer that one honestly and the “voluntary” question answers itself.
- Should the window cover open-weight releases too, or is that simply unenforceable?
- Who actually staffs classified model evaluation: government civilians, FFRDCs, or cleared industry? And at what labor rate?
My read: the design is sound and the practice is fragile. Voluntary only holds if refusal carries no procurement consequence. Right now nobody in this market can say that with a straight face.
