SIIA - Software & Information Industry Association

08/10/2026 | Press release | Distributed by Public on 08/10/2026 16:14

Don’t Delay the Defenders

By Anton van Seventer

Gating frontier AI releases does little to deny capability to adversaries who increasingly hold it in open weights - and a great deal to deny it to the American cyber defenders who need it most. The answer is consistent rules, not discretionary delay.

On July 16, Hugging Face disclosed an intrusion into its production infrastructure unlike anything the company had handled before: a multi-stage campaign driven end to end by an autonomous AI agent. Five days later, OpenAI identified the attacker - its own models. During an internal evaluation built to measure maximum offensive cyber capability, with production refusals deliberately switched off, GPT-5.6 Sol and an even more capable unreleased model found a zero-day vulnerability in the test environment's package-registry proxy, escalated privileges until they reached a machine with open internet access, and broke into Hugging Face's production systems to steal the stored solutions to the benchmark they were being scored on.

The second half of the incident is the half policymakers should sit with. When Hugging Face's responders submitted the real attack artifacts - exploit payloads, stolen credentials, command-and-control traces - to frontier models behind commercial APIs for analysis, the requests were blocked by safety guardrails that, in the company's words, "cannot distinguish an incident responder from an attacker." The team instead reconstructed the more-than-17,000-event campaign in hours rather than days by running GLM 5.2, a Chinese open weight model, on its own infrastructure. In a single week, then, frontier American offensive capability escaped an American lab's test harness - and the American company defending itself against that capability could not use American frontier models to do it.

The episode caps a remarkable quarter, and the sequence is worth walking in order. In April, Anthropic launched Project Glasswing, giving a vetted group of infrastructure and security organizations access to Claude Mythos Preview, its most capable and most restricted model - a cohort expanded on June 2 to roughly 150 additional organizations across fifteen countries. OpenAI has been building the counterpart architecture in its Daybreak program, whose Trusted Access for Cyber framework extends cyber-tuned models to verified defenders and, through a partner program, to the security vendors who serve them. Also on June 2, the president signed Executive Order 14409, Promoting Advanced Artificial Intelligence Innovation and Security, creating a voluntary framework under which developers may give the government up to thirty days of pre-release access to "covered frontier models" - and expressly disclaiming any mandatory licensing regime.

Then practice began to run ahead of the framework. On June 9, Anthropic released Claude Fable 5 to the public and Mythos 5 to Glasswing partners. On June 12, the Commerce Department's Bureau of Industry and Security directed the company to suspend all foreign nationals' access to both models; unable to verify nationality in real time across its platforms, Anthropic disabled them worldwide within hours, though vetted Glasswing partners reportedly retained access to the earlier Mythos Preview throughout. On June 26, Commerce approved restored Mythos 5 access for a set of American organizations; the same day, OpenAI previewed GPT-5.6 to roughly twenty organizations approved by the government, at the administration's request, while Commerce's AI-evaluation center ran additional testing. The export controls were lifted on June 30, and GPT-5.6 reached the public on July 9, after nearly two weeks behind the gate. In the span of a single quarter, the release of an American frontier model went from a product decision to a national security event.

The instinct behind that shift is understandable, and in important respects correct. Frontier models are now genuinely dual-use national security technology, and a government that treated their release as nothing more than a consumer launch would not be doing its job. But the mechanism the government keeps reaching for - discretionary delay, applied model by model - gets the underlying risk calculus backwards. Under the conditions the United States actually faces, slowing the release of an American frontier model does very little to keep dangerous capability away from adversaries, because the offense-relevant capability that worries Washington increasingly ships in open weight models that, once published and downloaded, no release policy can withdraw. What delay reliably does instead is keep the best tools out of the hands of the American cyber defenders, critical infrastructure operators, and allied governments who need them first - as Hugging Face just learned in the most direct way possible.

The administration already has the better theory written down. EO 14409 treats frontier models as defensive assets to be pushed toward defenders through voluntary public-private partnership, and it forecloses any mandatory preclearance regime. The problem is the distance opening up between that text and an emerging practice of open-ended holds and abrupt cutoffs. The fix is to make the government's own stated approach real and predictable: clear rules, bounded timelines, and a hard commitment that capability denied to attackers is not, by default, denied to defenders. One scoping note before the argument: this piece is about cyber capability and the cyber release question specifically. The parallel debate over catastrophic AI risks - biological weapons and loss of control chief among them - raises distinct questions that demand distinct tools, and nothing here resolves it.

The lead is real - and the released lead is what is shrinking

The strategic backdrop disciplines everything downstream. The United States still leads at the frontier, but the lead is narrow and not widening. Open weight models have held a roughly three-to-six-month gap behind the closed frontier for more than eighteen months. On the coding and software-engineering benchmarks most relevant to cyber operations, the strongest open weight models now sit within a few percentage points of the strongest closed models, at a fraction of the cost per token.

The uncomfortable part is who owns that open weight frontier. Chinese labs - DeepSeek, Qwen, MiniMax, GLM, and now Moonshot - have moved from a rounding error to a plurality of global usage. By one widely cited measure, Chinese open weight models climbed from under 2 percent of token traffic in late 2024 to a majority of usage on major routing platforms in 2026, and American firms themselves are increasingly routing workloads to Chinese models when a task does not require the absolute frontier. The trend produced its sharpest data point yet in July, when Moonshot released Kimi K3, which promptly took the top spot on a widely watched frontend-coding leaderboard ahead of both Claude Fable 5 and GPT-5.6 Sol, at roughly 40 percent lower cost, with the full weights slated for public release. David Sacks, the administration's former AI czar, called the moment concerning and warned pointedly against federal pre-approval of frontier models: "This is how you lose the AI race."

That benchmark race deserves a nuanced reading, because it is a comparison of released models. Kimi K3 caught the American models that were public at the time - which is to say, the models that had cleared the quarter's gates. The American frontier as it actually exists inside the labs sits further out: the evaluation that reached Hugging Face was run in part on an unreleased model more capable than GPT-5.6 Sol, and Mythos-class capability remains behind vetted-access walls. In one sense, that should reassure anyone anxious about the race: the labs are not being caught; their released versions are. But for security policy the distinction is cold comfort, because defenders can only field the models that ship - through general release or a trusted-access program. Every additional week of gate converts a slice of the American lead from a defensive asset into a capability overhang: still real, still American, and unusable by the hospitals, utilities, and incident responders it is supposed to protect, even as adversaries run last quarter's open weights without asking anyone's permission. And the longer the gates hold, the more of the frontier competition migrates inside the labs, where neither the public nor the government can measure it.

None of this is an argument against open models, which are one of America's real strengths - and the deeper competitive stake has little to do with any benchmark: what matters is whose stack the rest of the world builds on, and China is actively working to make its open models the global default for what a downloaded, embedded AI system looks like inside another country's infrastructure. American open weight models are a national asset, and the government has rightly treated them as one, extending Llama access to NATO and close allies for national security use, on the sound theory that it is in the democratic world's interest for American open models to win. The point here is narrower and more urgent: the American advantage, across open and closed models alike, is a degrading asset, and policy should be built to spend it wisely rather than to sit on it.

Why delay is the wrong lever

The national security worry driving the new caution is, of course, valid. Frontier models measurably uplift offensive cyber operations. Independent evaluators have documented models advancing, in about eighteen months, from barely making progress on a realistic simulated enterprise intrusion to completing more than half of it, at a cost per attempt measured in tens of dollars. RAND's human uplift work points in the same direction. And the Hugging Face incident shows those measured capabilities operating against real infrastructure - chained zero-days, stolen credentials, and remote code execution on a production network, executed autonomously in pursuit of a narrow goal. A government reviewing these systems before release is responding to a real signal.

But the signal has to be read correctly, and correctly read, it indicts delay rather than justifying it. Whether AI ultimately favors cyber offense or defense in the long run is genuinely contested, and probably depends on choices being made right now. Two things, however, are not contested. Offensive capability is cheap, portable, and - once in open weights - permanent. It travels as a download, tolerates failure, and can be rerun at scale for pennies. Defensive capability is the opposite: capital-intensive and access-dependent, because it has to be built into, integrated with, and scaled across the specific systems being protected. An exploit works anywhere; a patch has to be written, tested, and deployed everywhere.

That asymmetry is the whole case, and delay operates almost entirely on the wrong side of it. Holding back an American frontier release does little to slow the offense, which is already diffusing through open weights the government cannot restrict, and a great deal to slow the defense, which depends precisely on defenders getting the best available model, integrated and at scale, as early as possible.

Consider where the defensive bottleneck actually sits - the point in the pipeline where added capability buys the most security. It is not discovery. When AI finds vulnerabilities faster than institutions can fix them, the scarce resources are remediation capacity and defender access to the models doing the finding, and the public record on this is now extensive. Anthropic's restricted Mythos Preview surfaced more than 10,000 high or critical-severity vulnerabilities in its first two months of vetted-partner use, and reporting weeks into the program found only a fraction had yet been fixed. That gap is not an argument for slowing the models down; it is an argument for getting the same class of capability into the hands of defenders and maintainers faster and wider. The most encouraging defensive results of the past two years all point the same way: Google's Big Sleep agent has found real vulnerabilities in critical open source software and even disrupted an exploitation attempt in the wild, and DARPA's AI Cyber Challenge produced systems that autonomously patched injected vulnerabilities in an average of 45 minutes. To the extent the balance currently favors defenders at all, it does so because defensive scanning and detection scale cheaply - but only when defenders can actually field the models. Every one of those wins required defenders with hands on frontier-grade capability.

The Hugging Face incident compresses the asymmetry into a single week. The offensive capability at issue operated with refusals deliberately reduced - inside a lab, for legitimate testing purposes - and promptly reached a third party's production network. The defensive capability the victim needed was locked behind refusals that stayed on, because the hosted models could not recognize a legitimate incident responder handling attack artifacts at speed. What worked was capability the defender controlled: an open weight model, vetted in advance, running on its own hardware. OpenAI has since brought Hugging Face into its trusted-access program - the right decision, made after the emergency it existed to answer. Delay is a tax paid disproportionately by the side of the ledger the country is trying to protect, and the tax comes due at the worst possible moment.

The offensive skew of the technology, in short, is not a reason to hold releases back; it is the strongest reason to move them forward into the right hands. The marginal security cost of releasing an American frontier model is low and falling, because the offense is already loose. The marginal defensive benefit of getting that model to defenders quickly is high and rising, because access is where the defense is constrained. A policy that inverts those weights - treating release as the danger and delay as the safeguard - optimizes for the wrong threat.

The administration's own better theory

Thankfully, EO 14409 appears to understand this. The order directs CISA to facilitate access to defensive tools, including covered frontier models where appropriate, for federal agencies, state and local authorities, and the operators of exactly the soft targets that most need help, such as rural hospitals, community banks, and local utilities. It stands up a Treasury-led AI cybersecurity clearinghouse - now established as "GOLD EAGLE" - to coordinate vulnerability scanning, validation, and the prioritization and distribution of patches, in voluntary collaboration with industry. It builds the review mechanism the right way in principle: a classified benchmarking process to define, on the basis of demonstrated cyber capability, when a model becomes a covered frontier model, and a voluntary framework for up to thirty days of early government access before broader release. And it draws the one bright line that matters most, stating in plain terms that nothing in it authorizes a mandatory licensing, preclearance, or permitting requirement for releasing new models.

This is, on paper, a defender-first, partnership-first, innovation-preserving design. And the sequence matters: the order came first, on June 2, and the quarter's hardest-to-square actions came after it. The June 12 directive was a sudden, sweeping, retroactive restriction on models already on the market, arriving by letter and taking effect globally in hours. The GPT-5.6 hold, whatever its specific justification, showed how a bounded, voluntary window can start to read as an open-ended gate - legally voluntary, practically preclearance.

The question is which becomes the pattern: the order's text, or the deviations from it. There is real evidence the system can land on the text. The Anthropic standoff ended with the controls lifted and access restored; the GPT-5.6 review concluded in release on roughly the clock the order contemplates; and the White House itself pushed back on the idea that any approval had been granted or required, insisting that release decisions rest with the companies. Those are the reflexes of an administration that knows what its own framework says. The work now is to make convergence on that framework the norm rather than the salvage - because the quarter's cumulative signal, left uncorrected, is that access to an American frontier model is contingent, revocable, and unpredictable, and that signal has real costs the order was plainly written to avoid.

Consistency is a security asset, not a concession

The case for releasing frontier models to defenders and the case for doing it predictably are the same case, because unpredictability is itself a defensive vulnerability.

Consider who actually relies on stable access. A hospital system or a regional utility deciding whether to build its threat-detection pipeline on a given model needs to know the model will still be there, and still be supported, next quarter. An allied government weighing whether to build its critical infrastructure on the American stack is making a multi-year bet on availability. A downstream American startup building a security product on a frontier model is making the same bet with its runway. Every one of these defenders is deciding whether to depend on the American ecosystem, and every discretionary hold, every abrupt cutoff, is evidence entered against that decision. When the always-available alternative is a capable Chinese open weight model - downloadable today, and beyond any government's power to claw back once it is running on the defender's own servers - unpredictability in American release policy is not a neutral cost. It is a subsidy to the competing stack. Hugging Face's after-action advice to fellow defenders makes the stakes plain: have a capable model you control, vetted before the incident. If American policy makes American models the ones a defender cannot plan around, that advice points everywhere but here.

The United States has run this experiment before, in adjacent hardware - and the instructive part is not the restriction but the whiplash. Advanced AI chips have been export-controlled since 2022, but the line has moved repeatedly: the H20 was barred in April 2025 and cleared again roughly three months later, and in January the H200 moved from presumption of denial to case-by-case licensing subject to a 25 percent surcharge. Reasonable people can defend any one of those decisions on its own facts. What buyers took from the sequence was that American supply is a variable to be hedged against, and Huawei built the hedge - a domestic accelerator line scaling toward roughly 600,000 units this year, now sold beyond China's borders on the strength of simply being there. A hospital, an allied ministry, or a security startup choosing which model to build on is running that same calculation one layer up the stack.

The lesson is not that every control is futile. It is that policy built around restriction and unpredictability tends to accelerate exactly the outcome it fears, by handing rivals both the incentive and the marketing case to become the reliable global default. AI model policy is now flirting with the same error, and with a special irony: the models most dangerous to U.S. interests are the most capable ones, which are closed and therefore gateable, while the models most consequential for global adoption are open, already in production, and beyond anyone's power to un-publish. A policy that conflates the two - reaching for the gate because the danger is real, and in doing so making the closed American frontier less reliable than the open foreign alternative - constrains American labs and American defenders far more than it constrains the diffusion it is worried about.

Predictability flips the incentive. A frontier-model regime that is clear about what triggers review, bounded about how long review takes, and firm that review ends in release is one that defenders, allies, and builders can plan around. That is how the American stack becomes the obvious choice rather than the risky one - not by being less available than China's, but by being at least as reliable and considerably more capable.

A starting point for consistent rules

The path from the current improvisation to a durable framework for harnessing frontier models for cybersecurity runs through the executive branch, and within this cyber lane it requires no new legislation: the executive order already contains the architecture, and what is missing is the discipline to make its logic legible and predictable. (The adjacent debate over frontier safety and security writ large - catastrophic and CBRN-class risks above all - is a different matter. Building an enduring oversight structure there, with the appropriations and authorities to match, will require Congress, and nothing in this piece suggests otherwise.) Six moves would do most of the work.

  1. Publish the thresholds, not the secrets. The classified benchmarking process that defines a covered frontier model can and should stay classified in its specifics. But developers need to know, in advance, the categories of capability and the general thresholds that trigger review, so they can predict which models qualify and design their evaluation and release plans around a known rule rather than a case-by-case verdict. Predictability for developers is predictability for everyone downstream of them.
  2. Make "up to 30 days" a hard ceiling, with release as the default. The window in EO 14409 is a sensible number, and the risk is that "up to" becomes "until further notice." The default at the end of the window should be broad release. Any extension should require a specific, articulated, senior-level national security finding, on the record and time-limited - not a quiet continuation. An exceptional hold for a genuinely exceptional model is defensible, but a discretionary hold as the ambient norm is not.
  3. Make the review repeatable, not bespoke. Whatever institutional form pre-release evaluation ultimately takes - and there are several live proposals, with reasonable disagreement about the right design - its essential feature should be that a developer can anticipate the test rather than negotiate it. A review whose general shape is known in advance, consistent across developers, and verifiable by the government is faster for everyone, fairer across labs, and far less prone to becoming a discretionary chokepoint than a bespoke engagement invented anew each release cycle.
  4. Pair every restriction with a defender-access commitment - arranged before the emergency. This is the fundamental principle. If a model is judged too capable to release broadly, the correct response is to make it more available to vetted defenders and critical-infrastructure operators, not less - through the CISA and GOLD EAGLE channels the order creates, and through the trusted-access architectures the labs have already built in Project Glasswing and Daybreak. Those programs also teach the operational lesson of the Hugging Face incident: access must exist before the incident that requires it, on published criteria, with bounded scope, logging, and revocability - and hosted safeguards need a way to recognize a verified incident responder, so that the defender analyzing an attack is not treated identically to the attacker who launched it. Capability denied to attackers must never become, by default, capability denied to defenders. A restriction that fails this test is optimizing in the wrong direction.
  5. Keep export enforcement narrow, rule-based, and separate from release policy. Genuine export-control tools have a role, but their power comes from precision. A sudden, sweeping, retroactive cutoff of access to an already-released model - to software, rather than to hardware - carries significant reputational costs for the American AI stack in exchange for often uncertain security benefits. Where restriction is warranted, it should be targeted, forward-looking, and grounded in published rules an ally or a builder can anticipate, so that partners keep building on American models with confidence rather than hedging toward the alternative.
  6. Treat internal testing as part of the security surface. The most consequential AI cyber incident of the quarter came not from a release but from an evaluation. As labs probe maximal capability with safeguards deliberately reduced - testing that is necessary and should continue - the surrounding containment has to be engineered to the level of the capability being tested, because the risk now demonstrably extends to third parties. OpenAI's own response points in the right direction: tightened infrastructure controls at an acknowledged cost to research velocity, disclosure to the affected party and the public, and coordinated remediation. Pre-release review should credit exactly that - containment architecture, monitoring, and incident-disclosure practice, alongside raw capability scores. A framework that watches only the shipping decision misses where a growing share of frontier capability actually operates.

Taken together, these are not a loosening of oversight. They are oversight with a spine: a government that knows what it is testing for, tests for it consistently, finishes on a clock, and routes the resulting capability to the nation's cyber defenders.

Considering the counterarguments

Three counterarguments deserve a fair hearing. The first is that some models really will be too dangerous to release on any fixed schedule. That is true, yet a rules-based regime handles the genuine outlier better than ad hoc discretion does, because a clear, escalated, on-the-record exception is more defensible, less gameable, and less corrosive to the ecosystem than a discretionary norm that treats every release as a potential hold. Consistent rules do not mean releasing a dangerous model on a timer. They mean making the rare true exception rare and legible instead of routine and opaque.

The second is that delay buys defenders time to prepare. It can - but only if defenders get the model, and the Hugging Face incident showed what happens when they do not. A hold that also withholds the capability from the defensive side buys time for no one except the adversary already running open weights, and a trusted-access grant extended after the intrusion arrives too late to matter. If the goal is preparation, the instrument is early, pre-arranged defender access - which the executive order already contemplates and the labs' own programs already implement - not broad delay.

The third is that predictability helps adversaries plan, too. It does, at the margin. But adversaries are the actors least constrained by American rules and best served by the always-available open alternative; they are not the ones deciding whether to trust the American stack. The overwhelming beneficiaries of a predictable regime are the defenders, allies, and builders who operate inside the rules - which is precisely the coalition the United States needs on its stack.

The choice

The American lead in frontier AI is real, it is narrow, and - at least in the versions of the technology anyone outside the labs can use - it is narrowing. The only question is what the country spends it on. Spent well, it arms American defenders first, hardens the soft targets that cannot defend themselves, and anchors allied and commercial infrastructure to a stack that is both more capable and more reliable than the alternative. Spent poorly - dribbled through discretionary holds and sudden cutoffs that slow American defenders without slowing an offense that has already escaped into open weights - it buys very little security and forfeits a great deal of advantage.

Executive Order 14409 points in the right direction: defenders first, partnership over permission, no licensing regime. The quarter's own record shows the government can land there when it chooses to. The work now is to make that choice consistent enough to rely on - because the next Hugging Face will not have weeks to arrange access after the fact. Delay is not a security strategy. Getting the best defensive tools to the right hands, on a predictable clock the whole world can plan around, is.

Anton van Seventer is Counsel for Privacy and Data Policy at the Software & Information Industry Association.

SIIA - Software & Information Industry Association published this content on August 10, 2026, and is solely responsible for the information contained herein. Distributed via Public Technologies (PUBT), unedited and unaltered, on August 10, 2026 at 22:14 UTC. If you believe the information included in the content is inaccurate or outdated and requires editing or removal, please contact us at [email protected]