Wispr Flow turns your voice into clean, ready-to-send writing — speak naturally, it strips the filler and fixes the punctuation. I've used it daily since February 2026 to build this newsletter. Read the full review →

On June 2, an executive order gave the federal government 60 days to build a process for testing whether frontier AI models can find and exploit software vulnerabilities.

The deadline was August 1. The framework was finished. On August 4 the White House walked a room full of AI companies through it — Meta, Nvidia, Microsoft, OpenAI, Anthropic, and a number of smaller firms.

Then it declined to publish it.

The criteria are classified. The labs being tested know what they are measured against. Regulators, security researchers, and the businesses deploying these models do not.

What the order actually requires

The executive order is public even though the framework built from it is not, and it is worth reading plainly.

Executive Order 14409, signed June 2, directs officials to develop and maintain a classified benchmarking process to assess the advanced cyber capabilities of AI models. Classified is not an interpretation of the order. It is the order.

Alongside it sits the voluntary side: developers would give the government access to covered frontier models for up to 30 days before release, subject to confidentiality, cybersecurity, insider-risk, and intellectual property protections.

And then the sentence that defines the whole arrangement. The order states that nothing in it authorizes the creation of a mandatory governmental licensing, preclearance, or permitting requirement for the development, publication, release, or distribution of new AI models.

That is not boilerplate. It is the load-bearing beam. There is no permit and no gate a model must clear before shipping. A lab that does not want a federal review can decline and release anyway. You can read the executive order itself here →

1,000+ Claude Prompts Top Professionals Actually Use at Work

Claude can be your analyst, editor, and strategist.

But most professionals are using it to fix grammar.

These 1,000+ Claude prompts take it from grammar tool to your most powerful AI work assistant.

Sign up for Superhuman AI and get:

  • 1,000+ ready-to-use Claude prompts to get real work done in minutes — researched, tested, and used by professionals at Google, Microsoft, and NASA

  • Superhuman AI newsletter (4 min daily) so you keep learning new AI tools and skills to stay ahead in your career — the prompts are just the beginning

Why a lab would say yes to something optional

A voluntary program with no penalty for refusing sounds like a program nobody joins. The reason the labs are in the room is ten weeks old.

On June 12, Anthropic received a government letter at 5:21pm Eastern — an export-control directive suspending all access to its Fable 5 and Mythos 5 models by any foreign national, anywhere in the world, including its own foreign-national employees.

There was no practical way to comply selectively. The company disabled both models for every customer on earth.

By Anthropic's account, the letter cited national security authorities but did not provide specific details of the concern. The dispute later surfaced as a disagreement over a jailbreak technique: the administration considered it severe, Anthropic said the demonstration it reviewed surfaced a small number of previously known, minor vulnerabilities. The restrictions were lifted around July 1.

That is the incentive. Not a fine, not a license — the memory of a flagship model going dark globally, overnight, with no published criteria and no appeals process. Against that, a voluntary 30-day review looks less like regulation and more like insurance.

The category it appears to leave out

According to reporting by Politico, citing three people familiar with the discussions, the framework broadly exempts open-weight and open-source models and would primarily affect the most capable closed models from OpenAI, Anthropic, and Google.

Treat that as well-sourced but unconfirmed — it comes from anonymous accounts of a classified document, which is the only form this information can currently take.

If it holds, it matters more than the secrecy does. On July 27, eight days before that meeting, NVIDIA open-sourced NOOA as its first named technical contribution to a new Open Secure AI Alliance formed with the Linux Foundation. It is an Apache 2.0 Python framework that collapses an AI agent into a single Python class. It is deliberately model-agnostic — you point it at whatever model you like.

That is the shape of the gap. A governance framework aimed at a handful of closed frontier models, and a tooling ecosystem built so the model underneath is an interchangeable part.

Wispr Flow turns your voice into clean, ready-to-send writing — speak naturally, it strips the filler and fixes the punctuation. I've used it daily since February 2026 to build this newsletter. Read the full review →

Why this lands on your desk

You are probably not submitting a frontier model for federal review. The relevance is downstream, and it is concrete.

When a vendor tells you their model has been evaluated for security risk, you cannot check what that sentence means. Not because the vendor is hiding something — because the standard itself is classified. The usual move for a diligence question, going and reading the benchmark, is unavailable.

Enterprises are used to assurance they can inspect: SOC 2 reports, penetration test summaries, published benchmarks, model cards. Those are all still there. Underneath them now sits a government evaluation that is real, is happening, and is unreadable.

The consequence is not that you should distrust the models. It is that this particular claim cannot carry weight in a risk assessment, because it cannot be examined. If a vendor cites federal review as a reason to be comfortable, the honest response is that you have no way to evaluate that, and it should be scored as unverifiable rather than as a pass.

The takeaway

A classified vetting framework is better than no framework. The capability it targets — models that can find and exploit software vulnerabilities at scale — is one of the few AI risks that is concrete, measurable, and already partly here.

But it creates something genuinely new: accountability without transparency. The labs know the test. The government knows the test. Everyone whose business depends on the outcome is asked to accept it on trust.

That may be the right trade. It is impossible to say from outside, which is precisely the point. What you can do is stop treating it as assurance — know which of your AI risk claims you can verify and which you cannot, and never let the second category quietly get filed as the first.

The site version has the full breakdown, a table separating what is on the public record from what stays classified, and the copy-paste prompt that runs the audit on your own stack.

AI Insights That Turn Your Data Into Bigger Profits

Your store is full of hidden growth opportunities. StoreClaw analyzes your Shopify and Amazon data, identifies the highest-impact actions, and helps you grow revenue while improving margins. Start free today with 10,000 bonus tokens. No credit card required.

— Jerry
AI Super Simplified