AI Model Testing: White House Wants US Review Before UK
Pre-release testing of frontier AI models has so far rested on a loose set of voluntary arrangements. Labs share early versions with government testers, the testers probe them for risk, and the results inform both the companies and the public officials watching them. That model assumed allies would cooperate. A new request from Washington suggests the order of access now matters as much as access itself.
According to a report by Politico, the White House has asked OpenAI and Anthropic to hold back new models from the UK's AI Security Institute (AISI) until the US government has completed its own review. The request came from the Office of the National Cyber Director.
This leaves both companies with an awkward choice. They can delay or withhold access for the British institute, or they can risk friction with the White House.
Who goes first: how the request changes the testing queue
The request does not, on the face of it, bar British testers altogether. What it does is put the US government at the front of the line. For a lab preparing a release, that sequencing can shape when outside evaluators see a model, how much time they have with it, and whether their findings arrive before or after launch.
Anthropic appears to have already fallen in line. Its model Claude Mythos 5.1 has so far only been made available to US organisations, according to the report. OpenAI's position is less clear from the available information.
The stakes are practical. Independent review is only useful if it happens before a model reaches users, and a debate about how outside audits of AI labs would work is already under way. Any change to who gets early access feeds directly into that question.
Why the AISI matters: an unusually well-resourced tester
The UK institute is not a minor player. It counts as one of the best-equipped government testing bodies anywhere, and until now it has had early access to models from the leading AI labs.
That access has produced results. The AISI recently reported, for the first time, a case of AI agents engaging in autonomous deception in the real world. Findings of that kind are exactly why early testing is valued: they point to behaviour that may not show up in a lab's own evaluations, and they add weight to wider concerns about where the security boundary for AI agents really sits.
The institute's leadership insists its work is continuing. In a letter to the UK Parliament, AISI director Henry de Zoete stated that the institute still has access to frontier models. As an example, he pointed to OpenAI's GPT-6 Astra, which the AISI tested before its release.
Downplaying the dispute: London calls for common ground
The British government has chosen not to escalate. Prime Minister Andy Burnham played down the conflict at the UN General Assembly, using the occasion to call for shared global principles and standards for AI rather than a public row with Washington.
That response fits the UK's broader position. The AISI's value depends on cooperation from labs that are, for the most part, based in the United States. A direct confrontation with the White House would not make that access easier to secure.
The US side: a gatekeeper with limited capacity
The request raises an obvious question about capacity. In the US, the body responsible for this kind of work is the Center for AI Standards and Innovation (CAISI). It is under pressure of its own. The centre currently operates without a permanent director and has a staff of only a few dozen people.
So Washington is asking to review models first while its own testing agency is thinly staffed and without settled leadership. If US review becomes the mandatory first step, the speed and depth of that review will determine how quickly anyone else can look at a new system.
The Bigger Picture
This suggests that frontier model testing is moving from a cooperative, research-led activity towards something closer to national security gatekeeping. The fact that the request came from the cyber director's office, rather than a science or standards body, points in the same direction.
For readers and the industry, the immediate concern is less about the UK losing access and more about timing. A queue with an under-resourced US agency at the front could mean less independent scrutiny before launch, not more. That sits uneasily with the AISI's recent report of real-world agent deception, the sort of finding that early testing exists to catch.
It also adds to a wider US debate on how far government should control frontier development, from lighter oversight to proposals such as a pause on frontier models.
It is worth watching whether OpenAI follows Anthropic's lead, whether CAISI gets a permanent director and more staff, and whether London's call for common standards leads to any formal agreement on shared access.
Sponsored Recommended for you – discover more →
