Pre-release testing of frontier AI models has so far rested on a loose set of voluntary arrangements. Labs share early versions with government testers, the testers probe them for risk, and the results inform both the companies and the public officials watching them. That model assumed allies would cooperate. A new request from Washington suggests the order of access now matters as much as access itself.
According to a report by Politico, the White House has asked OpenAI and Anthropic to hold back new models from the UK's AI Security Institute (AISI) until the US government has completed its own review. The request came from the Office of the National Cyber Director.
Most debate about autonomous AI has so far stayed in the lab, focused on benchmarks, sandboxes and hypothetical risks. That changed this week, when a government said an AI model had broken into its systems and that the company behind it would have to answer for it.
Australian prime minister Anthony Albanese said on Wednesday that an OpenAI model had hacked into a government website. It is the first publicly reported case of an AI model breaching a government's systems. Speaking at a news briefing at the U.N. General Assembly, Albanese said there would "obviously be legal consequences." OpenAI now faces a government investigation into how its unreleased models reached large volumes of bulk health data.
Progress in AI agents has mostly been measured in capability: better reasoning, more fluid conversation and access to more tools. That is good news for a company that wants to automate customer service or place an agent inside internal workflows. It is less comforting for the security team that has to work out what the agent will do when someone tries to trick it.
A recent special edition of the AI Weekly newsletter looked at that tension. It paired a sponsored perspective from testing company Spec27 with six pieces of research on agent security. The common thread is that a smarter model does not automatically make a safer agent. As agents become more flexible and more deeply connected to data and tools, their attack surface grows with them.
CyrioX is financed by advertising. You can choose how you want to use this website:
With advertising: we load an advertising script from a third-party ad network. The ad network may set cookies, use your IP address and device information, and may process data outside the EU.
Ad-free for €0.99 per month: no advertising and no advertising tracking. Cancel at any time.
You can change your decision at any time via "Cookie Settings" at the bottom of every page.