OpenAI Halts Reasoning Theft, but Azure Stayed Exposed

OpenAI Halts Reasoning Theft, but Azure Stayed Exposed

OpenAI says it has shut down a coordinated attempt to extract the hidden reasoning of its models. Independent researchers, however, found that the same technique kept working on Microsoft Azure for weeks afterwards. The case shows how a security fix at the source can miss the cloud platforms that resell the same models.

What OpenAI says happened

In a blog post, OpenAI describes what it calls an "adversarial distillation" campaign. Distillation means training one model on another model's output. The most valuable target is the full chain of thought: the intermediate steps a reasoning model works through before it answers. Users normally see only the final answer.

OpenAI says these steps can contain information that is deliberately kept out of the answer, and that they can help a competitor rebuild a model's capabilities.

The timeline, according to OpenAI:

  • July 1: activity starts at low volume.
  • July 24-25: a spike of 16,000 requests from more than 4,000 users, all using a typical extraction pattern.
  • Further analysis: a network of more than 15,000 accounts with related patterns.
  • July 28: OpenAI says the network was fully shut down.

A footnote says these were attempted extractions, not necessarily successful ones. OpenAI links a core group to people associated with Moonshot AI, the maker of the Kimi language model. It adds that it is unclear whether every actor it observed traces back to one source. Anthropic recently reported similar attempts by Chinese AI companies.

How the trick worked

Providers return reasoning to customers only as encrypted data packets, which customers send back with follow-up requests. The attackers copied encrypted reasoning from one conversation and asked a model in a separate conversation to decrypt it and write it out.

Researcher Joachim Schaeffer and his team had already described this in a paper. Because the packets use shared keys, they can move between sessions, between users and even between different models from the same provider. A weaker, cheaper model from the same family can then act as a "decryption oracle" and print the stronger model's reasoning word for word.

OpenAI credits the researchers by name and says their findings helped it ship countermeasures faster. Those measures include banning fraudulent accounts, tightening sign-ups, closing the hole that allowed reuse of other users' encrypted reasoning, and screening streamed output so that it can be held back if it might reveal reasoning. OpenAI says it shared its findings through the Frontier Model Forum and government channels.

The cloud gap

On the same day as OpenAI's post, the researchers published an update. "We stole reasoning. Again," Schaeffer wrote on X.

Testing again on September 13, they found the attack blocked on OpenAI's and Anthropic's own APIs. On Azure, it worked against every OpenAI model they tried, including the new GPT-6 Astra, and against Anthropic models up to Sonnet 5. One attempt was enough to pull out the reasoning verbatim. "Same models, but different protections depending on which platform serves them," Schaeffer said.

There is also a simpler route, shown publicly by developer Can Bölük. The model is given a virtual notepad as a tool and told to write its reasoning there, which the user can then read. According to the researchers, this worked on every OpenAI model and on Opus 4.8 and Sonnet 5. Only Opus 5, Fable 5 and Fable 5.1 did not reveal their reasoning. The researchers say the output closely resembled the decryption results and would likely be just as useful for distillation.

They call the fixes so far piecemeal and superficial. Many rely on brittle matching of specific request patterns, and some reached cloud platforms only days later. GPT-6 Astra launched on third-party platforms without the protections. By the researchers' timeline, OpenAI added safeguards to the Azure endpoint on September 27. For Anthropic models, the extraction could no longer be reproduced on Azure from September 28.

Schaeffer argues that patches must cover every attack type and every cloud that hosts a model, or attackers will pick the weakest route. The paper goes further: cloud providers that do not enforce equivalent protections should not be allowed to serve reasoning models at all. Otherwise, the researchers write, open backdoors could let people sidestep export controls at the API level. OpenAI agrees that partner-hosted models need the same protection as its own services and says the work is not finished.

The Bigger Picture

The main lesson is architectural, not about one exploit. A model's security boundary is not the lab's API. It is every endpoint where the model runs. If a provider hardens its own service but a reseller lags by weeks, the weaker endpoint becomes the real boundary. Developers who buy model access through clouds such as Azure should not assume they get identical protections.

The case also suggests that filters tuned to specific request patterns are a weak defence on their own. The notepad method needed no decryption at all. Fixes that block one known pattern may simply push attackers to the next.

Distillation is a commercial issue and, increasingly, a policy one. The researchers' link to export controls ties it to the wider effort by Chinese labs to build capable models under hardware limits, as seen with Deepseek's tooling for Huawei Ascend chips. It is worth watching whether cloud platforms commit to matching lab-level safeguards at launch, whether regulators pick up the researchers' proposal, and whether OpenAI's expectation of more sophisticated attempts plays out as models improve.