OpenAI textGrain: ChatGPT Text Watermarks Hit the EU

OpenAI textGrain: ChatGPT Text Watermarks Hit the EU

OpenAI will start hiding a machine-readable signal inside text produced by ChatGPT and Codex for users in the European Union. Outside the chatbot, the approach is looser. Developers using the API anywhere in the world can choose whether to switch it on. This puts OpenAI on a different path from Anthropic, which applies watermarking to Claude everywhere.

Why the EU is the trigger

According to reports, the EU AI Act, the bloc's framework law for artificial intelligence, requires providers to mark AI-generated text in a way machines can read. OpenAI's response is a system called textGrain. It does not add visible labels or metadata. Instead, it nudges the model's word choices so that a statistical pattern sits in the output. A reader cannot see it, but a detector can look for it.

The idea is close to Claude's SynthID watermark, which is built on Google's open-source technology. OpenAI says textGrain performed as well as, or better than, competing methods in its own tests, including Google's SynthID for text. The company also plans to publish textGrain as open source so others can build on it.

The rollout has three parts:

  1. ChatGPT and Codex in the EU: watermarking turns on over the coming weeks.
  2. API worldwide: watermarking is opt-in, not mandatory.
  3. Cloud partners: API watermarking will also reach platforms such as Microsoft Azure in the coming weeks.

Anthropic's watermarking for Claude applies globally, no matter how people access its models. OpenAI's opt-in API model is the clearest difference between the two.

What the detection numbers show

OpenAI shared detection figures, and they depend heavily on how long the text is and what it is about. With the detector tuned to a target false-positive rate of 1 percent, it caught the watermark in about 95 percent of 400-token passages on psychology. At 200 tokens, the rate fell to roughly 80 percent.

Math content performed "substantially" worse, according to OpenAI. The reason is simple. When there are fewer valid ways to phrase something, the model has less room to vary its word choice, so the signal is weaker. Longer passages could give the detector more to work with, but results may still vary by topic. OpenAI has not offered data to back that up.

Editing is the bigger weakness. In OpenAI's figures, swapping just 10 percent of the words for synonyms in a 400-token passage drops detection from about 92 percent to 66 percent. Replace a quarter of the words and detection falls to 17 percent. In practice, a light rewrite is enough to defeat it.

That fragility may not be all bad for OpenAI. If its watermark were much harder to strip out, users who want to keep their ChatGPT use private might move to open-weight models instead.

Anthropic has not released detection rates for Claude's watermark, though it says its system holds up well against edits. OpenAI has published a technical report that goes into textGrain's design in detail.

Quality, and what a watermark does not prove

OpenAI says watermarking does not degrade output. It points to tests of its frontier model Astra, which showed no significant differences with the feature on or off across eight benchmarks, including GPQA Diamond, BrowseComp and DeepSWE. Those benchmarks measure reasoning, browsing and coding, though, not prose. They do not settle whether writing quality changes. Critics have raised the same question about Claude's watermark.

OpenAI is also careful about what a positive result means. A detected watermark says nothing about how much a human edited or contributed to the text. It does not establish ownership, assign responsibility, identify a user, or confirm that the content is accurate. The reverse holds too: no watermark does not prove a human wrote it. The text may be too short, edited, translated, or produced by a model the detector does not cover.

A gated detector

For now, the detector is not public. Selected researchers and specialist organizations can apply through a form, and OpenAI will decide case by case under the EU's Code of Practice, the voluntary guidance tied to the AI Act. Anthropic handles its detection API in a similar way.

The tool returns only one answer: whether an OpenAI watermark was found. It does not reveal users, prompts or conversations.

OpenAI's stated reason for the restriction is that the detector can wrongly flag unmarked text and can miss real watermarks. It says it will widen access "when we believe results can be interpreted responsibly," but gives no timeline. Its existing provenance tools for images and audio, including openai.com/verify and the Content Provenance API, stay publicly available.

Our Take

The pattern here is compliance shaped by geography. OpenAI watermarks where regulation demands it and leaves the choice to developers elsewhere. That suggests regulation, not product philosophy, is setting the floor for text provenance. For readers building on the API, the practical point is that watermarking outside the EU is now a setting you own.

The numbers also deserve caution. A signal that drops to 17 percent detection after replacing a quarter of the words is a weak tool for teachers, editors or platforms hoping to catch AI text. OpenAI itself says a result proves little either way. Anyone treating these detectors as lie detectors is likely to make mistakes.

It is worth watching whether the open-source release lets outsiders test textGrain against SynthID independently, and whether Anthropic publishes its own detection rates. Pressure on AI labs from regulators, such as the FTC probe of OpenAI and Anthropic, could also push other jurisdictions toward EU-style labeling rules. If that happens, today's opt-in default may not last.