- EU Watermarks: OpenAI plans invisible text patterns that detectors can recognize in eligible EU output from ChatGPT and its Codex coding agent.
- Developer Choice: Developers using OpenAI models in their own services can opt in worldwide; watermarking is off by default.
- Editing Limits: Copying preserves the pattern, but editing, translation and short passages can weaken detection.
- Detector Access: OpenAI initially limits its text detector to approved researchers and expert organizations.
OpenAI plans to add invisible watermarks to eligible text from ChatGPT and its Codex coding agent in the European Union, giving detectors a way to recognize its AI-generated output after the words are copied elsewhere. The EU rollout is planned for the coming weeks; developers using select OpenAI models in their own services can opt in worldwide from October 5, 2026.
The EU rollout covers eligible users across all plans. The API, through which developers use OpenAI models in their own services, keeps watermarking off by default and supports select models. Access to the detector requires separate approval, so choosing to generate marked text does not automatically let a developer check it.
OpenAI had already developed a word-choice watermark by 2024, when The Wall Street Journal reported that its release was still under debate.
How the Watermark Travels With Text
OpenAI’s method, called textGrain, puts the marker in the model’s choices of words and word fragments, rather than adding a visible symbol. Because the wording carries the pattern, it remains when someone copies and pastes the passage.
A technical report, co-written by OpenAI and researchers at the University of Pennsylvania and Yale University, explains how those choices become recognizable. As a language model produces one token, a word or part of a word, at a time, a secret key and the preceding text divide possible next tokens into groups and influence which group supplies the next token. Within the chosen group, the tokens keep their original relative probabilities.
For the example sentence “The morning was cold and bright,” the detector uses the key and preceding words to reconstruct a score for “cold.” It repeats that process across the passage and combines the scores to test for the watermark. The detector can do this from the text and matching key and settings, without running the model that wrote it.
textGrain also controls how much randomness generation gives up to create the signal. That matters when an application needs several different answers to the same prompt: some fixed-key methods can repeatedly choose the same answer. Its “entropy budget” limits the average loss of that randomness.
Longer Passages Are Easier to Detect
More words give the detector more opportunities to recognize the pattern, while tightly constrained answers give the model fewer choices. In OpenAI’s tests on psychology questions from the ELI5 question-answer dataset, it detected about 80% of watermarked 200-token passages and 95% of 400-token passages at a target false-positive rate of 1%. That rate describes incorrectly flagging unwatermarked text; mathematics responses were substantially harder to detect.
A separate test used watermarked, 400-token English ELI5 responses. Detection fell from about 92% for the unedited passages to 66% after 10% of their words were replaced with synonyms, and to 17% after 25% were replaced. Translation can also weaken detection. An undetected passage may therefore still contain AI-generated text.
OpenAI says the marker indicates that its system generated or processed part of a passage. It does not measure how much human judgment or editing contributed, and the company says it carries no association with a user, account, prompt or conversation.
Working Code Is a Different Measurement
OpenAI reports no meaningful performance difference in benchmark tests comparing watermarked and unwatermarked output from a single configuration of its Astra language model. On DeepSWE v1.1, a coding benchmark, it scored 72.80% without watermarking and 71.68% with it.
| Benchmark | Unwatermarked text (Astra, max) | Watermarked text (Astra, max) |
|---|---|---|
| Artificial Analysis Intelligence Index | 49.57 points | 49.76 points |
| AutomationBench | 34.09% | 34.86% |
| DeepSWE v1.1 | 72.80% | 71.68% |
| Terminal-Bench 4.0 | 53.90% | 56.06% |
| Terminal-Bench Science 0.1 | 56.90% | 60.00% |
| BrowseComp | 87.92% | 87.35% |
| HealthBench Professional | 64.27% | 64.60% |
| GPQA Diamond | 94.44% | 93.94% |
But successfully completing a coding task and leaving a detectable watermark are separately measured outcomes. Which Codex outputs count as eligible and whether it has code-specific detection rates remains unclear as of today.
A September 9 preprint by Alexander Nemecek, Vipin Chaudhary and Erman Ayday at Case Western Reserve University shows why the distinction matters. Their experiment with Google DeepMind’s public SynthID text-watermarking implementation tested two open-weight models, whose trained parameters are publicly available, on 364 programming problems, with ten samples per problem in each test condition. Watermarked code performed similarly to unwatermarked code on Gemma-2-9B, while correctness fell by 3.1 percentage points on Llama-3.1-8B.
The researchers scored the same programs with a SynthID detector that averages token-level watermark scores, known as the raw mean g-value method. The public implementation also offers a trained detector. The raw-mean detector barely separated marked from unmarked code: its AUROC scores were 0.55 and 0.57, close to the 0.5 chance reference. AUROC measures separation across detection thresholds. The outputs were generally short, and code offered few discretionary word choices.
Developers Face Different Defaults
Anthropic’s Claude assistant already uses text watermarks on supported models worldwide, including its API and supported cloud channels. Support varies by model; some Amazon Bedrock rollouts are scheduled to finish by October 12. Unlike OpenAI’s optional API setting, The New Stack’s comparison found no equivalent developer opt-out described by Anthropic.
Anthropic attributes its global approach to lacking a durable way to restrict watermarking by region. Its detection API is in private preview for eligible organizations, including media, researchers and certain enterprises. OpenAI initially grants access to approved researchers and expert organizations, with applications reviewed individually.
Google DeepMind’s SynthID also marks text from the Gemini app and web experience. Google had already open-sourced its text-watermarking implementation by May 2025, while OpenAI plans to release textGrain as open source. These are ways to build a marker into generation; a conventional AI-text classifier instead looks afterward for stylistic or statistical signs of machine writing.
What the EU Requires
The EU AI Act’s transparency obligations apply from August 2, 2026. Article 50 requires providers of systems generating synthetic content to mark their output in a machine-readable format and make it detectable as artificially generated or manipulated. Its robustness and reliability requirements are qualified by technical feasibility, and the marking duty has exceptions for standard editing or changes that do not substantially alter the supplied content.
The Code of Practice on Transparency of AI-generated Content is a voluntary tool for meeting those legal duties. A separate Article 50 disclosure duty applies to AI-generated or manipulated text published to inform the public on matters of public interest. It includes an exception where the text has undergone human review or editorial control and a person or organization holds editorial responsibility for publication.


