AI text watermarks are easy to misunderstand because nothing is hidden in the characters. There is no invisible symbol, no zero-width space and no secret metadata riding along in the words. What exists is a statistical fingerprint created while the text is being generated, which a matching detector can look for afterwards. That design has real strengths and some limits that are documented rather than speculative, and the limits are the part worth knowing if you are ever asked to rely on a detector’s verdict.
Updated October 2026.

What AI text watermarks actually are
NIST defines digital watermarking as embedding information into content, whether image, text, audio or video, typically while making it difficult to remove. For text the property being perturbed is the model’s own choice of words. A language model produces text one token at a time, each drawn from a probability distribution over the vocabulary. A watermarking scheme adjusts that distribution so that, where several words would do, the model prefers a particular subset determined by a secret key. NIST lists the families in use, including red and green list schemes, distortion-free schemes and semantic-invariant schemes, all applied during generation, and notes that some versions may degrade text quality.
Google DeepMind’s SynthID is the most widely deployed example and describes its text approach in the same terms: it adjusts the probability scores used during generation to produce a watermark that is not noticeable to the reader. The reference implementation published alongside the research modifies the model’s output logits using a configuration of keys, an n-gram length and sampling parameters, and offers two detectors, a weighted mean detector that needs no training and a more powerful Bayesian detector that must be trained on watermarked and unwatermarked examples. Detection produces a score between 0 and 1 rather than a verdict. The repository is candid that it exists for reference and research reproducibility, and that its hashing function provides no guarantees of cryptographic security. The underlying work was published in Nature as Scalable watermarking for identifying large language model outputs.
5 hard limits
- Short and predictable text carries no signal. NIST states that text watermarks generally cannot be embedded or detected reliably when the text has low entropy, which covers short answers, formulaic prose and anything with only one sensible wording.
- Paraphrasing works. Having a separate non-watermarked model paraphrase the output can often remove a text watermark with only minor degradation of quality. NIST notes paraphrasing substantially reduces detection accuracy for short texts and only slightly for longer ones, beyond about 400 words.
- Targeted editing defeats longer texts too. If an attacker knows which tokens a watermark prefers, paraphrasing can be aimed at exactly those, which NIST says can remove watermarks even from longer texts. A simpler trick works as well: ask the model to scatter an irrelevant word through the output, then strip it, breaking the correspondence between each adjustment and its preceding context.
- Watermarks can be forged. NIST reports research in which an attacker used their own non-watermarked model to generate text that detectors classified 80 percent of the time as bearing another model watermark. A positive result is therefore not proof of origin.
- False positives are the expensive error. Any detector has a nonzero false positive and false negative rate, and NIST observes that assessing human content as AI-generated can be extremely damaging, potentially causing major reputational harm or adverse treatment.
Content Credentials are a different mechanism
Provenance metadata is often discussed in the same breath as watermarking, and it works in the opposite way. In the C2PA specification a manifest is the set of information about an asset’s provenance, built from one or more assertions including content bindings, a single claim, and a claim signature. An assertion is a statement about the asset made by the signer or gathered when the claim is generated. The claim is a digitally signed, tamper-evident structure referencing those assertions, and a hard binding is one or more cryptographic hashes that uniquely identify the asset or part of it. The signature is applied with the signer’s private key.
The specification is careful about what that proves. Its guiding principle is that it should not provide value judgements about whether a given set of provenance data is good or bad, only whether the assertions can be validated as associated with the underlying asset, correctly formed and free from tampering. Validation tells you the record has not been altered and who signed it. Whether to trust that signer is a separate question. NIST adds the practical weakness: metadata recorded within a file can be stripped altogether, as it often is when files are shared, sometimes to deceive and sometimes for ordinary privacy reasons. Covert watermarks are designed to persist where metadata does not.
Detection without a watermark is worse
Tools that claim to identify machine-written text by its style alone are a weaker proposition than any of the above, because they are inferring from surface features rather than reading an embedded signal. NIST cites research under the title “GPT detectors are biased against non-native English writers”, which is the finding that should settle how such tools are used in practice. Combine that with the damage a false positive does to the person accused, and the responsible conclusion is that a styleometric detector score is evidence of nothing on its own. If you are trying to judge whether a particular claim is true rather than who wrote it, our guide to checking an AI answer is the more useful tool, and spotting AI generated images covers the visual equivalent.
What the law now requires
Article 50 of the EU AI Act requires providers of systems that generate synthetic content to ensure outputs are marked in a machine-readable format and detectable as artificially generated or manipulated, with technical solutions that are effective, interoperable, robust and reliable as far as technically feasible, taking account of cost and the state of the art. Deployers who create or alter image, audio or video content constituting a deepfake must disclose that it has been artificially generated or manipulated, with carve-outs for law enforcement and for artistic, creative, satirical and fictional works where disclosure must still happen without spoiling the work. The information has to be provided clearly and distinguishably at the latest at the time of first interaction or exposure. Our explainer on the EU AI Act sets out the wider timetable. The honest summary of AI text watermarks is that they are a useful signal at scale and a poor basis for accusing an individual.
Common questions
Is anything hidden in the characters of watermarked text? No. The watermark is created by adjusting the probability scores the model uses while generating, so it lives in the pattern of word choices rather than in invisible characters or metadata.
Can paraphrasing remove an AI text watermark? Often, yes. NIST notes that a separate non-watermarked model paraphrasing the output can remove the watermark with only minor quality loss, and that detection falls sharply for short texts.
Why do detectors fail on short passages? Because text watermarks generally cannot be embedded or detected reliably in low entropy text. A short or formulaic passage simply does not contain enough free word choices to carry a signal.
How is C2PA different from a watermark? C2PA attaches a signed manifest of assertions bound to the asset by cryptographic hashes. It proves the record has not been tampered with, not that the signer is trustworthy, and NIST notes such metadata is often stripped when files are shared.
Can a detector prove a student used AI? No. Detectors return probability scores, watermarks can be forged so that text is flagged as carrying another model mark, and NIST cites research finding that detectors are biased against non-native English writers.
Sources and further reading
Where the figures and rules above come from, so you can check them:
- Definition of digital watermarking, low entropy limits, paraphrasing, spoofing and false positives: NIST AI 100-4, Reducing Risks Posed by Synthetic Content
- How SynthID watermarks text by adjusting probability scores: Google DeepMind
- Reference implementation, detectors, scores and stated caveats: SynthID-Text repository, Google DeepMind
- Manifests, assertions, claims, claim signatures and hard bindings: C2PA Specification 2.1
- Article 50 marking and disclosure obligations: EU Artificial Intelligence Act text
Photo credit: Code on computer monitor (Unsplash) by Markus Spiske markusspiske, CC0, via Wikimedia Commons.
Join the discussion