# How Watermarks Track AI Generated Content - Computerphile

**URL:** <https://www.artofsm.art/t/how-watermarks-track-ai-generated-content-computerphile/23478>\
**Category:** Content Creators\
**Tags:** computerphile, law, security\
**Created:** [3 September 2026 13:30 UTC](https://www.artofsm.art/t/how-watermarks-track-ai-generated-content-computerphile/23478 "2026-09-03T13:30:04Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![artesia](https://www.artofsm.art/user_avatar/www.artofsm.art/artesia/32/36_2.png) [@artesia](https://www.artofsm.art/u/artesia)\
**Post date:** [3 September 2026 13:30 UTC](https://www.artofsm.art/t/how-watermarks-track-ai-generated-content-computerphile/23478/1 "2026-09-03T13:30:04Z")

</div>

[![](https://www.artofsm.art/uploads/default/original/3X/5/c/5ce2edcbc3820b5448d66521eea63feadd6aa275.jpeg "How Watermarks Track AI Generated Content - Computerphile") ](https://www.youtube.com/watch?v=kVXp6UNVPTo)

The video explains how watermarking AI-generated text embeds a subtle, secret signal within language model outputs by probabilistically biasing word choices, enabling reliable detection of AI-produced content without affecting text quality. While effective for natural language, watermarking is less reliable for code and can be partially evaded, serving primarily as a practical tool for identifying AI-generated content rather than an absolute safeguard.

---

<div class="post-metadata">

**Author:** ![artesia](https://www.artofsm.art/user_avatar/www.artofsm.art/artesia/32/36_2.png) [@artesia](https://www.artofsm.art/u/artesia)\
**Post date:** [3 September 2026 13:50 UTC](https://www.artofsm.art/t/how-watermarks-track-ai-generated-content-computerphile/23478/3 "2026-09-03T13:50:33Z")

</div>

The video discusses the emerging practice of watermarking AI-generated text, a method mandated by the EU and adopted by major AI companies like Anthropic and Google to help detect AI-produced content. Watermarking embeds a subtle, secret signal within the output of large language models (LLMs) without noticeably affecting the quality or meaning of the text. This approach is not primarily about preventing plagiarism but about enabling reliable detection of AI-generated content later, even if the text is edited or partially rewritten. The watermarking process relies on secret keys known only to the AI providers, ensuring the watermark’s security and making it difficult for outsiders to detect or remove without access to these secrets.

At the core of the watermarking technique is a clever probabilistic method that slightly biases the choice of words or tokens during text generation. Instead of always picking the most likely next word, the model runs a “tournament” among candidate tokens, scoring them based on a secret hash function that depends on the context and a secret key. This tournament selects tokens in a way that preserves the overall probability distribution of words, so the output remains natural and coherent. Over many tokens, this subtle bias creates a statistical pattern that can be detected later by reapplying the hash function and analyzing the distribution of scores, revealing whether the text was watermarked.

Detection works by examining the tokens in the generated text and calculating their tournament scores using the secret key. If the average score deviates significantly from what would be expected by chance (around 0.5), it indicates the presence of a watermark. The method requires analyzing a substantial amount of text—typically hundreds of tokens—to achieve statistical confidence. While minor edits or partial rewrites can weaken the watermark, removing it entirely would require extensive changes to the text, often degrading its quality or coherence. This makes watermarking effective for detecting AI-generated content in typical use cases, such as essays or articles, especially when no human intervention occurs.

Watermarking is less effective for highly structured outputs like code, where many tokens are deterministic and cannot be altered without breaking functionality. In code, watermarking mainly influences comments, variable names, or import order, resulting in a weaker and less reliable watermark. Detection in code requires analyzing large volumes of output, and it is easier to remove the watermark by editing non-essential parts. Despite these limitations, watermarking still provides a useful tool for identifying AI-generated code at scale, though it is more challenging than with open-ended natural language text.

The video also highlights potential ways to evade watermarking, such as inserting and then removing emojis between words to disrupt the watermark’s context or using open-source models that do not implement watermarking. However, these methods require significant effort and often result in degraded text quality. The watermarking system is designed as a deterrent and a detection aid rather than an absolute barrier. Ultimately, watermarking represents a practical compromise to help identify AI-generated content reliably while maintaining output quality and usability, with ongoing research and development to improve its robustness and applicability.
