How AI Text Watermarking Works: Secret Keys, Green/Red Word Choices, and Detection

✍️ OpenClawRadar📅 Published: August 15, 2026🔗 Source
Ad

AI text watermarking works by hiding marks in the choices between words, not in the text itself. Google has watermarked Gemini app and web text since 2024, and Claude models mark text at the model level as of August 2026. The marks survive copy-paste and are invisible to readers.

Watermarking via word-choice bias

When a language model writes, it picks each word by rolling weighted dice over a shortlist of candidates. Watermarking uses a secret key to color those candidates green or red, then nudges the dice slightly toward green. The text still reads naturally — red words can still win, just less often.

How detection works

With the secret key, you can re-color any text and count how many words are green. In unmarked text, about half the words will be green by chance. In watermarked text, the count is significantly higher. The detector doesn't read the text — it just counts green words over a sufficiently long run.

Ad

What editing does to the mark

The watermark lives in runs of untouched wording. Each word's color is derived from a short window of preceding words, so editing erases the mark exactly where the run breaks. Short texts are hard to call — a 1,500-word document with a mild tilt would flag at only ~55% green, which is why longer texts are more reliable.

Production schemes

  • Google SynthID: Uses a secret tournament instead of a simple nudge, preserving exact word probabilities.
  • Aaronson's scheme: Derives the dice-rolls themselves from the key.
  • Kirchenbauer et al. (2023): The classic green/red bias approach.

Check out the interactive visual guide for a hands-on demo of the process.

📖 Read the full source: HN AI Agents

Ad

👀 See Also

AppLovin Mediation Cipher Broken: Device Fingerprinting Bypasses ATT
Security

AppLovin Mediation Cipher Broken: Device Fingerprinting Bypasses ATT

Reverse-engineering revealed that AppLovin's custom cipher uses a constant salt + SDK key, a SplitMix64 PRNG, and no authentication. Decrypted requests carry ~50 device fields (hardware model, screen size, locale, boot time, etc.) even when ATT is denied, enabling deterministic re-identification across apps.

OpenClawRadar
Sandboxing Local AI Agents with Firecracker MicroVMs
Security

Sandboxing Local AI Agents with Firecracker MicroVMs

A developer created a sandbox that isolates AI agent execution inside Firecracker microVMs running Alpine Linux, addressing security concerns about agents running commands directly on the host machine. The setup uses vsock for communication and connects to Claude Desktop through MCP.

OpenClawRadar
Axios 1.14.1 compromised with malware, targets AI-assisted development workflows
Security

Axios 1.14.1 compromised with malware, targets AI-assisted development workflows

Axios version 1.14.1 has been compromised in a supply chain attack that silently pulls in [email protected], an obfuscated RAT dropper. Developers using AI coding assistants like Claude should immediately check their lockfiles and machines for infection.

OpenClawRadar
OpenClaw Security: The Hardened Baseline You Should Start With
Security

OpenClaw Security: The Hardened Baseline You Should Start With

Self-hosting OpenClaw doesn't automatically make it secure. A Reddit post details the hardened baseline config: local-only Gateway, per-peer DM isolation, deny runtime/fs/automation tool groups, exec locked down, and mention-gated groups.

OpenClawRadar