logo

The Watermark Removal Debate: What’s Technically Possible vs. What’s Ethical

Paperbleach
Paperbleach

05 Aug 2026

Can you remove AI watermarks? Technically, yes — in most realistic cases, paraphrasing, translation, or heavy editing will erode a text watermark below the detection threshold, and researchers have argued formally that no watermark can survive a determined attacker. That’s the uncomfortable engineering answer. The more interesting debate is the one it opens: if removal is this easy, what are watermarks actually for, and where is the line between legitimately editing your own draft and deliberately laundering machine output? The two questions — can you, and should you — have very different answers, and conflating them is how most coverage of this topic goes wrong.

Key takeaways

  • Text watermarks are statistical, not structural: they bias word choices in a key-dependent pattern, so anything that changes enough words weakens them.
  • Paraphrase attacks degrade every published scheme, and a 2023 impossibility result argues strong watermarking can’t exist against a motivated attacker.
  • Watermarks still make sense — they cheaply label the unedited majority of AI text, which is what transparency regulation actually demands.
  • Legality is mostly about what you do next: removal itself is rarely illegal, but using stripped text to deceive can be, and terms of service often prohibit it.
  • The ethical line runs through intent: rewriting a draft into genuine authorship is not the same act as stripping a mark to manufacture a false claim.

Can you remove AI watermarks? The technical answer

To understand why removal works, you have to understand what’s being removed. A text watermark isn’t a hidden character or a metadata tag — plain text can’t hold either. Schemes like Google’s SynthID-Text, published in *Nature* in 2024 and deployed across Gemini, work at generation time: the model’s word choices are nudged in a pattern determined by a secret key, so that watermarked text contains statistically more “preferred” tokens than chance would produce. A detector with the key counts those preferences and computes a confidence score. We walked through the mechanics in how SynthID’s statistical signatures actually work.

That design has an unavoidable consequence: the watermark lives in the words, so changing the words removes it — gradually, proportionally, and without any special tooling. Fix a few typos and the signal barely moves. Rewrite every other sentence and confidence drops sharply. Run the text through a paraphrasing model, or translate it into German and back, and what remains is usually indistinguishable from noise. Researchers at Maryland demonstrated paraphrase attacks against the Kirchenbauer watermark within months of its publication, and a 2023 result from Harvard and elsewhere — memorably titled “Watermarks in the Sand” — made the general case: against an attacker who can check output quality, no strong watermark survives, because the attacker can always walk the text away from the marked distribution while preserving meaning.

There’s no exotic hacking here. The removal tool is editing.

Why removal is easier than the marketing suggests

Watermark announcements tend to emphasize robustness — SynthID-Text survives “light editing,” and that’s true. But three structural facts keep the honest ceiling low.

The signal needs length. Watermark confidence accumulates over tokens. A few hundred words of unedited output detect reliably; a two-sentence excerpt often doesn’t. Anything that shortens, excerpts, or interleaves watermarked text with human writing dilutes the count.

Only cooperative models watermark. A watermark exists because the generator’s operator put it there. Open-weight models running on local hardware mark nothing, and a user who wants unmarked text can simply pick a model that doesn’t play along. The scheme covers exactly the population that isn’t trying to evade it.

Paraphrase is now free. The same ecosystem that produced watermarks produced high-quality paraphrasing models. Asking one AI to reword another’s output is a one-step, zero-skill operation — which means the practical cost of removal rounds to zero. We covered the empirical picture in whether AI watermarks survive paraphrasing and editing, and it hasn’t improved since: robustness to light edits, collapse under heavy ones.

So why do providers bother? Because the target was never the adversary. Most AI-generated text is pasted somewhere unedited, and a watermark labels that bulk case at almost no cost. The EU AI Act’s Article 50 asks providers to mark AI content in machine-readable form — it asks for coverage of the ordinary case, not cryptographic invincibility. Watermarks are a census tool being debated as if they were a lock.

The ethical fault lines

Here’s where the debate actually lives, because “possible” was never in serious doubt.

The case against removal is straightforward: watermarks are transparency infrastructure. They let platforms label synthetic content, researchers measure AI’s footprint, and readers know what they’re looking at. Deliberately stripping a mark — specifically so that someone will believe a machine’s output was human-authored — is manufacturing a false provenance claim. When the destination is an academic submission, a news outlet, a product review, or a court filing, that false claim has victims.

The case for nuance is just as real. Consider the writer who drafts with AI assistance, then substantially rewrites the result — restructuring arguments, adding their own evidence, changing most of the language. The watermark fades as a *side effect of authorship*, not as an act of concealment; disclosure-based policies at journals and universities generally treat that edited text as the writer’s own, with AI assistance declared. Consider also translation, accessibility rewriting, or an ESL writer reworking AI-polished sentences into their own voice. Mechanically, all of these “remove the watermark.” Ethically, none of them resemble running unedited output through a stripping service to cheat a policy.

That’s the honest resolution of the debate: the keystroke sequences overlap, so the technology can’t draw the line. Intent draws it. What claim are you making about the text once the mark is gone — and is that claim true? Removal in service of a true claim (“I wrote this, with AI assistance I’ve disclosed”) is editing. Removal in service of a false one (“no machine touched this”) is deception, and it stays deception no matter how easy the tooling makes it.

Legally, the picture is murkier but points the same direction. No general statute forbids editing text you generated. Exposure comes from downstream use — fraud, impersonation, academic misconduct rules, platform policies — and increasingly from provider terms of service that prohibit circumventing safety systems. The EU AI Act obligates *providers* to mark content; it does not criminalize *users* who edit it, though deployers in sensitive contexts carry disclosure duties of their own.

What this means in practice

If you’re evaluating text, the watermark lesson is the same one detection keeps teaching: absence of evidence is not evidence of authorship. Unmarked text is the overwhelming default — most models never watermark, and most marked text loses its signal in ordinary revision — so a missing watermark proves nothing, and a present one proves origin, not intent. Statistical review remains the only tool that works on arbitrary prose; try it on your own text to see a sentence-level read rather than a bare verdict, or see what each plan handles if you check documents at volume.

If you’re a writer using AI assistance, the durable posture is the boring one: edit until the text is genuinely yours, disclose where policy asks, and keep your drafts. The watermark question then answers itself — you’re not removing a mark to hide a machine; you’re revising toward a claim you can stand behind. For the wider map of marks, scores, and their failure modes, browse the rest of our writing on detection.

Frequently asked questions

Can you remove AI watermarks from text? Usually, yes — and often without trying. Text watermarks like SynthID-Text work by biasing word choices in a key-dependent pattern, so anything that changes enough words weakens the signal: heavy editing, paraphrasing, translation, or summarization. Research groups have shown repeatedly that paraphrase attacks degrade every published text watermarking scheme, and a 2023 impossibility result argues no strong watermark can survive a determined attacker with access to quality checks.

Is removing an AI watermark illegal? In most places, not in itself — there is no general law against editing text you generated. The legal exposure comes from context: the EU AI Act requires providers to mark AI content in machine-readable form, and using stripped content to defraud, impersonate, or mislead can violate existing fraud, consumer-protection, or platform rules. Terms of service are the nearer tripwire — some providers prohibit circumventing their safety systems, watermarks included.

Do AI watermarks survive normal editing? Light editing, mostly yes; heavy editing, mostly no. Schemes like SynthID-Text are designed with redundancy so fixing typos or swapping a few sentences leaves enough biased tokens to detect. But the signal is statistical, so its strength scales with how much watermarked text survives. Rewrite half the words, run it through a paraphraser, or translate it to another language and back, and detection confidence collapses toward a coin flip.

Why do AI companies bother with watermarks if they’re removable? Because watermarks aren’t built to stop determined adversaries — they’re built to label the honest majority of content cheaply. Most AI text is never edited at all, and a watermark catches that bulk case at near-zero cost, which is what transparency regulation like the EU AI Act actually asks for. Providers know the scheme fails against motivated removal; they ship it anyway because labeling ninety percent of unedited output has real value for platforms and researchers.

Is editing your own AI-assisted draft the same as removing a watermark? Mechanically it can have the same effect, which is exactly why intent is the ethical dividing line. A writer who substantially rewrites an AI draft into their own words is doing what disclosure-based policies expect — the watermark fading is a side effect of genuine authorship. Someone running unedited output through a stripping tool to pass it off as human-written is manufacturing deception. The keystrokes differ, and so does the honesty of the claim being made afterward.

The bottom line

The watermark removal debate is two debates wearing one name. The technical one is settled: text watermarks are statistical signals living in word choices, and enough rewording — by editor, paraphraser, or translator — removes them, a limit the impossibility results say is permanent. The ethical one can’t be settled by technology at all, because the same edits serve honest authorship and deliberate deception alike. What separates them is the claim you make when you’re done. Watermarks will keep labeling the unedited majority of AI text, and that’s genuinely useful; but for edited prose, the question was never really “is there a mark?” It was always “is the story you’re telling about this text true?”

Try it on your own text

Paste your draft into PaperBleach to humanize AI text so it reads naturally — then check your score against built-in AI detection. Free on your first run.