首页 > AI前沿 > RMCW: A Deletion-Robust Watermark Based on Reed--Muller Codes for Language Models

RMCW: A Deletion-Robust Watermark Based on Reed--Muller Codes for Language Models

arXiv机器学习 2026-10-02 13:05 5 阅读 查看原文

Large Language Model (LLM) watermarking provides a lightweight mechanism for identifying text generated by a specific model, but its robustness remains fragile under post-processing attacks.

Deletion attacks are particularly challenging because they shift token positions and break the alignment between observed tokens and their original watermark positions.

We propose Reed--Muller Code Watermarking (RMCW)

In contrast to global codeword recovery, RMCW searches for surviving local algebraic structure, leveraging the Reed--Solomon consistency induced by affine-line restrictions of Reed--Muller codewords.

During generation, RMCW injects a Reed--Muller structure into the sequence via a secret-keyed vocabulary partition.

During detection, it maps the given text to keyed vocabulary bins and tests local subsequences for low-degree Reed--Solomon consistency using Berlekamp--Welch tests.

Experiments

Experiments on C4 and ELI5 datasets with OPT-1.3B and Llama-3.1-8B-Instruct show that RMCW preserves strong clean-text detectability and outperforms or matches the baseline methods under several deletion and rewriting attacks.

Availability

Our code is available at https://github.com/BaichengDanny/RMCW.