One Prompt Is Enough: Watermark Laundering Through Foundation Image Models

TL;DR

One prompt enables watermark laundering through foundation image models, significantly reducing watermark reliability.

cs.CV 🔴 Advanced 2026-09-01 6 views
Jidong Yang Qi Li Wei Zong Yang-Wai Chow Willy Susilo Huaike Yu Chunpeng Wang Suo Gao
watermark laundering foundation image models black-box attacks visual fidelity semantic preservation

Key Findings

Methodology

The study introduces a method for watermark laundering using foundation image models. Attackers input a watermarked image with a single reconstruction prompt, and the model outputs a visually faithful image where the embedded watermark becomes unreliable. Evaluations were conducted using six OpenAI and Google image editing models, three watermark schemes, and 1,800 reconstructed outputs.

Key Results

  • OpenAI models demonstrated the strongest attack capability across all evaluated schemes, with DwtDct remaining vulnerable under high-fidelity reconstruction, BER approaching 0.5.
  • Nano Banana 2 showed DwtDct's vulnerability with a PSNR of 30.27 under high-fidelity reconstruction.
  • Prompt ablations indicated explicit removal instructions are unnecessary; the reconstruction pathway is the primary factor.

Significance

This research highlights the importance of foundation-model reconstruction as a missing robustness condition in invisible watermark evaluation. The single-prompt watermark laundering challenges traditional watermark attack methods, emphasizing the need to consider foundation-model reconstruction in watermark schemes.

Technical Contribution

The study is the first to utilize foundation image models' reconstruction capabilities for watermark laundering, proposing a new operational attack interface and demonstrating the vulnerability of watermark schemes under high-fidelity reconstruction. The method's effectiveness is validated through a joint BER and visual fidelity evaluation framework.

Novelty

This study is the first to propose watermark laundering using foundation image models, differing from prior regeneration attacks that require access to generative pipelines, revealing prompt-conditioned reconstruction as a unique operational attack interface.

Limitations

  • The method may be less effective in schemes sensitive to high-frequency residuals, as carriers and decoders differ across schemes.
  • Formal uncertainty testing was not conducted, potentially affecting the generality of the results.

Future Work

Future research could explore the impact of different foundation models' reconstruction capabilities on watermark laundering and how to enhance watermark schemes' robustness against such attacks.

AI Executive Summary

Invisible watermarking is crucial for image provenance, copyright protection, and AI-generated content identification. However, with the rise of foundation image models, traditional watermark evaluation methods face challenges. Researchers have proposed a novel method for watermark laundering using foundation image models, where attackers only need a single reconstruction prompt to make embedded watermark information unreliable. Experiments showed that OpenAI models demonstrated the strongest attack capability across all evaluated schemes, while Nano Banana 2 revealed DwtDct's vulnerability under high-fidelity reconstruction. Prompt ablations further confirmed that explicit removal instructions are unnecessary, with the reconstruction pathway being the primary factor. This finding emphasizes the importance of foundation-model reconstruction in invisible watermark evaluation and provides new insights for future watermark scheme design. Although the method may be less effective in schemes sensitive to high-frequency residuals, the potential threat posed by foundation-model reconstruction capabilities warrants further investigation. Future research could explore the impact of different foundation models' reconstruction capabilities on watermark laundering and how to enhance watermark schemes' robustness against such attacks.

Deep Analysis

Background

Invisible watermarking is essential for image provenance, copyright protection, and AI-generated content identification. Traditional watermark evaluation methods typically rely on fixed compression, filtering, geometric transformation, or denoising attacks. However, with the rise of foundation image models, the effectiveness of these methods is challenged. Researchers realized that attackers could use the reconstruction capabilities of foundation image models to achieve watermark laundering with a single prompt, making embedded watermark information unreliable.

Core Problem

The core problem is how to utilize the reconstruction capabilities of foundation image models for watermark laundering. Traditional watermark attack methods often require access to generative pipelines, while the black-box nature of foundation image models makes them a new operational attack interface. The challenge is to evaluate the effectiveness of this novel attack method and reveal its potential threat to watermark schemes.

Innovation

The core innovation of this study is the first use of foundation image models' reconstruction capabilities for watermark laundering. Unlike prior regeneration attacks that require access to generative pipelines, this method only needs a natural language prompt to effectively attack watermark information. This finding reveals prompt-conditioned reconstruction as a unique operational attack interface and emphasizes the need to consider foundation-model reconstruction in watermark schemes.

Methodology

  • �� Evaluated using six OpenAI and Google image editing models
  • �� Selected three representative watermark schemes: DwtDct, DwtDctSvd, and RivaGAN
  • �� Conducted experiments with 1,800 reconstructed outputs
  • �� Used a joint BER and visual fidelity evaluation framework
  • �� Conducted prompt ablation experiments to verify the impact of the reconstruction pathway

Experiments

The experimental design includes evaluations using six foundation image models and three watermark schemes. Researchers conducted experiments with 1,800 reconstructed outputs using a joint BER and visual fidelity evaluation framework. The experiments also included prompt ablation experiments to verify the impact of explicit removal instructions on watermark laundering.

Results

The experimental results showed that OpenAI models demonstrated the strongest attack capability across all evaluated schemes, while Nano Banana 2 revealed DwtDct's vulnerability under high-fidelity reconstruction. Prompt ablations further confirmed that explicit removal instructions are unnecessary, with the reconstruction pathway being the primary factor.

Applications

The study's application scenarios include image provenance verification, copyright protection, and AI-generated content identification. By revealing the potential threat of foundation image models' reconstruction capabilities to watermark schemes, the research provides new insights for future watermark scheme design.

Limitations & Outlook

The method may be less effective in schemes sensitive to high-frequency residuals, as carriers and decoders differ across schemes. Formal uncertainty testing was not conducted, potentially affecting the generality of the results.

Plain Language Accessible to non-experts

Imagine you have a painting with an invisible signature that proves it's your work. Now, someone uses a special camera to take a picture of the painting and magically recreates it. Although the painting looks identical, the invisible signature is gone. This is what the study does: using the reconstruction capabilities of foundation image models, a simple prompt can make the embedded watermark information unreliable. It's like washing away the signature with a new method while the painting itself remains unchanged.

ELI14 Explained like you're 14

Imagine you're playing a game with a hidden mark that proves you're the game's owner. Now, someone uses a special tool to remake the game. Although the game looks the same, the hidden mark is gone. That's the core of this study: using the reconstruction capabilities of foundation image models, a simple prompt can make the embedded watermark information unreliable. It's like washing away the mark with a new method while the game itself remains unchanged. Isn't that cool?

Glossary

Invisible Watermark

A mark embedded in an image to protect copyright or verify provenance, invisible to the naked eye.

Used for image provenance verification and copyright protection.

Foundation Image Model

A deep learning model capable of image editing and reconstruction.

Used to achieve watermark laundering through reconstruction capabilities.

Watermark Laundering

The process of making embedded watermark information unreliable through image reconstruction.

The core problem of the study.

BER (Bit Error Rate)

A metric measuring the accuracy of watermark information recovery; values closer to 0.5 indicate stronger attack effects.

Used to evaluate the effectiveness of watermark laundering.

High-Frequency Residual

The amount of change in high-frequency information in an image, used to analyze the vulnerability of watermark schemes.

Used to analyze the effectiveness of watermark laundering.

Open Questions Unanswered questions from this research

  • 1 How to enhance the robustness of watermark schemes against reconstruction attacks by foundation image models?
  • 2 What impact do different foundation models' reconstruction capabilities have on watermark laundering effectiveness?

Applications

Immediate Applications

Image Provenance Verification

By revealing the potential threat of foundation image models' reconstruction capabilities to watermark schemes, the research provides new insights for image provenance verification.

Long-term Vision

Copyright Protection

The study's findings could have a profound impact on future copyright protection technologies, driving the design of more robust watermark schemes.

Abstract

Invisible watermarks are typically evaluated against predefined perturbations such as compression, blur, noise, cropping, and denoising. Public foundation image models expose a distinct threat: an attacker can submit a watermarked image with a single reconstruction prompt and obtain a visually faithful output from which the invisible watermark can no longer be decoded reliably. We formalize this failure mode as watermark laundering and evaluate it using a joint payload-fidelity profile that combines bit error rate (BER) with visual and semantic preservation. Across six OpenAI and Google image editing models, three representative watermarking schemes, and 1,800 reconstructed outputs, we identify two complementary laundering regimes: OpenAI models produce the strongest payload disruption across the evaluated schemes, whereas Nano Banana 2 shows that DwtDct remains vulnerable under high-fidelity reconstruction. Prompt ablations show that no single removal-oriented instruction is necessary for payload disruption, indicating that the effect is primarily induced by the reconstruction pathway rather than by explicit attack wording. Comparisons with conventional attacks further show that prompt-conditioned reconstruction constitutes a distinct operational attack interface. These findings motivate foundation-model reconstruction as a missing robustness condition in invisible watermark evaluation.

cs.CV cs.AI cs.CR