Invisible watermarking has become the default answer to a genuinely hard problem: how do you establish provenance for AI-generated images at scale, across distribution channels you do not control? Google, Meta, OpenAI, and Amazon have all integrated watermarking into their generative pipelines, and enterprise compliance teams have started treating it as an integrity mechanism. The problem is that the research community has been quietly demonstrating that the mechanism is far weaker than the compliance framing implies, and the weakness is not a tuning problem. It is structural.
What Residual Transferability Actually Means for Provenance
Neural image watermarks work by embedding an imperceptible perturbation into an image during generation. A decoder later reads that perturbation and confirms the image's origin. The assumption baked into most enterprise deployments is that the perturbation is meaningfully tied to the image it was embedded in.
Residual-based forgery attacks break that assumption. An attacker extracts the watermark-bearing perturbation from a legitimately watermarked image, then transfers it to entirely unrelated content. The decoder reads the transferred perturbation and confirms a false provenance. The forged image passes as authenticated.
Dong et al. (Dong et al., HuggingFace 2026) formalise this as residual transferability (RT): a metric that quantifies how well watermark evidence remains decodable after transfer to a different image. Their findings show that RT varies substantially across watermarking systems, and that variation is not explained by training differences. Architecture is the determining factor.
Why Architecture Determines Forgery Resistance
The intuition here is important. A watermark residual that is strongly coupled to its original cover image will degrade when transferred, because the spatial and semantic context it depends on no longer exists in the target image. A residual that is weakly coupled to its cover image behaves more like a portable signal, which is precisely what an attacker needs.
Dong et al. identify two architectural mechanisms that strengthen cover-image dependence and suppress transferability. Systems that lack these mechanisms produce residuals that travel well across images, making forgery straightforward and requiring only one or a few watermarked samples to execute at scale.
The commercial implication is direct. If your watermarking vendor cannot explain how their architecture suppresses residual transferability, you do not have enough information to assess the forgery risk in your deployment. The question is not whether the decoder achieves high accuracy on clean images. The question is what happens when the residual is extracted and replanted.
The Gap Between Robustness and Security
Most watermarking systems are evaluated on robustness: can the watermark survive JPEG compression, resizing, colour adjustments, and other common image transformations? This is a legitimate engineering concern, but it measures the wrong threat model for provenance use cases.
Security, in this context, means resistance to adversarial forgery. A system can be highly robust against incidental distortions and simultaneously trivial to forge, because the same property that makes the residual durable under compression (low cover-image dependence) makes it transferable under attack. Robustness and security are in tension, not alignment.
This distinction matters for procurement and audit. When a vendor presents robustness benchmarks as evidence of security, they are answering a different question than the one your threat model requires. Engineering leaders evaluating watermarking tooling should request adversarial forgery evaluations specifically, not robustness curves.
CoverLock and the Limits of Retrofit Defences
For teams already running production systems where architectural redesign is not feasible, Dong et al. introduce CoverLock, a plug-and-play strategy that strengthens cover-image dependence without changing the underlying watermarking architecture. Across high-RT systems, CoverLock achieves a more favourable security-to-robustness trade-off than both traditional handcrafted defences and learned classifier-based defences.
This is useful, but the framing matters. CoverLock is a mitigation for a structural vulnerability, not a resolution of it. Applying it reduces RT without eliminating the underlying architectural weakness. That is an acceptable interim position for a system already in production, but it should not be treated as equivalent to building a low-RT architecture from the start.
The practical guidance here is to treat CoverLock-style defences as a risk reduction measure with a known ceiling, and to document that ceiling explicitly in your security posture. Regulators and auditors increasingly expect provenance mechanisms to be accompanied by adversarial threat assessments, not just accuracy metrics.
What CTOs Should Evaluate Before Treating Watermarking as a Provenance Guarantee
Watermarking can be a useful signal in a layered provenance strategy. It becomes a liability when it is treated as a standalone integrity control, because a forged watermark does not just fail to detect manipulation. It actively asserts false provenance, which is worse than no signal at all.
Before deploying any watermarking system as a compliance or integrity mechanism, the following questions warrant explicit answers from your vendor or internal team:
- What is the measured residual transferability of the architecture under adversarial conditions?
- Has the system been evaluated against one-shot or few-shot forgery attacks, not just incidental distortions?
- Does the architecture incorporate mechanisms that create cover-image dependence in the residual, and can those mechanisms be described concretely?
- If architectural redesign is not feasible, what plug-in defences are applied, and what is their documented security ceiling?
Watermarking is a provenance signal that requires an adversarial threat model to be assessed honestly. Treating it as a solved problem, before that assessment has been done, transfers the risk to the systems and decisions that depend on it downstream.
FAQs
Residual transferability (RT) measures how well the watermark signal extracted from one image remains decodable when transferred to an entirely different image. High RT means an attacker can forge provenance by transplanting a legitimate watermark onto unauthorised content. For enterprise deployments that rely on watermarking as an integrity or compliance signal, high RT means the system can be made to assert false provenance, which is a more dangerous failure mode than simply missing a detection.
Research by Dong et al. (HuggingFace 2026) shows that common training-side variations do not account for the large differences in RT across watermarking systems. Architecture is the primary determinant. Training adjustments may shift performance on robustness benchmarks without meaningfully reducing forgery risk. Addressing RT properly requires architectural changes that increase the dependence of the watermark residual on the specific cover image it was embedded in.
Robustness benchmarks measure whether a watermark survives incidental distortions like compression or resizing. They do not measure resistance to adversarial forgery attacks, which operate on a different threat model. A system can score well on robustness while remaining straightforward to forge, because low cover-image dependence aids both durability and transferability. Request adversarial forgery evaluations separately, and treat robustness figures as necessary but insufficient evidence of security.
CoverLock is a meaningful risk reduction measure for systems where architectural redesign is not practical. It strengthens cover-image dependence without requiring changes to the underlying watermarking architecture and outperforms traditional handcrafted defences on the security-to-robustness trade-off. However, it reduces rather than eliminates the architectural vulnerability. Teams applying it should document the residual risk explicitly and treat it as an interim position, not a permanent security resolution.
Regulators and auditors evaluating AI provenance mechanisms are increasingly looking for adversarial threat assessments alongside accuracy metrics. If your watermarking system has not been evaluated against forgery attacks and that gap is later identified during an audit, the compliance value of the control is undermined. More concretely, a system that can be made to assert false provenance may create liability rather than limit it, particularly in contexts where watermark evidence is used to make attribution or enforcement decisions.

