Learning Counterfactually Decoupled Attention for Open-World Model Attribution

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zheng, Yu, Gong, Boyang, Kong, Fanye, Duan, Yueqi, Yu, Bingyao, Zheng, Wenzhao, Chen, Lei, Lu, Jiwen, Zhou, Jie
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908427593187328
author Zheng, Yu
Gong, Boyang
Kong, Fanye
Duan, Yueqi
Yu, Bingyao
Zheng, Wenzhao
Chen, Lei
Lu, Jiwen
Zhou, Jie
author_facet Zheng, Yu
Gong, Boyang
Kong, Fanye
Duan, Yueqi
Yu, Bingyao
Zheng, Wenzhao
Chen, Lei
Lu, Jiwen
Zhou, Jie
contents In this paper, we propose a Counterfactually Decoupled Attention Learning (CDAL) method for open-world model attribution. Existing methods rely on handcrafted design of region partitioning or feature space, which could be confounded by the spurious statistical correlations and struggle with novel attacks in open-world scenarios. To address this, CDAL explicitly models the causal relationships between the attentional visual traces and source model attribution, and counterfactually decouples the discriminative model-specific artifacts from confounding source biases for comparison. In this way, the resulting causal effect provides a quantification on the quality of learned attention maps, thus encouraging the network to capture essential generation patterns that generalize to unseen source models by maximizing the effect. Extensive experiments on existing open-world model attribution benchmarks show that with minimal computational overhead, our method consistently improves state-of-the-art models by large margins, particularly for unseen novel attacks. Source code: https://github.com/yzheng97/CDAL.
format Preprint
id arxiv_https___arxiv_org_abs_2506_23074
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning Counterfactually Decoupled Attention for Open-World Model Attribution
Zheng, Yu
Gong, Boyang
Kong, Fanye
Duan, Yueqi
Yu, Bingyao
Zheng, Wenzhao
Chen, Lei
Lu, Jiwen
Zhou, Jie
Computer Vision and Pattern Recognition
Cryptography and Security
Machine Learning
In this paper, we propose a Counterfactually Decoupled Attention Learning (CDAL) method for open-world model attribution. Existing methods rely on handcrafted design of region partitioning or feature space, which could be confounded by the spurious statistical correlations and struggle with novel attacks in open-world scenarios. To address this, CDAL explicitly models the causal relationships between the attentional visual traces and source model attribution, and counterfactually decouples the discriminative model-specific artifacts from confounding source biases for comparison. In this way, the resulting causal effect provides a quantification on the quality of learned attention maps, thus encouraging the network to capture essential generation patterns that generalize to unseen source models by maximizing the effect. Extensive experiments on existing open-world model attribution benchmarks show that with minimal computational overhead, our method consistently improves state-of-the-art models by large margins, particularly for unseen novel attacks. Source code: https://github.com/yzheng97/CDAL.
title Learning Counterfactually Decoupled Attention for Open-World Model Attribution
topic Computer Vision and Pattern Recognition
Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2506.23074