Toward Accountable AI-Generated Content on Social Platforms: Steganographic Attribution and Multimodal Harm Detection

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Guan, Xinlei, Arosemena, David, Dhandu, Tejaswi, Huang, Kuan, Xu, Meng, Li, Miles Q., Shen, Bingyu, Qin, Ruiyang, Tida, Umamaheswara Rao, Li, Boyang
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914466441986048
author Guan, Xinlei
Arosemena, David
Dhandu, Tejaswi
Huang, Kuan
Xu, Meng
Li, Miles Q.
Shen, Bingyu
Qin, Ruiyang
Tida, Umamaheswara Rao
Li, Boyang
author_facet Guan, Xinlei
Arosemena, David
Dhandu, Tejaswi
Huang, Kuan
Xu, Meng
Li, Miles Q.
Shen, Bingyu
Qin, Ruiyang
Tida, Umamaheswara Rao
Li, Boyang
contents The rapid growth of generative AI has introduced new challenges in content moderation and digital forensics. In particular, benign AI-generated images can be paired with harmful or misleading text, creating difficult-to-detect misuse. This contextual misuse undermines the traditional moderation framework and complicates attribution, as synthetic images typically lack persistent metadata or device signatures. We introduce a steganography enabled attribution framework that embeds cryptographically signed identifiers into images at creation time and uses multimodal harmful content detection as a trigger for attribution verification. Our system evaluates five watermarking methods across spatial, frequency, and wavelet domains. It also integrates a CLIP-based fusion model for multimodal harmful-content detection. Experiments demonstrate that spread-spectrum watermarking, especially in the wavelet domain, provides strong robustness under blur distortions, and our multimodal fusion detector achieves an AUC-ROC of 0.99, enabling reliable cross-modal attribution verification. These components form an end-to-end forensic pipeline that enables reliable tracing of harmful deployments of AI-generated imagery, supporting accountability in modern synthetic media environments. Our code is available at GitHub: https://github.com/bli1/steganography
format Preprint
id arxiv_https___arxiv_org_abs_2604_10460
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Toward Accountable AI-Generated Content on Social Platforms: Steganographic Attribution and Multimodal Harm Detection
Guan, Xinlei
Arosemena, David
Dhandu, Tejaswi
Huang, Kuan
Xu, Meng
Li, Miles Q.
Shen, Bingyu
Qin, Ruiyang
Tida, Umamaheswara Rao
Li, Boyang
Computer Vision and Pattern Recognition
Artificial Intelligence
Cryptography and Security
Emerging Technologies
The rapid growth of generative AI has introduced new challenges in content moderation and digital forensics. In particular, benign AI-generated images can be paired with harmful or misleading text, creating difficult-to-detect misuse. This contextual misuse undermines the traditional moderation framework and complicates attribution, as synthetic images typically lack persistent metadata or device signatures. We introduce a steganography enabled attribution framework that embeds cryptographically signed identifiers into images at creation time and uses multimodal harmful content detection as a trigger for attribution verification. Our system evaluates five watermarking methods across spatial, frequency, and wavelet domains. It also integrates a CLIP-based fusion model for multimodal harmful-content detection. Experiments demonstrate that spread-spectrum watermarking, especially in the wavelet domain, provides strong robustness under blur distortions, and our multimodal fusion detector achieves an AUC-ROC of 0.99, enabling reliable cross-modal attribution verification. These components form an end-to-end forensic pipeline that enables reliable tracing of harmful deployments of AI-generated imagery, supporting accountability in modern synthetic media environments. Our code is available at GitHub: https://github.com/bli1/steganography
title Toward Accountable AI-Generated Content on Social Platforms: Steganographic Attribution and Multimodal Harm Detection
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Cryptography and Security
Emerging Technologies
url https://arxiv.org/abs/2604.10460