Stego Battlefield: Evaluating Image Steganography Attacks and Steganalysis Defenses

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Sun, Zhen, Zhang, Zongmin, Sheng, Leyi, Liu, Yule, Liao, Yifan, Li, Ke, Zheng, Xinhu, Wei, Jiaheng, Yang, Wenyuan, He, Xinlei
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917467001978880
author Sun, Zhen
Zhang, Zongmin
Sheng, Leyi
Liu, Yule
Liao, Yifan
Li, Ke
Zheng, Xinhu
Wei, Jiaheng
Yang, Wenyuan
He, Xinlei
author_facet Sun, Zhen
Zhang, Zongmin
Sheng, Leyi
Liu, Yule
Liao, Yifan
Li, Ke
Zheng, Xinhu
Wei, Jiaheng
Yang, Wenyuan
He, Xinlei
contents Image steganography is widely used to protect user privacy and enable covert communication. However, it can also be abused by the adversary as a covert channel to bypass content moderation, disseminate harmful semantics, and even hide malicious instructions in images to elicit dangerous outputs from large models, posing a practical security risk that continues to evolve. To address the lack of a unified and systematic evaluation framework, we propose SADBench, a systematic benchmark that assesses the adversary's ability to inject harmful secrets via steganography and the defender's ability to detect such threats through steganalysis. Crucially, SADBench comprises $4$ core tasks, namely steganography attack capability evaluation, steganalysis defense capability evaluation, efficiency evaluation, and transferability evaluation. It evaluates both image-payload and text-payload steganography across diverse cover distributions, utilizing harmful visual semantics and toxic instructions to simulate malicious attacks. Across a broad set of attacks and detectors, SADBench reveals that (i) INN and autoencoder-based methods demonstrate superior stability compared to other architectures, (ii) in-domain detection is near-perfect and cheaper than generation, (iii) a critical asymmetry exists in transferability where attacks robustly generalize to new distributions while detectors fail to adapt, and (iv) real-world threats persist on social media, where payloads either survive minimal compression or effectively adapt to aggressive compression via simulated training. Overall, SADBench establishes a systematic, reproducible, and extensible framework to quantify risks, paving the way for measurable and security-driven advancements in steganography defense.
format Preprint
id arxiv_https___arxiv_org_abs_2605_05789
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Stego Battlefield: Evaluating Image Steganography Attacks and Steganalysis Defenses
Sun, Zhen
Zhang, Zongmin
Sheng, Leyi
Liu, Yule
Liao, Yifan
Li, Ke
Zheng, Xinhu
Wei, Jiaheng
Yang, Wenyuan
He, Xinlei
Cryptography and Security
Computer Vision and Pattern Recognition
Image steganography is widely used to protect user privacy and enable covert communication. However, it can also be abused by the adversary as a covert channel to bypass content moderation, disseminate harmful semantics, and even hide malicious instructions in images to elicit dangerous outputs from large models, posing a practical security risk that continues to evolve. To address the lack of a unified and systematic evaluation framework, we propose SADBench, a systematic benchmark that assesses the adversary's ability to inject harmful secrets via steganography and the defender's ability to detect such threats through steganalysis. Crucially, SADBench comprises $4$ core tasks, namely steganography attack capability evaluation, steganalysis defense capability evaluation, efficiency evaluation, and transferability evaluation. It evaluates both image-payload and text-payload steganography across diverse cover distributions, utilizing harmful visual semantics and toxic instructions to simulate malicious attacks. Across a broad set of attacks and detectors, SADBench reveals that (i) INN and autoencoder-based methods demonstrate superior stability compared to other architectures, (ii) in-domain detection is near-perfect and cheaper than generation, (iii) a critical asymmetry exists in transferability where attacks robustly generalize to new distributions while detectors fail to adapt, and (iv) real-world threats persist on social media, where payloads either survive minimal compression or effectively adapt to aggressive compression via simulated training. Overall, SADBench establishes a systematic, reproducible, and extensible framework to quantify risks, paving the way for measurable and security-driven advancements in steganography defense.
title Stego Battlefield: Evaluating Image Steganography Attacks and Steganalysis Defenses
topic Cryptography and Security
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.05789