BlackMirror: Black-Box Backdoor Detection for Text-to-Image Models via Instruction-Response Deviation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Feiran, Xu, Qianqian, Bao, Shilong, Yang, Zhiyong, Zhao, Xilin, Cao, Xiaochun, Huang, Qingming
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915839167430656
author Li, Feiran
Xu, Qianqian
Bao, Shilong
Yang, Zhiyong
Zhao, Xilin
Cao, Xiaochun
Huang, Qingming
author_facet Li, Feiran
Xu, Qianqian
Bao, Shilong
Yang, Zhiyong
Zhao, Xilin
Cao, Xiaochun
Huang, Qingming
contents This paper investigates the challenging task of detecting backdoored text-to-image models under black-box settings and introduces a novel detection framework BlackMirror. Existing approaches typically rely on analyzing image-level similarity, under the assumption that backdoor-triggered generations exhibit strong consistency across samples. However, they struggle to generalize to recently emerging backdoor attacks, where backdoored generations can appear visually diverse. BlackMirror is motivated by an observation: across backdoor attacks, {only partial semantic patterns within the generated image are steadily manipulated, while the rest of the content remains diverse or benign. Accordingly, BlackMirror consists of two components: MirrorMatch, which aligns visual patterns with the corresponding instructions to detect semantic deviations; and MirrorVerify, which evaluates the stability of these deviations across varied prompts to distinguish true backdoor behavior from benign responses. BlackMirror is a general, training-free framework that can be deployed as a plug-and-play module in Model-as-a-Service (MaaS) applications. Comprehensive experiments demonstrate that BlackMirror achieves accurate detection across a wide range of attacks. Code is available at https://github.com/Ferry-Li/BlackMirror.
format Preprint
id arxiv_https___arxiv_org_abs_2603_05921
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle BlackMirror: Black-Box Backdoor Detection for Text-to-Image Models via Instruction-Response Deviation
Li, Feiran
Xu, Qianqian
Bao, Shilong
Yang, Zhiyong
Zhao, Xilin
Cao, Xiaochun
Huang, Qingming
Computer Vision and Pattern Recognition
Artificial Intelligence
This paper investigates the challenging task of detecting backdoored text-to-image models under black-box settings and introduces a novel detection framework BlackMirror. Existing approaches typically rely on analyzing image-level similarity, under the assumption that backdoor-triggered generations exhibit strong consistency across samples. However, they struggle to generalize to recently emerging backdoor attacks, where backdoored generations can appear visually diverse. BlackMirror is motivated by an observation: across backdoor attacks, {only partial semantic patterns within the generated image are steadily manipulated, while the rest of the content remains diverse or benign. Accordingly, BlackMirror consists of two components: MirrorMatch, which aligns visual patterns with the corresponding instructions to detect semantic deviations; and MirrorVerify, which evaluates the stability of these deviations across varied prompts to distinguish true backdoor behavior from benign responses. BlackMirror is a general, training-free framework that can be deployed as a plug-and-play module in Model-as-a-Service (MaaS) applications. Comprehensive experiments demonstrate that BlackMirror achieves accurate detection across a wide range of attacks. Code is available at https://github.com/Ferry-Li/BlackMirror.
title BlackMirror: Black-Box Backdoor Detection for Text-to-Image Models via Instruction-Response Deviation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2603.05921