EntropyScan: Towards Model-level Backdoor Detection in LVLMs via Visual Attention Entropy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ge, Xuanyu, Wang, Zhongqi, Zhang, Jie, Shan, Shiguang, Chen, Xilin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918503376748544
author Ge, Xuanyu
Wang, Zhongqi
Zhang, Jie
Shan, Shiguang
Chen, Xilin
author_facet Ge, Xuanyu
Wang, Zhongqi
Zhang, Jie
Shan, Shiguang
Chen, Xilin
contents Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities across various tasks, yet they remain vulnerable to backdoor attacks. Existing defense methods predominantly focus on sample-level defense, which relies on the knowledge of training data or triggers. However, identifying whether a given model is backdoored remains a critical but unexplored task. To fill this gap, we propose EntropyScan, a lightweight and trigger-agnostic method for model-level backdoor detection in LVLMs. We first observe that backdoor injection disrupts the cross-modal alignment, resulting in pronounced structural anomalies in visual attention allocation on benign samples. Based on this insight, EntropyScan detects the backdoor models by quantifying such attention deviations. Specifically, it extracts visual attention distributions from the initial layers of the Large Language Model (LLM) and applies Tsallis entropy to capture these structural distortions. By employing a reference-anchored Z-score normalization on a small set of benign samples, it effectively identifies the backdoored model. Extensive experiments across two LVLMs architectures and three advanced attack scenarios show that EntropyScan achieves an F1 score of 98.5% in average and an AUC of 96.6%. Our code will be publicly available soon.
format Preprint
id arxiv_https___arxiv_org_abs_2605_15711
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EntropyScan: Towards Model-level Backdoor Detection in LVLMs via Visual Attention Entropy
Ge, Xuanyu
Wang, Zhongqi
Zhang, Jie
Shan, Shiguang
Chen, Xilin
Computer Vision and Pattern Recognition
Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities across various tasks, yet they remain vulnerable to backdoor attacks. Existing defense methods predominantly focus on sample-level defense, which relies on the knowledge of training data or triggers. However, identifying whether a given model is backdoored remains a critical but unexplored task. To fill this gap, we propose EntropyScan, a lightweight and trigger-agnostic method for model-level backdoor detection in LVLMs. We first observe that backdoor injection disrupts the cross-modal alignment, resulting in pronounced structural anomalies in visual attention allocation on benign samples. Based on this insight, EntropyScan detects the backdoor models by quantifying such attention deviations. Specifically, it extracts visual attention distributions from the initial layers of the Large Language Model (LLM) and applies Tsallis entropy to capture these structural distortions. By employing a reference-anchored Z-score normalization on a small set of benign samples, it effectively identifies the backdoored model. Extensive experiments across two LVLMs architectures and three advanced attack scenarios show that EntropyScan achieves an F1 score of 98.5% in average and an AUC of 96.6%. Our code will be publicly available soon.
title EntropyScan: Towards Model-level Backdoor Detection in LVLMs via Visual Attention Entropy
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.15711