Backdoor Cleaning without External Guidance in MLLM Fine-tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rong, Xuankun, Huang, Wenke, Liang, Jian, Bi, Jinhe, Xiao, Xun, Li, Yiming, Du, Bo, Ye, Mang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915298873966592
author Rong, Xuankun
Huang, Wenke
Liang, Jian
Bi, Jinhe
Xiao, Xun
Li, Yiming
Du, Bo
Ye, Mang
author_facet Rong, Xuankun
Huang, Wenke
Liang, Jian
Bi, Jinhe
Xiao, Xun
Li, Yiming
Du, Bo
Ye, Mang
contents Multimodal Large Language Models (MLLMs) are increasingly deployed in fine-tuning-as-a-service (FTaaS) settings, where user-submitted datasets adapt general-purpose models to downstream tasks. This flexibility, however, introduces serious security risks, as malicious fine-tuning can implant backdoors into MLLMs with minimal effort. In this paper, we observe that backdoor triggers systematically disrupt cross-modal processing by causing abnormal attention concentration on non-semantic regions--a phenomenon we term attention collapse. Based on this insight, we propose Believe Your Eyes (BYE), a data filtering framework that leverages attention entropy patterns as self-supervised signals to identify and filter backdoor samples. BYE operates via a three-stage pipeline: (1) extracting attention maps using the fine-tuned model, (2) computing entropy scores and profiling sensitive layers via bimodal separation, and (3) performing unsupervised clustering to remove suspicious samples. Unlike prior defenses, BYE equires no clean supervision, auxiliary labels, or model modifications. Extensive experiments across various datasets, models, and diverse trigger types validate BYE's effectiveness: it achieves near-zero attack success rates while maintaining clean-task performance, offering a robust and generalizable solution against backdoor threats in MLLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16916
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Backdoor Cleaning without External Guidance in MLLM Fine-tuning
Rong, Xuankun
Huang, Wenke
Liang, Jian
Bi, Jinhe
Xiao, Xun
Li, Yiming
Du, Bo
Ye, Mang
Cryptography and Security
Computer Vision and Pattern Recognition
Multimodal Large Language Models (MLLMs) are increasingly deployed in fine-tuning-as-a-service (FTaaS) settings, where user-submitted datasets adapt general-purpose models to downstream tasks. This flexibility, however, introduces serious security risks, as malicious fine-tuning can implant backdoors into MLLMs with minimal effort. In this paper, we observe that backdoor triggers systematically disrupt cross-modal processing by causing abnormal attention concentration on non-semantic regions--a phenomenon we term attention collapse. Based on this insight, we propose Believe Your Eyes (BYE), a data filtering framework that leverages attention entropy patterns as self-supervised signals to identify and filter backdoor samples. BYE operates via a three-stage pipeline: (1) extracting attention maps using the fine-tuned model, (2) computing entropy scores and profiling sensitive layers via bimodal separation, and (3) performing unsupervised clustering to remove suspicious samples. Unlike prior defenses, BYE equires no clean supervision, auxiliary labels, or model modifications. Extensive experiments across various datasets, models, and diverse trigger types validate BYE's effectiveness: it achieves near-zero attack success rates while maintaining clean-task performance, offering a robust and generalizable solution against backdoor threats in MLLMs.
title Backdoor Cleaning without External Guidance in MLLM Fine-tuning
topic Cryptography and Security
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.16916