DPL: Decoupled Prototype Learning for Enhancing Robustness of Vision-Language Transformers to Missing Modalities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Jueqing, Qi, Yuanyuan, Yang, Xiaohao, Niu, Shuaicheng, Ke, Fucai, Zhou, Shujie, Tan, Wei, Lin, Jionghao, Buntine, Wray, Rezatofighi, Hamid, Du, Lan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909904813424640
author Lu, Jueqing
Qi, Yuanyuan
Yang, Xiaohao
Niu, Shuaicheng
Ke, Fucai
Zhou, Shujie
Tan, Wei
Lin, Jionghao
Buntine, Wray
Rezatofighi, Hamid
Du, Lan
author_facet Lu, Jueqing
Qi, Yuanyuan
Yang, Xiaohao
Niu, Shuaicheng
Ke, Fucai
Zhou, Shujie
Tan, Wei
Lin, Jionghao
Buntine, Wray
Rezatofighi, Hamid
Du, Lan
contents The performance of Visio-Language Transformers drops sharply when an input modality (e.g., image) is missing, because the model is forced to make predictions using incomplete information. Existing missing-aware prompt methods help reduce this degradation, but they still rely on conventional prediction heads (e.g., a Fully-Connected layer) that compute class scores in the same way regardless of which modality is present or absent. We introduce Decoupled Prototype Learning (DPL), a new prediction head architecture that explicitly adjusts its decision process to the observed input modalities. For each class, DPL selects a set of prototypes specific to the current missing-modality cases (image-missing, text-missing, or mixed-missing). Each prototype is then decomposed into image-specific and text-specific components, enabling the head to make decisions that depend on the information actually present. This adaptive design allows DPL to handle inputs with missing modalities more effectively while remaining fully compatible with existing prompt-based frameworks. Extensive experiments on MM-IMDb, UPMC Food-101, and Hateful Memes demonstrate that DPL outperforms state-of-the-art approaches across all widely used multimodal imag-text datasets and various missing cases.
format Preprint
id arxiv_https___arxiv_org_abs_2505_08283
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DPL: Decoupled Prototype Learning for Enhancing Robustness of Vision-Language Transformers to Missing Modalities
Lu, Jueqing
Qi, Yuanyuan
Yang, Xiaohao
Niu, Shuaicheng
Ke, Fucai
Zhou, Shujie
Tan, Wei
Lin, Jionghao
Buntine, Wray
Rezatofighi, Hamid
Du, Lan
Machine Learning
Computer Vision and Pattern Recognition
The performance of Visio-Language Transformers drops sharply when an input modality (e.g., image) is missing, because the model is forced to make predictions using incomplete information. Existing missing-aware prompt methods help reduce this degradation, but they still rely on conventional prediction heads (e.g., a Fully-Connected layer) that compute class scores in the same way regardless of which modality is present or absent. We introduce Decoupled Prototype Learning (DPL), a new prediction head architecture that explicitly adjusts its decision process to the observed input modalities. For each class, DPL selects a set of prototypes specific to the current missing-modality cases (image-missing, text-missing, or mixed-missing). Each prototype is then decomposed into image-specific and text-specific components, enabling the head to make decisions that depend on the information actually present. This adaptive design allows DPL to handle inputs with missing modalities more effectively while remaining fully compatible with existing prompt-based frameworks. Extensive experiments on MM-IMDb, UPMC Food-101, and Hateful Memes demonstrate that DPL outperforms state-of-the-art approaches across all widely used multimodal imag-text datasets and various missing cases.
title DPL: Decoupled Prototype Learning for Enhancing Robustness of Vision-Language Transformers to Missing Modalities
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.08283