EPE-P: Evidence-based Parameter-efficient Prompting for Multimodal Learning with Missing Modalities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Zhe, Lin, Xun, Cui, Yawen, Yu, Zitong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929645364969472
author Chen, Zhe
Lin, Xun
Cui, Yawen
Yu, Zitong
author_facet Chen, Zhe
Lin, Xun
Cui, Yawen
Yu, Zitong
contents Missing modalities are a common challenge in real-world multimodal learning scenarios, occurring during both training and testing. Existing methods for managing missing modalities often require the design of separate prompts for each modality or missing case, leading to complex designs and a substantial increase in the number of parameters to be learned. As the number of modalities grows, these methods become increasingly inefficient due to parameter redundancy. To address these issues, we propose Evidence-based Parameter-Efficient Prompting (EPE-P), a novel and parameter-efficient method for pretrained multimodal networks. Our approach introduces a streamlined design that integrates prompting information across different modalities, reducing complexity and mitigating redundant parameters. Furthermore, we propose an Evidence-based Loss function to better handle the uncertainty associated with missing modalities, improving the model's decision-making. Our experiments demonstrate that EPE-P outperforms existing prompting-based methods in terms of both effectiveness and efficiency. The code is released at https://github.com/Boris-Jobs/EPE-P_MLLMs-Robustness.
format Preprint
id arxiv_https___arxiv_org_abs_2412_17677
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EPE-P: Evidence-based Parameter-efficient Prompting for Multimodal Learning with Missing Modalities
Chen, Zhe
Lin, Xun
Cui, Yawen
Yu, Zitong
Computer Vision and Pattern Recognition
Missing modalities are a common challenge in real-world multimodal learning scenarios, occurring during both training and testing. Existing methods for managing missing modalities often require the design of separate prompts for each modality or missing case, leading to complex designs and a substantial increase in the number of parameters to be learned. As the number of modalities grows, these methods become increasingly inefficient due to parameter redundancy. To address these issues, we propose Evidence-based Parameter-Efficient Prompting (EPE-P), a novel and parameter-efficient method for pretrained multimodal networks. Our approach introduces a streamlined design that integrates prompting information across different modalities, reducing complexity and mitigating redundant parameters. Furthermore, we propose an Evidence-based Loss function to better handle the uncertainty associated with missing modalities, improving the model's decision-making. Our experiments demonstrate that EPE-P outperforms existing prompting-based methods in terms of both effectiveness and efficiency. The code is released at https://github.com/Boris-Jobs/EPE-P_MLLMs-Robustness.
title EPE-P: Evidence-based Parameter-efficient Prompting for Multimodal Learning with Missing Modalities
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.17677