PDMP: Rethinking Balanced Multimodal Learning via Performance-Dominant Modality Prioritization
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Shicai, Luo, Chunbo, Zhu, Qiang, Luo, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
One-stage Modality Distillation for Incomplete Multimodal Learning
by: Wei, Shicai, et al.
Published: (2023)
by: Wei, Shicai, et al.
Published: (2023)
Improving Multimodal Learning via Imbalanced Learning
by: Wei, Shicai, et al.
Published: (2025)
by: Wei, Shicai, et al.
Published: (2025)
Boosting Multimodal Learning via Disentangled Gradient Learning
by: Wei, Shicai, et al.
Published: (2025)
by: Wei, Shicai, et al.
Published: (2025)
Robust Multimodal Learning via Representation Decoupling
by: Wei, Shicai, et al.
Published: (2024)
by: Wei, Shicai, et al.
Published: (2024)
Scale Decoupled Distillation
by: Luo, Shicai Wei Chunbo Luo Yang
Published: (2024)
by: Luo, Shicai Wei Chunbo Luo Yang
Published: (2024)
Unbiased Dynamic Multimodal Fusion
by: Wei, Shicai, et al.
Published: (2026)
by: Wei, Shicai, et al.
Published: (2026)
See-Saw Modality Balance: See Gradient, and Sew Impaired Vision-Language Balance to Mitigate Dominant Modality Bias
by: Kwon, JuneHyoung, et al.
Published: (2025)
by: Kwon, JuneHyoung, et al.
Published: (2025)
Robust Multimodal Semantic Segmentation with Balanced Modality Contributions
by: Tan, Jiaqi, et al.
Published: (2025)
by: Tan, Jiaqi, et al.
Published: (2025)
Weakly Supervised Multimodal Temporal Forgery Localization via Multitask Learning
by: Xu, Wenbo, et al.
Published: (2025)
by: Xu, Wenbo, et al.
Published: (2025)
Rethinking Visual Token Reduction in LVLMs Under Cross-Modal Misalignment
by: Xu, Rui, et al.
Published: (2025)
by: Xu, Rui, et al.
Published: (2025)
Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers
by: Lv, Zhengyao, et al.
Published: (2025)
by: Lv, Zhengyao, et al.
Published: (2025)
Modality-Balanced Learning for Multimedia Recommendation
by: Zhang, Jinghao, et al.
Published: (2024)
by: Zhang, Jinghao, et al.
Published: (2024)
CrossVL: Complexity-Aware Feature Routing and Paired Curriculum for Cross-View Vision-Language Detection
by: Liu, Zhipeng, et al.
Published: (2026)
by: Liu, Zhipeng, et al.
Published: (2026)
Label What Matters: Modality-Balanced and Difficulty-Aware Multimodal Active Learning
by: Zeng, Yuqiao, et al.
Published: (2026)
by: Zeng, Yuqiao, et al.
Published: (2026)
Rethinking Multi-Modal Object Detection from the Perspective of Mono-Modality Feature Learning
by: Zhao, Tianyi, et al.
Published: (2025)
by: Zhao, Tianyi, et al.
Published: (2025)
MILES: Modality-Informed Learning Rate Scheduler for Balancing Multimodal Learning
by: Guerra-Manzanares, Alejandro, et al.
Published: (2025)
by: Guerra-Manzanares, Alejandro, et al.
Published: (2025)
GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking
by: Zhan, Yufei, et al.
Published: (2025)
by: Zhan, Yufei, et al.
Published: (2025)
Multimodal Cancer Survival Analysis via Hypergraph Learning with Cross-Modality Rebalance
by: Qu, Mingcheng, et al.
Published: (2025)
by: Qu, Mingcheng, et al.
Published: (2025)
Inference-Time Dynamic Modality Selection for Incomplete Multimodal Classification
by: Du, Siyi, et al.
Published: (2026)
by: Du, Siyi, et al.
Published: (2026)
OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering
by: Jia, Yiduo, et al.
Published: (2026)
by: Jia, Yiduo, et al.
Published: (2026)
Rethinking Multimodal Learning from the Perspective of Mitigating Classification Ability Disproportion
by: Jiang, QingYuan, et al.
Published: (2025)
by: Jiang, QingYuan, et al.
Published: (2025)
OilSAM2: Memory-Augmented SAM2 for Scalable SAR Oil Spill Detection
by: Chen, Shuaiyu, et al.
Published: (2026)
by: Chen, Shuaiyu, et al.
Published: (2026)
OSDMamba: Enhancing Oil Spill Detection from Remote Sensing Images Using Selective State Space Model
by: Chen, Shuaiyu, et al.
Published: (2025)
by: Chen, Shuaiyu, et al.
Published: (2025)
Learning to Balance: Decoupled Siamese Diffusion Transformer for Reference-Based Remote Sensing Image Super-Resolution
by: Luo, Bin, et al.
Published: (2026)
by: Luo, Bin, et al.
Published: (2026)
Asynchronous Multimodal Video Sequence Fusion via Learning Modality-Exclusive and -Agnostic Representations
by: Yang, Dingkang, et al.
Published: (2024)
by: Yang, Dingkang, et al.
Published: (2024)
Learning A Robust RGB-Thermal Detector for Extreme Modality Imbalance
by: Tian, Chao, et al.
Published: (2025)
by: Tian, Chao, et al.
Published: (2025)
CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation
by: Du, Yuxuan, et al.
Published: (2025)
by: Du, Yuxuan, et al.
Published: (2025)
Beyond Class Tokens: LLM-guided Dominant Property Mining for Few-shot Classification
by: Zhuo, Wei, et al.
Published: (2025)
by: Zhuo, Wei, et al.
Published: (2025)
Long-Tailed Out-of-Distribution Detection: Prioritizing Attention to Tail
by: He, Yina, et al.
Published: (2024)
by: He, Yina, et al.
Published: (2024)
Rethinking Multiple Instance Learning: Developing an Instance-Level Classifier via Weakly-Supervised Self-Training
by: Ma, Yingfan, et al.
Published: (2024)
by: Ma, Yingfan, et al.
Published: (2024)
AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model
by: Zheng, Dian, et al.
Published: (2025)
by: Zheng, Dian, et al.
Published: (2025)
Balanced Multi-modal Federated Learning via Cross-Modal Infiltration
by: Fan, Yunfeng, et al.
Published: (2023)
by: Fan, Yunfeng, et al.
Published: (2023)
Parameter-Efficient Modality-Balanced Symmetric Fusion for Multimodal Remote Sensing Semantic Segmentation
by: Li, Haocheng, et al.
Published: (2026)
by: Li, Haocheng, et al.
Published: (2026)
Improving Multimodal Sentiment Analysis via Modality Optimization and Dynamic Primary Modality Selection
by: Yang, Dingkang, et al.
Published: (2025)
by: Yang, Dingkang, et al.
Published: (2025)
MARE: Multimodal Alignment and Reinforcement for Explainable Deepfake Detection via Vision-Language Models
by: Xu, Wenbo, et al.
Published: (2026)
by: Xu, Wenbo, et al.
Published: (2026)
Rethinking Learned Image Compression: Context is All You Need
by: Luo, Jixiang
Published: (2024)
by: Luo, Jixiang
Published: (2024)
NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
by: Tian, Changyao, et al.
Published: (2025)
by: Tian, Changyao, et al.
Published: (2025)
Multimodal LLMs under Pairwise Modalities
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
How Does the Textual Information Affect the Retrieval of Multimodal In-Context Learning?
by: Luo, Yang, et al.
Published: (2024)
by: Luo, Yang, et al.
Published: (2024)
Progressive Representation Learning for Multimodal Sentiment Analysis with Incomplete Modalities
by: Bao, Jindi, et al.
Published: (2026)
by: Bao, Jindi, et al.
Published: (2026)
Similar Items
-
One-stage Modality Distillation for Incomplete Multimodal Learning
by: Wei, Shicai, et al.
Published: (2023) -
Improving Multimodal Learning via Imbalanced Learning
by: Wei, Shicai, et al.
Published: (2025) -
Boosting Multimodal Learning via Disentangled Gradient Learning
by: Wei, Shicai, et al.
Published: (2025) -
Robust Multimodal Learning via Representation Decoupling
by: Wei, Shicai, et al.
Published: (2024) -
Scale Decoupled Distillation
by: Luo, Shicai Wei Chunbo Luo Yang
Published: (2024)