Robust Multimodal Learning via Representation Decoupling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wei, Shicai, Luo, Yang, Wang, Yuji, Luo, Chunbo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scale Decoupled Distillation
von: Luo, Shicai Wei Chunbo Luo Yang
Veröffentlicht: (2024)
von: Luo, Shicai Wei Chunbo Luo Yang
Veröffentlicht: (2024)
Improving Multimodal Learning via Imbalanced Learning
von: Wei, Shicai, et al.
Veröffentlicht: (2025)
von: Wei, Shicai, et al.
Veröffentlicht: (2025)
Boosting Multimodal Learning via Disentangled Gradient Learning
von: Wei, Shicai, et al.
Veröffentlicht: (2025)
von: Wei, Shicai, et al.
Veröffentlicht: (2025)
One-stage Modality Distillation for Incomplete Multimodal Learning
von: Wei, Shicai, et al.
Veröffentlicht: (2023)
von: Wei, Shicai, et al.
Veröffentlicht: (2023)
PDMP: Rethinking Balanced Multimodal Learning via Performance-Dominant Modality Prioritization
von: Wei, Shicai, et al.
Veröffentlicht: (2026)
von: Wei, Shicai, et al.
Veröffentlicht: (2026)
CrossVL: Complexity-Aware Feature Routing and Paired Curriculum for Cross-View Vision-Language Detection
von: Liu, Zhipeng, et al.
Veröffentlicht: (2026)
von: Liu, Zhipeng, et al.
Veröffentlicht: (2026)
Robust Multimodal Learning for Ophthalmic Disease Grading via Disentangled Representation
von: Wang, Xinkun, et al.
Veröffentlicht: (2025)
von: Wang, Xinkun, et al.
Veröffentlicht: (2025)
Contrastive Representation Distillation via Multi-Scale Feature Decoupling
von: Wang, Cuipeng, et al.
Veröffentlicht: (2025)
von: Wang, Cuipeng, et al.
Veröffentlicht: (2025)
CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook
von: Chen, Zeyu, et al.
Veröffentlicht: (2026)
von: Chen, Zeyu, et al.
Veröffentlicht: (2026)
Decoupling the Image Perception and Multimodal Reasoning for Reasoning Segmentation with Digital Twin Representations
von: Li, Yizhen, et al.
Veröffentlicht: (2025)
von: Li, Yizhen, et al.
Veröffentlicht: (2025)
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
von: Wu, Chengyue, et al.
Veröffentlicht: (2024)
von: Wu, Chengyue, et al.
Veröffentlicht: (2024)
SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model
von: Yu, Chunlin, et al.
Veröffentlicht: (2024)
von: Yu, Chunlin, et al.
Veröffentlicht: (2024)
Distilling Vision-Language Foundation Models: A Data-Free Approach via Prompt Diversification
von: Xuan, Yunyi, et al.
Veröffentlicht: (2024)
von: Xuan, Yunyi, et al.
Veröffentlicht: (2024)
How Does the Textual Information Affect the Retrieval of Multimodal In-Context Learning?
von: Luo, Yang, et al.
Veröffentlicht: (2024)
von: Luo, Yang, et al.
Veröffentlicht: (2024)
Toward Robust Early Detection of Alzheimer's Disease via an Integrated Multimodal Learning Approach
von: Chen, Yifei, et al.
Veröffentlicht: (2024)
von: Chen, Yifei, et al.
Veröffentlicht: (2024)
ACD-CLIP: Decoupling Representation and Dynamic Fusion for Zero-Shot Anomaly Detection
von: Ma, Ke, et al.
Veröffentlicht: (2025)
von: Ma, Ke, et al.
Veröffentlicht: (2025)
Composed Vision-Language Retrieval for Skin Cancer Case Search via Joint Alignment of Global and Local Representations
von: Wang, Yuheng, et al.
Veröffentlicht: (2026)
von: Wang, Yuheng, et al.
Veröffentlicht: (2026)
D-GAP: Improving Out-of-Domain Robustness via Dataset-Agnostic and Gradient-Guided Augmentation in Frequency and Pixel Spaces
von: Wang, Ruoqi, et al.
Veröffentlicht: (2025)
von: Wang, Ruoqi, et al.
Veröffentlicht: (2025)
Enhancing Visual Token Representations for Video Large Language Models via Training-Free Spatial-Temporal Pooling and Gridding
von: Luo, Bingjun, et al.
Veröffentlicht: (2026)
von: Luo, Bingjun, et al.
Veröffentlicht: (2026)
IDRL: An Individual-Aware Multimodal Depression-Related Representation Learning Framework for Depression Diagnosis
von: Wang, Chongxiao, et al.
Veröffentlicht: (2026)
von: Wang, Chongxiao, et al.
Veröffentlicht: (2026)
CURVE: Learning Causality-Inspired Invariant Representations for Robust Scene Understanding via Uncertainty-Guided Regularization
von: Liang, Yue, et al.
Veröffentlicht: (2026)
von: Liang, Yue, et al.
Veröffentlicht: (2026)
Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
von: Li, Hengzhuang, et al.
Veröffentlicht: (2025)
von: Li, Hengzhuang, et al.
Veröffentlicht: (2025)
ARIW-Framework: Adaptive Robust Iterative Watermarking Framework
von: Wu, Shaowu, et al.
Veröffentlicht: (2025)
von: Wu, Shaowu, et al.
Veröffentlicht: (2025)
Robust Domain Generalization for Multi-modal Object Recognition
von: Qiao, Yuxin, et al.
Veröffentlicht: (2024)
von: Qiao, Yuxin, et al.
Veröffentlicht: (2024)
M$^2$IV: Towards Efficient and Fine-grained Multimodal In-Context Learning via Representation Engineering
von: Li, Yanshu, et al.
Veröffentlicht: (2025)
von: Li, Yanshu, et al.
Veröffentlicht: (2025)
MCRL4OR: Multimodal Contrastive Representation Learning for Off-Road Environmental Perception
von: Yang, Yi, et al.
Veröffentlicht: (2025)
von: Yang, Yi, et al.
Veröffentlicht: (2025)
Cheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension and Generation
von: Zhang, Yichen, et al.
Veröffentlicht: (2026)
von: Zhang, Yichen, et al.
Veröffentlicht: (2026)
Learning Robust Intervention Representations with Delta Embeddings
von: Alimisis, Panagiotis, et al.
Veröffentlicht: (2025)
von: Alimisis, Panagiotis, et al.
Veröffentlicht: (2025)
VideoSEAL: Mitigating Evidence Misalignment in Agentic Long Video Understanding by Decoupling Answer Authority
von: Qiu, Chenhao, et al.
Veröffentlicht: (2026)
von: Qiu, Chenhao, et al.
Veröffentlicht: (2026)
Adaptive Disentangled Representation Learning for Incomplete Multi-View Multi-Label Classification
von: Li, Quanjiang, et al.
Veröffentlicht: (2026)
von: Li, Quanjiang, et al.
Veröffentlicht: (2026)
MAJORScore: A Novel Metric for Evaluating Multimodal Relevance via Joint Representation
von: Du, Zhicheng, et al.
Veröffentlicht: (2025)
von: Du, Zhicheng, et al.
Veröffentlicht: (2025)
Learning to Merge Tokens via Decoupled Embedding for Efficient Vision Transformers
von: Lee, Dong Hoon, et al.
Veröffentlicht: (2024)
von: Lee, Dong Hoon, et al.
Veröffentlicht: (2024)
Unified Multimodal Understanding via Byte-Pair Visual Encoding
von: Zhang, Wanpeng, et al.
Veröffentlicht: (2025)
von: Zhang, Wanpeng, et al.
Veröffentlicht: (2025)
Multimodal Prompt Decoupling Attack on the Safety Filters in Text-to-Image Models
von: Peng, Xingkai, et al.
Veröffentlicht: (2025)
von: Peng, Xingkai, et al.
Veröffentlicht: (2025)
Learning Modality-Aware Representations: Adaptive Group-wise Interaction Network for Multimodal MRI Synthesis
von: Song, Tao, et al.
Veröffentlicht: (2024)
von: Song, Tao, et al.
Veröffentlicht: (2024)
Reconstruction-Driven Multimodal Representation Learning for Automated Media Understanding
von: Benhammou, Yassir, et al.
Veröffentlicht: (2025)
von: Benhammou, Yassir, et al.
Veröffentlicht: (2025)
Multimodal Continual Learning with MLLMs from Multi-scenario Perspectives
von: Jiang, Kai, et al.
Veröffentlicht: (2025)
von: Jiang, Kai, et al.
Veröffentlicht: (2025)
VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding
von: Zhang, Zhihong, et al.
Veröffentlicht: (2025)
von: Zhang, Zhihong, et al.
Veröffentlicht: (2025)
Learning to Rematch Mismatched Pairs for Robust Cross-Modal Retrieval
von: Han, Haochen, et al.
Veröffentlicht: (2024)
von: Han, Haochen, et al.
Veröffentlicht: (2024)
MMIF-AMIN: Adaptive Loss-Driven Multi-Scale Invertible Dense Network for Multimodal Medical Image Fusion
von: Luo, Tao, et al.
Veröffentlicht: (2025)
von: Luo, Tao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Scale Decoupled Distillation
von: Luo, Shicai Wei Chunbo Luo Yang
Veröffentlicht: (2024) -
Improving Multimodal Learning via Imbalanced Learning
von: Wei, Shicai, et al.
Veröffentlicht: (2025) -
Boosting Multimodal Learning via Disentangled Gradient Learning
von: Wei, Shicai, et al.
Veröffentlicht: (2025) -
One-stage Modality Distillation for Incomplete Multimodal Learning
von: Wei, Shicai, et al.
Veröffentlicht: (2023) -
PDMP: Rethinking Balanced Multimodal Learning via Performance-Dominant Modality Prioritization
von: Wei, Shicai, et al.
Veröffentlicht: (2026)