To Align or Not to Align: Strategic Multimodal Representation Alignment for Optimal Performance
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fang, Wanlong, Zhang, Tianle, Chan, Alvin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cross-Space Synergy: A Unified Framework for Multimodal Emotion Recognition in Conversation
von: Lyu, Xiaosen, et al.
Veröffentlicht: (2025)
von: Lyu, Xiaosen, et al.
Veröffentlicht: (2025)
Multimodal Representation Learning and Fusion
von: Jin, Qihang, et al.
Veröffentlicht: (2025)
von: Jin, Qihang, et al.
Veröffentlicht: (2025)
Enhancing Modality Representation and Alignment for Multimodal Cold-start Active Learning
von: Shen, Meng, et al.
Veröffentlicht: (2024)
von: Shen, Meng, et al.
Veröffentlicht: (2024)
From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images
von: Chen, Yiming, et al.
Veröffentlicht: (2025)
von: Chen, Yiming, et al.
Veröffentlicht: (2025)
FoleyGRAM: Video-to-Audio Generation with GRAM-Aligned Multimodal Encoders
von: Gramaccioni, Riccardo Fosco, et al.
Veröffentlicht: (2025)
von: Gramaccioni, Riccardo Fosco, et al.
Veröffentlicht: (2025)
GuardAlign: Test-time Safety Alignment in Multimodal Large Language Models
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
Zero-Shot Relational Learning for Multimodal Knowledge Graphs
von: Cai, Rui, et al.
Veröffentlicht: (2024)
von: Cai, Rui, et al.
Veröffentlicht: (2024)
Aligning Audio-Visual Joint Representations with an Agentic Workflow
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
Balanced Multimodal Learning: An Unidirectional Dynamic Interaction Perspective
von: Wang, Shijie, et al.
Veröffentlicht: (2025)
von: Wang, Shijie, et al.
Veröffentlicht: (2025)
MemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification
von: Shah, Siddhant Bikram, et al.
Veröffentlicht: (2024)
von: Shah, Siddhant Bikram, et al.
Veröffentlicht: (2024)
Token-Level Contrastive Learning with Modality-Aware Prompting for Multimodal Intent Recognition
von: Zhou, Qianrui, et al.
Veröffentlicht: (2023)
von: Zhou, Qianrui, et al.
Veröffentlicht: (2023)
Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers
von: Wang, Juncheng, et al.
Veröffentlicht: (2025)
von: Wang, Juncheng, et al.
Veröffentlicht: (2025)
HyperFusion: Hierarchical Multimodal Ensemble Learning for Social Media Popularity Prediction
von: Ye, Liliang, et al.
Veröffentlicht: (2025)
von: Ye, Liliang, et al.
Veröffentlicht: (2025)
Stable Multimodal Graph Unlearning via Feature-Dimension Aware Quantile Selection
von: Zhou, Jingjing, et al.
Veröffentlicht: (2026)
von: Zhou, Jingjing, et al.
Veröffentlicht: (2026)
Design-MLLM: A Reinforcement Alignment Framework for Verifiable and Aesthetic Interior Design
von: Yang, Yuxuan, et al.
Veröffentlicht: (2026)
von: Yang, Yuxuan, et al.
Veröffentlicht: (2026)
Principled Multimodal Representation Learning
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
Hyper-modal Imputation Diffusion Embedding with Dual-Distillation for Federated Multimodal Knowledge Graph Completion
von: Zhang, Ying, et al.
Veröffentlicht: (2025)
von: Zhang, Ying, et al.
Veröffentlicht: (2025)
Exploring Modality Disruption in Multimodal Fake News Detection
von: Liu, Moyang, et al.
Veröffentlicht: (2025)
von: Liu, Moyang, et al.
Veröffentlicht: (2025)
MCAD: Multimodal Context-Aware Audio Description Generation For Soccer
von: Chaudhary, Lipisha, et al.
Veröffentlicht: (2025)
von: Chaudhary, Lipisha, et al.
Veröffentlicht: (2025)
MCIGLE: Multimodal Exemplar-Free Class-Incremental Graph Learning
von: You, Haochen, et al.
Veröffentlicht: (2025)
von: You, Haochen, et al.
Veröffentlicht: (2025)
Deconfounded Reasoning for Multimodal Fake News Detection via Causal Intervention
von: Liu, Moyang, et al.
Veröffentlicht: (2025)
von: Liu, Moyang, et al.
Veröffentlicht: (2025)
Hybrid Feedback-Guided Optimal Learning for Wireless Interactive Panoramic Scene Delivery
von: Wu, Xiaoyi, et al.
Veröffentlicht: (2026)
von: Wu, Xiaoyi, et al.
Veröffentlicht: (2026)
Concept Drift Guided LayerNorm Tuning for Efficient Multimodal Metaphor Identification
von: Qian, Wenhao, et al.
Veröffentlicht: (2025)
von: Qian, Wenhao, et al.
Veröffentlicht: (2025)
Multimodal Methods for Analyzing Learning and Training Environments: A Systematic Literature Review
von: Cohn, Clayton, et al.
Veröffentlicht: (2024)
von: Cohn, Clayton, et al.
Veröffentlicht: (2024)
GAME-ON: Graph Attention Network based Multimodal Fusion for Fake News Detection
von: Dhawan, Mudit, et al.
Veröffentlicht: (2022)
von: Dhawan, Mudit, et al.
Veröffentlicht: (2022)
OneLLM: One Framework to Align All Modalities with Language
von: Han, Jiaming, et al.
Veröffentlicht: (2023)
von: Han, Jiaming, et al.
Veröffentlicht: (2023)
Calibrated Multimodal Representation Learning with Missing Modalities
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
Aligning Agentic World Models via Knowledgeable Experience Learning
von: Ren, Baochang, et al.
Veröffentlicht: (2026)
von: Ren, Baochang, et al.
Veröffentlicht: (2026)
PrismSSL: One Interface, Many Modalities; A Single-Interface Library for Multimodal Self-Supervised Learning
von: Shirian, Melika, et al.
Veröffentlicht: (2025)
von: Shirian, Melika, et al.
Veröffentlicht: (2025)
E2LVLM:Evidence-Enhanced Large Vision-Language Model for Multimodal Out-of-Context Misinformation Detection
von: Wu, Junjie, et al.
Veröffentlicht: (2025)
von: Wu, Junjie, et al.
Veröffentlicht: (2025)
4D Multimodal Co-attention Fusion Network with Latent Contrastive Alignment for Alzheimer's Diagnosis
von: Wei, Yuxiang, et al.
Veröffentlicht: (2025)
von: Wei, Yuxiang, et al.
Veröffentlicht: (2025)
ActAlign: Zero-Shot Fine-Grained Video Classification via Language-Guided Sequence Alignment
von: Aghdam, Amir, et al.
Veröffentlicht: (2025)
von: Aghdam, Amir, et al.
Veröffentlicht: (2025)
InfoMAE: Pair-Efficient Cross-Modal Alignment for Multimodal Time-Series Sensing Signals
von: Kimura, Tomoyoshi, et al.
Veröffentlicht: (2025)
von: Kimura, Tomoyoshi, et al.
Veröffentlicht: (2025)
Training Data Efficiency in Multimodal Process Reward Models
von: Li, Jinyuan, et al.
Veröffentlicht: (2026)
von: Li, Jinyuan, et al.
Veröffentlicht: (2026)
Improving Visual Representation Alignment Generation with GRPO
von: Mo, Shentong, et al.
Veröffentlicht: (2026)
von: Mo, Shentong, et al.
Veröffentlicht: (2026)
V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2026)
von: Lin, Yan-Bo, et al.
Veröffentlicht: (2026)
Video Soundtrack Generation by Aligning Emotions and Temporal Boundaries
von: Sulun, Serkan, et al.
Veröffentlicht: (2025)
von: Sulun, Serkan, et al.
Veröffentlicht: (2025)
Hierarchical Aligned Multimodal Learning for NER on Tweet Posts
von: Liu, Peipei, et al.
Veröffentlicht: (2023)
von: Liu, Peipei, et al.
Veröffentlicht: (2023)
SyMuPe: Affective and Controllable Symbolic Music Performance
von: Borovik, Ilya, et al.
Veröffentlicht: (2025)
von: Borovik, Ilya, et al.
Veröffentlicht: (2025)
ChemDFM-X: Towards Large Multimodal Model for Chemistry
von: Zhao, Zihan, et al.
Veröffentlicht: (2024)
von: Zhao, Zihan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Cross-Space Synergy: A Unified Framework for Multimodal Emotion Recognition in Conversation
von: Lyu, Xiaosen, et al.
Veröffentlicht: (2025) -
Multimodal Representation Learning and Fusion
von: Jin, Qihang, et al.
Veröffentlicht: (2025) -
Enhancing Modality Representation and Alignment for Multimodal Cold-start Active Learning
von: Shen, Meng, et al.
Veröffentlicht: (2024) -
From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images
von: Chen, Yiming, et al.
Veröffentlicht: (2025) -
FoleyGRAM: Video-to-Audio Generation with GRAM-Aligned Multimodal Encoders
von: Gramaccioni, Riccardo Fosco, et al.
Veröffentlicht: (2025)