DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qian, Chengxuan, Xing, Shuo, Li, Shawn, Zhao, Yue, Tu, Zhengzhong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference Optimization
von: Xing, Shuo, et al.
Veröffentlicht: (2025)
von: Xing, Shuo, et al.
Veröffentlicht: (2025)
DPU: Dynamic Prototype Updating for Multimodal Out-of-Distribution Detection
von: Li, Shawn, et al.
Veröffentlicht: (2024)
von: Li, Shawn, et al.
Veröffentlicht: (2024)
DeRA: Decoupled Representation Alignment for Video Tokenization
von: Guo, Pengbo, et al.
Veröffentlicht: (2025)
von: Guo, Pengbo, et al.
Veröffentlicht: (2025)
Representation Space Constrained Learning with Modality Decoupling for Multimodal Object Detection
von: Shao, YiKang, et al.
Veröffentlicht: (2025)
von: Shao, YiKang, et al.
Veröffentlicht: (2025)
DynCIM: Dynamic Curriculum for Imbalanced Multimodal Learning
von: Qian, Chengxuan, et al.
Veröffentlicht: (2025)
von: Qian, Chengxuan, et al.
Veröffentlicht: (2025)
AlignGen: Boosting Personalized Image Generation with Cross-Modality Prior Alignment
von: Lin, Yiheng, et al.
Veröffentlicht: (2025)
von: Lin, Yiheng, et al.
Veröffentlicht: (2025)
Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models
von: Zhu, Tinghui, et al.
Veröffentlicht: (2024)
von: Zhu, Tinghui, et al.
Veröffentlicht: (2024)
Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models
von: Xiao, Junyuan, et al.
Veröffentlicht: (2026)
von: Xiao, Junyuan, et al.
Veröffentlicht: (2026)
OpenEMMA: Open-Source Multimodal Model for End-to-End Autonomous Driving
von: Xing, Shuo, et al.
Veröffentlicht: (2024)
von: Xing, Shuo, et al.
Veröffentlicht: (2024)
AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment
von: Li, Yan, et al.
Veröffentlicht: (2024)
von: Li, Yan, et al.
Veröffentlicht: (2024)
Learning Modality Knowledge Alignment for Cross-Modality Transfer
von: Ma, Wenxuan, et al.
Veröffentlicht: (2024)
von: Ma, Wenxuan, et al.
Veröffentlicht: (2024)
Hierarchical Cross-Modal Alignment for Open-Vocabulary 3D Object Detection
von: Zhao, Youjun, et al.
Veröffentlicht: (2025)
von: Zhao, Youjun, et al.
Veröffentlicht: (2025)
Phrase Decoupling Cross-Modal Hierarchical Matching and Progressive Position Correction for Visual Grounding
von: Xie, Minghong, et al.
Veröffentlicht: (2024)
von: Xie, Minghong, et al.
Veröffentlicht: (2024)
Guiding Cross-Modal Representations with MLLM Priors via Preference Alignment
von: Zhao, Pengfei, et al.
Veröffentlicht: (2025)
von: Zhao, Pengfei, et al.
Veröffentlicht: (2025)
Correlation-Decoupled Knowledge Distillation for Multimodal Sentiment Analysis with Incomplete Modalities
von: Li, Mingcheng, et al.
Veröffentlicht: (2024)
von: Li, Mingcheng, et al.
Veröffentlicht: (2024)
Charts Are Not Images: On the Challenges of Scientific Chart Editing
von: Li, Shawn, et al.
Veröffentlicht: (2025)
von: Li, Shawn, et al.
Veröffentlicht: (2025)
CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook
von: Chen, Zeyu, et al.
Veröffentlicht: (2026)
von: Chen, Zeyu, et al.
Veröffentlicht: (2026)
Progressive Representation Learning for Multimodal Sentiment Analysis with Incomplete Modalities
von: Bao, Jindi, et al.
Veröffentlicht: (2026)
von: Bao, Jindi, et al.
Veröffentlicht: (2026)
StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval
von: Wang, Shaokun, et al.
Veröffentlicht: (2026)
von: Wang, Shaokun, et al.
Veröffentlicht: (2026)
Enhancing CLIP Robustness via Cross-Modality Alignment
von: Zhu, Xingyu, et al.
Veröffentlicht: (2025)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2025)
Cross-Modal-Domain Generalization Through Semantically Aligned Discrete Representations
von: Sen, Souptik, et al.
Veröffentlicht: (2026)
von: Sen, Souptik, et al.
Veröffentlicht: (2026)
Region-R1: Reinforcing Query-Side Region Cropping for Multi-Modal Re-Ranking
von: Hu, Chan-Wei, et al.
Veröffentlicht: (2026)
von: Hu, Chan-Wei, et al.
Veröffentlicht: (2026)
Decoupled Cross-Modal Alignment Network for Text-RGBT Person Retrieval and A High-Quality Benchmark
von: Deng, Yifei, et al.
Veröffentlicht: (2025)
von: Deng, Yifei, et al.
Veröffentlicht: (2025)
Demystifying the Visual Quality Paradox in Multimodal Large Language Models
von: Xing, Shuo, et al.
Veröffentlicht: (2025)
von: Xing, Shuo, et al.
Veröffentlicht: (2025)
Decoupled Hierarchical Distillation for Multimodal Emotion Recognition
von: Li, Yong, et al.
Veröffentlicht: (2026)
von: Li, Yong, et al.
Veröffentlicht: (2026)
On the Value of Cross-Modal Misalignment in Multimodal Representation Learning
von: Cai, Yichao, et al.
Veröffentlicht: (2025)
von: Cai, Yichao, et al.
Veröffentlicht: (2025)
HFMF: Hierarchical Fusion Meets Multi-Stream Models for Deepfake Detection
von: Mehta, Anant, et al.
Veröffentlicht: (2025)
von: Mehta, Anant, et al.
Veröffentlicht: (2025)
Open-set Cross Modal Generalization via Multimodal Unified Representation
von: Huang, Hai, et al.
Veröffentlicht: (2025)
von: Huang, Hai, et al.
Veröffentlicht: (2025)
Cross-Modal Scene Semantic Alignment for Image Complexity Assessment
von: Luo, Yuqing, et al.
Veröffentlicht: (2025)
von: Luo, Yuqing, et al.
Veröffentlicht: (2025)
Anisotropic Modality Align
von: Yu, Xiaomin, et al.
Veröffentlicht: (2026)
von: Yu, Xiaomin, et al.
Veröffentlicht: (2026)
Improving Medical Visual Representation Learning with Pathological-level Cross-Modal Alignment and Correlation Exploration
von: Wang, Jun, et al.
Veröffentlicht: (2025)
von: Wang, Jun, et al.
Veröffentlicht: (2025)
T2T-VICL: Unlocking the Boundaries of Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs
von: Xia, Shao-Jun, et al.
Veröffentlicht: (2025)
von: Xia, Shao-Jun, et al.
Veröffentlicht: (2025)
Enhancing Multimodal Unified Representations for Cross Modal Generalization
von: Huang, Hai, et al.
Veröffentlicht: (2024)
von: Huang, Hai, et al.
Veröffentlicht: (2024)
Calibrated Multimodal Representation Learning with Missing Modalities
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
PROMISE: Prompt-Attentive Hierarchical Contrastive Learning for Robust Cross-Modal Representation with Missing Modalities
von: Chen, Jiajun, et al.
Veröffentlicht: (2025)
von: Chen, Jiajun, et al.
Veröffentlicht: (2025)
Decoupling Common and Unique Representations for Multimodal Self-supervised Learning
von: Wang, Yi, et al.
Veröffentlicht: (2023)
von: Wang, Yi, et al.
Veröffentlicht: (2023)
TextAlign: Preference Alignment for Text Rendering with Hierarchical Rewards
von: Cui, Mingxuan, et al.
Veröffentlicht: (2026)
von: Cui, Mingxuan, et al.
Veröffentlicht: (2026)
Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Model
von: Feng, Qianhan, et al.
Veröffentlicht: (2024)
von: Feng, Qianhan, et al.
Veröffentlicht: (2024)
VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization
von: Li, Mingxiao, et al.
Veröffentlicht: (2025)
von: Li, Mingxiao, et al.
Veröffentlicht: (2025)
Toward Robust Incomplete Multimodal Sentiment Analysis via Hierarchical Representation Learning
von: Li, Mingcheng, et al.
Veröffentlicht: (2024)
von: Li, Mingcheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference Optimization
von: Xing, Shuo, et al.
Veröffentlicht: (2025) -
DPU: Dynamic Prototype Updating for Multimodal Out-of-Distribution Detection
von: Li, Shawn, et al.
Veröffentlicht: (2024) -
DeRA: Decoupled Representation Alignment for Video Tokenization
von: Guo, Pengbo, et al.
Veröffentlicht: (2025) -
Representation Space Constrained Learning with Modality Decoupling for Multimodal Object Detection
von: Shao, YiKang, et al.
Veröffentlicht: (2025) -
DynCIM: Dynamic Curriculum for Imbalanced Multimodal Learning
von: Qian, Chengxuan, et al.
Veröffentlicht: (2025)