Understanding the Emergence of Multimodal Representation Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tjandrasuwita, Megan, Ekbote, Chanakya, Ziyin, Liu, Liang, Paul Pu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO Training
von: Dai, Wei, et al.
Veröffentlicht: (2025)
von: Dai, Wei, et al.
Veröffentlicht: (2025)
SCATR: Simple Calibrated Test-Time Ranking
von: Shyamal, Divya, et al.
Veröffentlicht: (2026)
von: Shyamal, Divya, et al.
Veröffentlicht: (2026)
What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov Chains
von: Ekbote, Chanakya, et al.
Veröffentlicht: (2025)
von: Ekbote, Chanakya, et al.
Veröffentlicht: (2025)
MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation
von: Ekbote, Chanakya, et al.
Veröffentlicht: (2025)
von: Ekbote, Chanakya, et al.
Veröffentlicht: (2025)
OmniSapiens: A Foundation Model for Social Behavior Processing via Heterogeneity-Aware Relative Policy Optimization
von: Ong, Keane, et al.
Veröffentlicht: (2026)
von: Ong, Keane, et al.
Veröffentlicht: (2026)
PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts
von: Li, Hengzhi, et al.
Veröffentlicht: (2025)
von: Li, Hengzhi, et al.
Veröffentlicht: (2025)
NeuroAtlas: Benchmarking Foundation Models for Clinical EEG and Brain-Computer Interfaces
von: Kontras, Konstantinos, et al.
Veröffentlicht: (2026)
von: Kontras, Konstantinos, et al.
Veröffentlicht: (2026)
MultiMed: Massively Multimodal and Multitask Medical Understanding
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping
von: Shan, Xiaojun, et al.
Veröffentlicht: (2025)
von: Shan, Xiaojun, et al.
Veröffentlicht: (2025)
Three Mechanisms of Feature Learning in a Linear Network
von: Xu, Yizhou, et al.
Veröffentlicht: (2024)
von: Xu, Yizhou, et al.
Veröffentlicht: (2024)
Compositional Generalization via Forced Rendering of Disentangled Latents
von: Liang, Qiyao, et al.
Veröffentlicht: (2025)
von: Liang, Qiyao, et al.
Veröffentlicht: (2025)
Remove Symmetries to Control Model Expressivity and Improve Optimization
von: Ziyin, Liu, et al.
Veröffentlicht: (2024)
von: Ziyin, Liu, et al.
Veröffentlicht: (2024)
Noise Balance and Stationary Distribution of Stochastic Gradient Descent
von: Ziyin, Liu, et al.
Veröffentlicht: (2023)
von: Ziyin, Liu, et al.
Veröffentlicht: (2023)
EvoLlama: Enhancing LLMs' Understanding of Proteins via Multimodal Structure and Sequence Representations
von: Liu, Nuowei, et al.
Veröffentlicht: (2024)
von: Liu, Nuowei, et al.
Veröffentlicht: (2024)
Decipher the Modality Gap in Multimodal Contrastive Learning: From Convergent Representations to Pairwise Alignment
von: Yi, Lingjie, et al.
Veröffentlicht: (2025)
von: Yi, Lingjie, et al.
Veröffentlicht: (2025)
On the Spectral Geometry of Cross-Modal Representations: A Functional Map Diagnostic for Multimodal Alignment
von: Sarkar, Krisanu
Veröffentlicht: (2026)
von: Sarkar, Krisanu
Veröffentlicht: (2026)
Enhancing Modality Representation and Alignment for Multimodal Cold-start Active Learning
von: Shen, Meng, et al.
Veröffentlicht: (2024)
von: Shen, Meng, et al.
Veröffentlicht: (2024)
Multimodal Representation Learning Conditioned on Semantic Relations
von: Qiao, Yang, et al.
Veröffentlicht: (2025)
von: Qiao, Yang, et al.
Veröffentlicht: (2025)
Gramian Multimodal Representation Learning and Alignment
von: Cicchetti, Giordano, et al.
Veröffentlicht: (2024)
von: Cicchetti, Giordano, et al.
Veröffentlicht: (2024)
Robust Multimodal Representation Learning in Healthcare
von: Zhu, Xiaoguang, et al.
Veröffentlicht: (2026)
von: Zhu, Xiaoguang, et al.
Veröffentlicht: (2026)
Multi-Way Representation Alignment
von: Achara, Akshit, et al.
Veröffentlicht: (2026)
von: Achara, Akshit, et al.
Veröffentlicht: (2026)
Learning Chemical Reaction Representation with Reactant-Product Alignment
von: Zeng, Kaipeng, et al.
Veröffentlicht: (2024)
von: Zeng, Kaipeng, et al.
Veröffentlicht: (2024)
Representation Alignment Rests on Linear Structure
von: Bangachev, Kiril, et al.
Veröffentlicht: (2026)
von: Bangachev, Kiril, et al.
Veröffentlicht: (2026)
An Optimization Algorithm for Multimodal Data Alignment
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
CAP: Controllable Alignment Prompting for Unlearning in LLMs
von: Wang, Zhaokun, et al.
Veröffentlicht: (2026)
von: Wang, Zhaokun, et al.
Veröffentlicht: (2026)
Understanding the Learning Dynamics of Alignment with Human Feedback
von: Im, Shawn, et al.
Veröffentlicht: (2024)
von: Im, Shawn, et al.
Veröffentlicht: (2024)
Towards a Learning Theory of Representation Alignment
von: Insulla, Francesco, et al.
Veröffentlicht: (2025)
von: Insulla, Francesco, et al.
Veröffentlicht: (2025)
A Self-guided Multimodal Approach to Enhancing Graph Representation Learning for Alzheimer's Diseases
von: Wang, Zhepeng, et al.
Veröffentlicht: (2024)
von: Wang, Zhepeng, et al.
Veröffentlicht: (2024)
Seeking the Sufficiency and Necessity Causal Features in Multimodal Representation Learning
von: Chen, Boyu, et al.
Veröffentlicht: (2024)
von: Chen, Boyu, et al.
Veröffentlicht: (2024)
Diffusion Model with Representation Alignment for Protein Inverse Folding
von: Wang, Chenglin, et al.
Veröffentlicht: (2024)
von: Wang, Chenglin, et al.
Veröffentlicht: (2024)
FANoise: Singular Value-Adaptive Noise Modulation for Robust Multimodal Representation Learning
von: Li, Jiaoyang, et al.
Veröffentlicht: (2025)
von: Li, Jiaoyang, et al.
Veröffentlicht: (2025)
Causal Debiasing Medical Multimodal Representation Learning with Missing Modalities
von: Zhu, Xiaoguang, et al.
Veröffentlicht: (2025)
von: Zhu, Xiaoguang, et al.
Veröffentlicht: (2025)
Understanding Dementia Speech Alignment with Diffusion-Based Image Generation
von: Mansi, et al.
Veröffentlicht: (2025)
von: Mansi, et al.
Veröffentlicht: (2025)
Noisy-Pair Robust Representation Alignment for Positive-Unlabeled Learning
von: Zhao, Hengwei, et al.
Veröffentlicht: (2025)
von: Zhao, Hengwei, et al.
Veröffentlicht: (2025)
Representational Alignment with Chemical Induced Fit for Molecular Relational Learning
von: Zhang, Peiliang, et al.
Veröffentlicht: (2025)
von: Zhang, Peiliang, et al.
Veröffentlicht: (2025)
Vision-Based Generic Potential Function for Policy Alignment in Multi-Agent Reinforcement Learning
von: Ma, Hao, et al.
Veröffentlicht: (2025)
von: Ma, Hao, et al.
Veröffentlicht: (2025)
Knowledge-enhanced Multimodal ECG Representation Learning with Arbitrary-Lead Inputs
von: Liu, Che, et al.
Veröffentlicht: (2025)
von: Liu, Che, et al.
Veröffentlicht: (2025)
Hierarchical Learning for Maze Navigation: Emergence of Mental Representations via Second-Order Learning
von: Manir, Shalima Binta, et al.
Veröffentlicht: (2025)
von: Manir, Shalima Binta, et al.
Veröffentlicht: (2025)
Thermodynamic Irreversibility of Training Algorithms
von: Ziyin, Liu, et al.
Veröffentlicht: (2026)
von: Ziyin, Liu, et al.
Veröffentlicht: (2026)
Align Your Query: Representation Alignment for Multimodality Medical Object Detection
von: Seo, Ara, et al.
Veröffentlicht: (2025)
von: Seo, Ara, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO Training
von: Dai, Wei, et al.
Veröffentlicht: (2025) -
SCATR: Simple Calibrated Test-Time Ranking
von: Shyamal, Divya, et al.
Veröffentlicht: (2026) -
What One Cannot, Two Can: Two-Layer Transformers Provably Represent Induction Heads on Any-Order Markov Chains
von: Ekbote, Chanakya, et al.
Veröffentlicht: (2025) -
MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation
von: Ekbote, Chanakya, et al.
Veröffentlicht: (2025) -
OmniSapiens: A Foundation Model for Social Behavior Processing via Heterogeneity-Aware Relative Policy Optimization
von: Ong, Keane, et al.
Veröffentlicht: (2026)