Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information Decomposition
Fuente:
arXiv
Saved in:
| Main Authors: | Fang, Wanlong, Zhang, Tianle, Tao, Wen, Chan, Alvin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can LLMs Reason Over Non-Text Modalities in a Training-Free Manner? A Case Study with In-Context Representation Learning
by: Zhang, Tianle, et al.
Published: (2025)
by: Zhang, Tianle, et al.
Published: (2025)
To Align or Not to Align: Strategic Multimodal Representation Alignment for Optimal Performance
by: Fang, Wanlong, et al.
Published: (2025)
by: Fang, Wanlong, et al.
Published: (2025)
When Safe Unimodal Inputs Collide: Optimizing Reasoning Chains for Cross-Modal Safety in Multimodal Large Language Models
by: Cai, Wei, et al.
Published: (2025)
by: Cai, Wei, et al.
Published: (2025)
Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security
by: Fang, Xiang, et al.
Published: (2026)
by: Fang, Xiang, et al.
Published: (2026)
Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue Models
by: Chen, Yifu, et al.
Published: (2026)
by: Chen, Yifu, et al.
Published: (2026)
Efficient Table Retrieval and Understanding with Multimodal Large Language Models
by: Xu, Zhuoyan, et al.
Published: (2026)
by: Xu, Zhuoyan, et al.
Published: (2026)
Auditing Language Model Unlearning via Information Decomposition
by: Goel, Anmol, et al.
Published: (2026)
by: Goel, Anmol, et al.
Published: (2026)
Single Image Reflection Separation via Dual Prior Interaction Transformer
by: Huang, Yue, et al.
Published: (2025)
by: Huang, Yue, et al.
Published: (2025)
Towards Robust Multi-Modal Reasoning via Model Selection
by: Liu, Xiangyan, et al.
Published: (2023)
by: Liu, Xiangyan, et al.
Published: (2023)
Reinforcing Language Agents via Policy Optimization with Action Decomposition
by: Wen, Muning, et al.
Published: (2024)
by: Wen, Muning, et al.
Published: (2024)
How Creative Are Large Language Models in Generating Molecules?
by: Tao, Wen, et al.
Published: (2026)
by: Tao, Wen, et al.
Published: (2026)
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
by: Zhang, Jihai, et al.
Published: (2025)
by: Zhang, Jihai, et al.
Published: (2025)
OptMerge: Unifying Multimodal LLM Capabilities and Modalities via Model Merging
by: Wei, Yongxian, et al.
Published: (2025)
by: Wei, Yongxian, et al.
Published: (2025)
The Triangle of Similarity: A Multi-Faceted Framework for Comparing Neural Network Representations
by: Sirikova, Olha, et al.
Published: (2026)
by: Sirikova, Olha, et al.
Published: (2026)
Cross-Modal Consistency in Multimodal Large Language Models
by: Zhang, Xiang, et al.
Published: (2024)
by: Zhang, Xiang, et al.
Published: (2024)
MAD: Modality-Adaptive Decoding for Mitigating Cross-Modal Hallucinations in Multimodal Large Language Models
by: Chung, Sangyun, et al.
Published: (2026)
by: Chung, Sangyun, et al.
Published: (2026)
Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model
by: Li, Tianle, et al.
Published: (2025)
by: Li, Tianle, et al.
Published: (2025)
Towards Visually Grounded Multimodal Summarization via Cross-Modal Transformer and Gated Attention
by: Ali, Abid, et al.
Published: (2026)
by: Ali, Abid, et al.
Published: (2026)
Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanism
by: Chen, Tao, et al.
Published: (2026)
by: Chen, Tao, et al.
Published: (2026)
Towards Robust Argumentative Essay Understanding via TIDE: An Interactive Framework with Trial and Debate
by: Yin, Zheqin, et al.
Published: (2026)
by: Yin, Zheqin, et al.
Published: (2026)
Towards Formalizing Spuriousness of Biased Datasets Using Partial Information Decomposition
by: Halder, Barproda, et al.
Published: (2024)
by: Halder, Barproda, et al.
Published: (2024)
GR-MG: Leveraging Partially Annotated Data via Multi-Modal Goal-Conditioned Policy
by: Li, Peiyan, et al.
Published: (2024)
by: Li, Peiyan, et al.
Published: (2024)
TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem Understanding
by: Ku, Max, et al.
Published: (2025)
by: Ku, Max, et al.
Published: (2025)
Diagnosing and Mitigating Modality Interference in Multimodal Large Language Models
by: Cai, Rui, et al.
Published: (2025)
by: Cai, Rui, et al.
Published: (2025)
Conformal Information Pursuit for Interactively Guiding Large Language Models
by: Chan, Kwan Ho Ryan, et al.
Published: (2025)
by: Chan, Kwan Ho Ryan, et al.
Published: (2025)
Understanding eGFR Trajectories and Kidney Function Decline via Large Multimodal Models
by: Li, Chih-Yuan, et al.
Published: (2024)
by: Li, Chih-Yuan, et al.
Published: (2024)
Debiased Multimodal Understanding for Human Language Sequences
by: Xu, Zhi, et al.
Published: (2024)
by: Xu, Zhi, et al.
Published: (2024)
Measuring Copyright Risks of Large Language Model via Partial Information Probing
by: Zhao, Weijie, et al.
Published: (2024)
by: Zhao, Weijie, et al.
Published: (2024)
Partial Information Decomposition for Data Interpretability and Feature Selection
by: Westphal, Charles, et al.
Published: (2024)
by: Westphal, Charles, et al.
Published: (2024)
MLlm-DR: Towards Explainable Depression Recognition with MultiModal Large Language Models
by: Zhang, Wei, et al.
Published: (2025)
by: Zhang, Wei, et al.
Published: (2025)
Cross-Modal Unlearning via Influential Neuron Path Editing in Multimodal Large Language Models
by: Li, Kunhao, et al.
Published: (2025)
by: Li, Kunhao, et al.
Published: (2025)
STaR: Towards Effective and Stable Table Reasoning via Slow-Thinking Large Language Models
by: Zhang, Huajian, et al.
Published: (2025)
by: Zhang, Huajian, et al.
Published: (2025)
MultiSHAP: A Shapley-Based Framework for Explaining Cross-Modal Interactions in Multimodal AI Models
by: Wang, Zhanliang, et al.
Published: (2025)
by: Wang, Zhanliang, et al.
Published: (2025)
Photon: Speedup Volume Understanding with Efficient Multimodal Large Language Models
by: Fang, Chengyu, et al.
Published: (2026)
by: Fang, Chengyu, et al.
Published: (2026)
Understanding Chain-of-Thought in Large Language Models via Topological Data Analysis
by: Li, Chenghao, et al.
Published: (2025)
by: Li, Chenghao, et al.
Published: (2025)
Beyond Text-Dominance: Understanding Modality Preference of Omni-modal Large Language Models
by: Yan, Xinru, et al.
Published: (2026)
by: Yan, Xinru, et al.
Published: (2026)
Towards Theoretical Understandings of Self-Consuming Generative Models
by: Fu, Shi, et al.
Published: (2024)
by: Fu, Shi, et al.
Published: (2024)
Modality-Guided Mixture of Graph Experts with Entropy-Triggered Routing for Multimodal Recommendation
by: Dai, Ji, et al.
Published: (2026)
by: Dai, Ji, et al.
Published: (2026)
Towards Unified Multimodal Financial Forecasting: Integrating Sentiment Embeddings and Market Indicators via Cross-Modal Attention
by: Khanna, Sarthak, et al.
Published: (2025)
by: Khanna, Sarthak, et al.
Published: (2025)
Can Multimodal Large Language Models Truly Understand Small Objects?
by: Han, Fujun, et al.
Published: (2026)
by: Han, Fujun, et al.
Published: (2026)
Similar Items
-
Can LLMs Reason Over Non-Text Modalities in a Training-Free Manner? A Case Study with In-Context Representation Learning
by: Zhang, Tianle, et al.
Published: (2025) -
To Align or Not to Align: Strategic Multimodal Representation Alignment for Optimal Performance
by: Fang, Wanlong, et al.
Published: (2025) -
When Safe Unimodal Inputs Collide: Optimizing Reasoning Chains for Cross-Modal Safety in Multimodal Large Language Models
by: Cai, Wei, et al.
Published: (2025) -
Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security
by: Fang, Xiang, et al.
Published: (2026) -
Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue Models
by: Chen, Yifu, et al.
Published: (2026)