Second-Order Multi-Level Variance Correction for Modality Competition in Multimodal Models
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Yishun, Armour, Wes |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Offline Semantic Guidance for Efficient Vision-Language-Action Policy Distillation
by: Shi, Jin, et al.
Published: (2026)
by: Shi, Jin, et al.
Published: (2026)
UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding
by: Xu, Chenkai, et al.
Published: (2025)
by: Xu, Chenkai, et al.
Published: (2025)
The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models
by: Paruchuri, Akshay, et al.
Published: (2026)
by: Paruchuri, Akshay, et al.
Published: (2026)
From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding
by: Zou, Heqing, et al.
Published: (2024)
by: Zou, Heqing, et al.
Published: (2024)
Modelling Visual Semantics via Image Captioning to extract Enhanced Multi-Level Cross-Modal Semantic Incongruity Representation with Attention for Multimodal Sarcasm Detection
by: Aggarwal, Sajal, et al.
Published: (2024)
by: Aggarwal, Sajal, et al.
Published: (2024)
Towards Multimodal Video Paragraph Captioning Models Robust to Missing Modality
by: Chen, Sishuo, et al.
Published: (2024)
by: Chen, Sishuo, et al.
Published: (2024)
MammothModa: Multi-Modal Large Language Model
by: She, Qi, et al.
Published: (2024)
by: She, Qi, et al.
Published: (2024)
Towards Understanding Deep Learning Model in Image Recognition via Coverage Test
by: Li, Wenkai, et al.
Published: (2025)
by: Li, Wenkai, et al.
Published: (2025)
Spatial Transcriptomics as Images for Large-Scale Pretraining
by: Zhu, Yishun, et al.
Published: (2026)
by: Zhu, Yishun, et al.
Published: (2026)
Gram-Anchored Prompt Learning for Vision-Language Models via Second-Order Statistics
by: Chen, Minglei, et al.
Published: (2026)
by: Chen, Minglei, et al.
Published: (2026)
Multimodal Emotion Recognition with Vision-language Prompting and Modality Dropout
by: QI, Anbin, et al.
Published: (2024)
by: QI, Anbin, et al.
Published: (2024)
A Second-Order Perspective on Pruning at Initialization and Knowledge Transfer
by: Iurada, Leonardo, et al.
Published: (2025)
by: Iurada, Leonardo, et al.
Published: (2025)
Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion
by: Rawal, Ishaan Singh, et al.
Published: (2023)
by: Rawal, Ishaan Singh, et al.
Published: (2023)
Guided and Variance-Corrected Fusion with One-shot Style Alignment for Large-Content Image Generation
by: Sun, Shoukun, et al.
Published: (2024)
by: Sun, Shoukun, et al.
Published: (2024)
Explaining Caption-Image Interactions in CLIP Models with Second-Order Attributions
by: Möller, Lucas, et al.
Published: (2024)
by: Möller, Lucas, et al.
Published: (2024)
Multi-Modal Parameter-Efficient Fine-tuning via Graph Neural Network
by: Cheng, Bin, et al.
Published: (2024)
by: Cheng, Bin, et al.
Published: (2024)
SDIGLM: Leveraging Large Language Models and Multi-Modal Chain of Thought for Structural Damage Identification
by: Zhang, Yunkai, et al.
Published: (2025)
by: Zhang, Yunkai, et al.
Published: (2025)
PARTONOMY: Large Multimodal Models with Part-Level Visual Understanding
by: Blume, Ansel, et al.
Published: (2025)
by: Blume, Ansel, et al.
Published: (2025)
Unveiling Ontological Commitment in Multi-Modal Foundation Models
by: Keser, Mert, et al.
Published: (2024)
by: Keser, Mert, et al.
Published: (2024)
DiffCrossGait: Trajectory-Level Alignment for 2D-3D Cross-Modal Gait Recognition via Latent Diffusion
by: Lu, Zhiyang, et al.
Published: (2026)
by: Lu, Zhiyang, et al.
Published: (2026)
Think While Watching: Online Streaming Segment-Level Memory for Multi-Turn Video Reasoning in Multimodal Large Language Models
by: Wang, Lu, et al.
Published: (2026)
by: Wang, Lu, et al.
Published: (2026)
PulseMind: A Multi-Modal Medical Model for Real-World Clinical Diagnosis
by: Xu, Jiao, et al.
Published: (2026)
by: Xu, Jiao, et al.
Published: (2026)
Missing Modality Prediction for Unpaired Multimodal Learning via Joint Embedding of Unimodal Models
by: Kim, Donggeun, et al.
Published: (2024)
by: Kim, Donggeun, et al.
Published: (2024)
A Unified Model for Longitudinal Multi-Modal Multi-View Prediction with Missingness
by: Chen, Boqi, et al.
Published: (2024)
by: Chen, Boqi, et al.
Published: (2024)
Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation
by: Liang, Yupu, et al.
Published: (2025)
by: Liang, Yupu, et al.
Published: (2025)
Enhancing Multimodal Unified Representations for Cross Modal Generalization
by: Huang, Hai, et al.
Published: (2024)
by: Huang, Hai, et al.
Published: (2024)
GAIS: Frame-Level Gated Audio-Visual Integration with Semantic Variance-Scaled Perturbation for Text-Video Retrieval
by: Yang, Bowen, et al.
Published: (2025)
by: Yang, Bowen, et al.
Published: (2025)
Multi-Modal Foundation Models for Computational Pathology: A Survey
by: Li, Dong, et al.
Published: (2025)
by: Li, Dong, et al.
Published: (2025)
Beyond Cross-Modal Alignment: Measuring and Leveraging Modality Gap in Vision-Language Models
by: Yan, Hanqi, et al.
Published: (2025)
by: Yan, Hanqi, et al.
Published: (2025)
Robust Multimodal Large Language Models Against Modality Conflict
by: Zhang, Zongmeng, et al.
Published: (2025)
by: Zhang, Zongmeng, et al.
Published: (2025)
Towards a Multimodal Large Language Model with Pixel-Level Insight for Biomedicine
by: Huang, Xiaoshuang, et al.
Published: (2024)
by: Huang, Xiaoshuang, et al.
Published: (2024)
MediRound: Multi-Round Entity-Level Reasoning Segmentation in Medical Images
by: Tong, Qinyue, et al.
Published: (2025)
by: Tong, Qinyue, et al.
Published: (2025)
Team LEYA in 10th ABAW Competition: Multimodal Ambivalence/Hesitancy Recognition Approach
by: Ryumina, Elena, et al.
Published: (2026)
by: Ryumina, Elena, et al.
Published: (2026)
Team RAS in 10th ABAW Competition: Multimodal Valence and Arousal Estimation Approach
by: Ryumina, Elena, et al.
Published: (2026)
by: Ryumina, Elena, et al.
Published: (2026)
MIAR: Modality Interaction and Alignment Representation Fuison for Multimodal Emotion
by: Zhu, Jichao, et al.
Published: (2026)
by: Zhu, Jichao, et al.
Published: (2026)
The Role of Visual Modality in Multimodal Mathematical Reasoning: Challenges and Insights
by: Liu, Yufang, et al.
Published: (2025)
by: Liu, Yufang, et al.
Published: (2025)
Preserving Cross-Modal Stability for Visual Unlearning in Multimodal Scenarios
by: Li, Jinghan Xu Yuyang Zhang Qixuan Cai Jiancheng Chen Keqiu
Published: (2025)
by: Li, Jinghan Xu Yuyang Zhang Qixuan Cai Jiancheng Chen Keqiu
Published: (2025)
Revisiting Multi-Modal LLM Evaluation
by: Lu, Jian, et al.
Published: (2024)
by: Lu, Jian, et al.
Published: (2024)
Region-Level Context-Aware Multimodal Understanding
by: Wei, Hongliang, et al.
Published: (2025)
by: Wei, Hongliang, et al.
Published: (2025)
D3: Training-Free AI-Generated Video Detection Using Second-Order Features
by: Zheng, Chende, et al.
Published: (2025)
by: Zheng, Chende, et al.
Published: (2025)
Similar Items
-
Offline Semantic Guidance for Efficient Vision-Language-Action Policy Distillation
by: Shi, Jin, et al.
Published: (2026) -
UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding
by: Xu, Chenkai, et al.
Published: (2025) -
The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models
by: Paruchuri, Akshay, et al.
Published: (2026) -
From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding
by: Zou, Heqing, et al.
Published: (2024) -
Modelling Visual Semantics via Image Captioning to extract Enhanced Multi-Level Cross-Modal Semantic Incongruity Representation with Attention for Multimodal Sarcasm Detection
by: Aggarwal, Sajal, et al.
Published: (2024)