MiMIC: Mitigating Visual Modality Collapse in Universal Multimodal Retrieval While Avoiding Semantic Misalignment
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Juan, Ding, Chuanghao, Zhang, Xujie, Nguyen, Cam-Tu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SynthDoc: Bilingual Documents Synthesis for Visual Document Understanding
by: Ding, Chuanghao, et al.
Published: (2024)
by: Ding, Chuanghao, et al.
Published: (2024)
Mitigating Semantic Collapse in Partially Relevant Video Retrieval
by: Moon, WonJun, et al.
Published: (2025)
by: Moon, WonJun, et al.
Published: (2025)
MiMIC: Multi-Modal Indian Earnings Calls Dataset to Predict Stock Prices
by: Ghosh, Sohom, et al.
Published: (2025)
by: Ghosh, Sohom, et al.
Published: (2025)
Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal Manipulations
by: Shen, Jinjie, et al.
Published: (2025)
by: Shen, Jinjie, et al.
Published: (2025)
A Closer Look at Multimodal Representation Collapse
by: Chaudhuri, Abhra, et al.
Published: (2025)
by: Chaudhuri, Abhra, et al.
Published: (2025)
Federated Cross-Modal Retrieval with Missing Modalities via Semantic Routing and Adapter Personalization
by: Zhou, Hefeng, et al.
Published: (2026)
by: Zhou, Hefeng, et al.
Published: (2026)
MissBench: Benchmarking Multimodal Affective Analysis under Imbalanced Missing Modalities
by: Pham, Tien Anh, et al.
Published: (2026)
by: Pham, Tien Anh, et al.
Published: (2026)
Beyond Semantic Priors: Mitigating Optimization Collapse for Generalizable Visual Forensics
by: Liu, Jipeng, et al.
Published: (2026)
by: Liu, Jipeng, et al.
Published: (2026)
StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic Segmentation
by: Li, Bingyu, et al.
Published: (2024)
by: Li, Bingyu, et al.
Published: (2024)
Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering
by: Hong, Yuyang, et al.
Published: (2025)
by: Hong, Yuyang, et al.
Published: (2025)
Multimodal Contextualized Support for Enhancing Video Retrieval System
by: Nguyen-Le, Quoc-Bao, et al.
Published: (2024)
by: Nguyen-Le, Quoc-Bao, et al.
Published: (2024)
LLandMark: A Multi-Agent Framework for Landmark-Aware Multimodal Interactive Video Retrieval
by: Phung, Minh-Chi, et al.
Published: (2026)
by: Phung, Minh-Chi, et al.
Published: (2026)
Modelling Visual Semantics via Image Captioning to extract Enhanced Multi-Level Cross-Modal Semantic Incongruity Representation with Attention for Multimodal Sarcasm Detection
by: Aggarwal, Sajal, et al.
Published: (2024)
by: Aggarwal, Sajal, et al.
Published: (2024)
Combating Visual Neglect and Semantic Drift in Large Multimodal Models for Enhanced Cross-Modal Retrieval
by: Zhang, Guosheng, et al.
Published: (2026)
by: Zhang, Guosheng, et al.
Published: (2026)
Explaining and Mitigating the Modality Gap in Contrastive Multimodal Learning
by: Yaras, Can, et al.
Published: (2024)
by: Yaras, Can, et al.
Published: (2024)
The Role of Visual Modality in Multimodal Mathematical Reasoning: Challenges and Insights
by: Liu, Yufang, et al.
Published: (2025)
by: Liu, Yufang, et al.
Published: (2025)
Preserving Cross-Modal Stability for Visual Unlearning in Multimodal Scenarios
by: Li, Jinghan Xu Yuyang Zhang Qixuan Cai Jiancheng Chen Keqiu
Published: (2025)
by: Li, Jinghan Xu Yuyang Zhang Qixuan Cai Jiancheng Chen Keqiu
Published: (2025)
Leveraging Retrieval Augment Approach for Multimodal Emotion Recognition Under Missing Modalities
by: Fan, Qi, et al.
Published: (2024)
by: Fan, Qi, et al.
Published: (2024)
Mitigating Visual Hallucinations via Semantic Curriculum Preference Optimization in MLLMs
by: Li, Yuanshuai, et al.
Published: (2025)
by: Li, Yuanshuai, et al.
Published: (2025)
HV-Attack: Hierarchical Visual Attack for Multimodal Retrieval Augmented Generation
by: Luo, Linyin, et al.
Published: (2025)
by: Luo, Linyin, et al.
Published: (2025)
AlignZeg: Mitigating Objective Misalignment for Zero-shot Semantic Segmentation
by: Ge, Jiannan, et al.
Published: (2024)
by: Ge, Jiannan, et al.
Published: (2024)
Semantic-Consistent Bidirectional Contrastive Hashing for Noisy Multi-Label Cross-Modal Retrieval
by: Peng, Likang, et al.
Published: (2025)
by: Peng, Likang, et al.
Published: (2025)
VideoSEAL: Mitigating Evidence Misalignment in Agentic Long Video Understanding by Decoupling Answer Authority
by: Qiu, Chenhao, et al.
Published: (2026)
by: Qiu, Chenhao, et al.
Published: (2026)
Diagnosing and Mitigating Modality Interference in Multimodal Large Language Models
by: Cai, Rui, et al.
Published: (2025)
by: Cai, Rui, et al.
Published: (2025)
Leveraging Hierarchical Image-Text Misalignment for Universal Fake Image Detection
by: Zhang, Daichi, et al.
Published: (2025)
by: Zhang, Daichi, et al.
Published: (2025)
On the Value of Cross-Modal Misalignment in Multimodal Representation Learning
by: Cai, Yichao, et al.
Published: (2025)
by: Cai, Yichao, et al.
Published: (2025)
Rethinking Visual Token Reduction in LVLMs Under Cross-Modal Misalignment
by: Xu, Rui, et al.
Published: (2025)
by: Xu, Rui, et al.
Published: (2025)
Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
by: Hao, Jitai, et al.
Published: (2025)
by: Hao, Jitai, et al.
Published: (2025)
Projecting Gaussian Ellipsoids While Avoiding Affine Projection Approximation
by: Qi, Han, et al.
Published: (2024)
by: Qi, Han, et al.
Published: (2024)
Rad-VLSM: A Cross-Modal Framework with Semantics-Assisted Prompting for Medical Segmentation and Diagnosis
by: Zhang, Fengyi, et al.
Published: (2026)
by: Zhang, Fengyi, et al.
Published: (2026)
IDMR: Towards Instance-Driven Precise Visual Correspondence in Multimodal Retrieval
by: Liu, Bangwei, et al.
Published: (2025)
by: Liu, Bangwei, et al.
Published: (2025)
Unifying Visual and Semantic Feature Spaces with Diffusion Models for Enhanced Cross-Modal Alignment
by: Zheng, Yuze, et al.
Published: (2024)
by: Zheng, Yuze, et al.
Published: (2024)
OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal Retrieval
by: Yang, Wei, et al.
Published: (2025)
by: Yang, Wei, et al.
Published: (2025)
StableAnimator++: Overcoming Pose Misalignment and Face Distortion for Human Image Animation
by: Tu, Shuyuan, et al.
Published: (2025)
by: Tu, Shuyuan, et al.
Published: (2025)
STSA: Spatial-Temporal Semantic Alignment for Visual Dubbing
by: Ding, Zijun, et al.
Published: (2025)
by: Ding, Zijun, et al.
Published: (2025)
Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination
by: Zheng, Haojie, et al.
Published: (2024)
by: Zheng, Haojie, et al.
Published: (2024)
Robustness Evaluation of OCR-based Visual Document Understanding under Multi-Modal Adversarial Attacks
by: Tien, Dong Nguyen, et al.
Published: (2025)
by: Tien, Dong Nguyen, et al.
Published: (2025)
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion
by: Mistretta, Marco, et al.
Published: (2025)
by: Mistretta, Marco, et al.
Published: (2025)
Adaptive Video Distillation: Mitigating Oversaturation and Temporal Collapse in Few-Step Generation
by: You, Yuyang, et al.
Published: (2026)
by: You, Yuyang, et al.
Published: (2026)
Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks
by: Matsuishi, Koki, et al.
Published: (2025)
by: Matsuishi, Koki, et al.
Published: (2025)
Similar Items
-
SynthDoc: Bilingual Documents Synthesis for Visual Document Understanding
by: Ding, Chuanghao, et al.
Published: (2024) -
Mitigating Semantic Collapse in Partially Relevant Video Retrieval
by: Moon, WonJun, et al.
Published: (2025) -
MiMIC: Multi-Modal Indian Earnings Calls Dataset to Predict Stock Prices
by: Ghosh, Sohom, et al.
Published: (2025) -
Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal Manipulations
by: Shen, Jinjie, et al.
Published: (2025) -
A Closer Look at Multimodal Representation Collapse
by: Chaudhuri, Abhra, et al.
Published: (2025)