Combating Visual Neglect and Semantic Drift in Large Multimodal Models for Enhanced Cross-Modal Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Guosheng, Liu, Linkai, Wang, Keyao, Yue, Haixiao, Tan, Zhiwen, Tan, Xiao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models
by: Zhang, Guosheng, et al.
Published: (2025)
by: Zhang, Guosheng, et al.
Published: (2025)
From Intuition to Investigation: A Tool-Augmented Reasoning MLLM Framework for Generalizable Face Anti-Spoofing
by: Zhang, Haoyuan, et al.
Published: (2026)
by: Zhang, Haoyuan, et al.
Published: (2026)
ALoRE: Efficient Visual Adaptation via Aggregating Low Rank Experts
by: Du, Sinan, et al.
Published: (2024)
by: Du, Sinan, et al.
Published: (2024)
Robust Multimodal Semantic Segmentation with Balanced Modality Contributions
by: Tan, Jiaqi, et al.
Published: (2025)
by: Tan, Jiaqi, et al.
Published: (2025)
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning
by: Gao, Tianhong, et al.
Published: (2025)
by: Gao, Tianhong, et al.
Published: (2025)
Visual Attention Drifts,but Anchors Hold:Mitigating Hallucination in Multimodal Large Language Models via Cross-Layer Visual Anchors
by: Yang, Chengxu, et al.
Published: (2026)
by: Yang, Chengxu, et al.
Published: (2026)
Unifying Visual and Semantic Feature Spaces with Diffusion Models for Enhanced Cross-Modal Alignment
by: Zheng, Yuze, et al.
Published: (2024)
by: Zheng, Yuze, et al.
Published: (2024)
Controlling Decision Drift in Multimodal Sentiment Analysis with Missing Modalities
by: Chen, Chenglizhao, et al.
Published: (2026)
by: Chen, Chenglizhao, et al.
Published: (2026)
Parameter-Efficient Modality-Balanced Symmetric Fusion for Multimodal Remote Sensing Semantic Segmentation
by: Li, Haocheng, et al.
Published: (2026)
by: Li, Haocheng, et al.
Published: (2026)
Leveraging Modality Tags for Enhanced Cross-Modal Video Retrieval
by: Fragomeni, Adriano, et al.
Published: (2025)
by: Fragomeni, Adriano, et al.
Published: (2025)
PhysLLM: Harnessing Large Language Models for Cross-Modal Remote Physiological Sensing
by: Xie, Yiping, et al.
Published: (2025)
by: Xie, Yiping, et al.
Published: (2025)
MiMIC: Mitigating Visual Modality Collapse in Universal Multimodal Retrieval While Avoiding Semantic Misalignment
by: Li, Juan, et al.
Published: (2026)
by: Li, Juan, et al.
Published: (2026)
Visual Foundation Models Boost Cross-Modal Unsupervised Domain Adaptation for 3D Semantic Segmentation
by: Xu, Jingyi, et al.
Published: (2024)
by: Xu, Jingyi, et al.
Published: (2024)
Enhancing Cross-Modal Medical Image Segmentation through Compositionality
by: Eijpe, Aniek, et al.
Published: (2024)
by: Eijpe, Aniek, et al.
Published: (2024)
Modelling Visual Semantics via Image Captioning to extract Enhanced Multi-Level Cross-Modal Semantic Incongruity Representation with Attention for Multimodal Sarcasm Detection
by: Aggarwal, Sajal, et al.
Published: (2024)
by: Aggarwal, Sajal, et al.
Published: (2024)
Federated Cross-Modal Retrieval with Missing Modalities via Semantic Routing and Adapter Personalization
by: Zhou, Hefeng, et al.
Published: (2026)
by: Zhou, Hefeng, et al.
Published: (2026)
CAR-MFL: Cross-Modal Augmentation by Retrieval for Multimodal Federated Learning with Missing Modalities
by: Poudel, Pranav, et al.
Published: (2024)
by: Poudel, Pranav, et al.
Published: (2024)
Semantic-Preserving Cross-Style Visual Reasoning for Robust Multi-Modal Understanding in Large Vision-Language Models
by: Nakayama, Aya, et al.
Published: (2025)
by: Nakayama, Aya, et al.
Published: (2025)
MambaSOD: Dual Mamba-Driven Cross-Modal Fusion Network for RGB-D Salient Object Detection
by: Zhan, Yue, et al.
Published: (2024)
by: Zhan, Yue, et al.
Published: (2024)
Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models
by: He, Hulingxiao, et al.
Published: (2026)
by: He, Hulingxiao, et al.
Published: (2026)
Prototype-Enhanced Confidence Modeling for Cross-Modal Medical Image-Report Retrieval
by: Gowda, Shreyank N, et al.
Published: (2025)
by: Gowda, Shreyank N, et al.
Published: (2025)
Enhancing Audio-Visual Spiking Neural Networks through Semantic-Alignment and Cross-Modal Residual Learning
by: He, Xiang, et al.
Published: (2025)
by: He, Xiang, et al.
Published: (2025)
RMMSS: Towards Advanced Robust Multi-Modal Semantic Segmentation with Hybrid Prototype Distillation and Feature Selection
by: Tan, Jiaqi, et al.
Published: (2025)
by: Tan, Jiaqi, et al.
Published: (2025)
StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic Segmentation
by: Li, Bingyu, et al.
Published: (2024)
by: Li, Bingyu, et al.
Published: (2024)
Cross-Modal Scene Semantic Alignment for Image Complexity Assessment
by: Luo, Yuqing, et al.
Published: (2025)
by: Luo, Yuqing, et al.
Published: (2025)
Cross-Layer Vision Smoothing: Enhancing Visual Understanding via Sustained Focus on Key Objects in Large Vision-Language Models
by: Zhao, Jianfei, et al.
Published: (2025)
by: Zhao, Jianfei, et al.
Published: (2025)
Self-Enhanced Image Clustering with Cross-Modal Semantic Consistency
by: Li, Zihan, et al.
Published: (2025)
by: Li, Zihan, et al.
Published: (2025)
TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models
by: Tan, Xudong, et al.
Published: (2025)
by: Tan, Xudong, et al.
Published: (2025)
Dynamic Modality Scheduling for Multimodal Large Models via Confidence, Uncertainty, and Semantic Consistency
by: Tanaka, Hiroshi, et al.
Published: (2025)
by: Tanaka, Hiroshi, et al.
Published: (2025)
Preserving Cross-Modal Stability for Visual Unlearning in Multimodal Scenarios
by: Li, Jinghan Xu Yuyang Zhang Qixuan Cai Jiancheng Chen Keqiu
Published: (2025)
by: Li, Jinghan Xu Yuyang Zhang Qixuan Cai Jiancheng Chen Keqiu
Published: (2025)
MUVR: A Multi-Modal Untrimmed Video Retrieval Benchmark with Multi-Level Visual Correspondence
by: Feng, Yue, et al.
Published: (2025)
by: Feng, Yue, et al.
Published: (2025)
Mario: Multimodal Graph Reasoning with Large Language Models
by: Sun, Yuanfu, et al.
Published: (2026)
by: Sun, Yuanfu, et al.
Published: (2026)
Enhanced Motion Forecasting with Plug-and-Play Multimodal Large Language Models
by: Luo, Katie, et al.
Published: (2025)
by: Luo, Katie, et al.
Published: (2025)
Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models
by: Dong, Xinpeng, et al.
Published: (2026)
by: Dong, Xinpeng, et al.
Published: (2026)
The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio
by: Leng, Sicong, et al.
Published: (2024)
by: Leng, Sicong, et al.
Published: (2024)
Manipulating Multimodal Agents via Cross-Modal Prompt Injection
by: Wang, Le, et al.
Published: (2025)
by: Wang, Le, et al.
Published: (2025)
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval
by: Huang, Jinghao, et al.
Published: (2025)
by: Huang, Jinghao, et al.
Published: (2025)
Semantic-Enhanced Cross-Modal Place Recognition for Robust Robot Localization
by: Lin, Yujia, et al.
Published: (2025)
by: Lin, Yujia, et al.
Published: (2025)
CrossWeaver: Cross-modal Weaving for Arbitrary-Modality Semantic Segmentation
by: Zhang, Zelin, et al.
Published: (2026)
by: Zhang, Zelin, et al.
Published: (2026)
IDMR: Towards Instance-Driven Precise Visual Correspondence in Multimodal Retrieval
by: Liu, Bangwei, et al.
Published: (2025)
by: Liu, Bangwei, et al.
Published: (2025)
Similar Items
-
Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models
by: Zhang, Guosheng, et al.
Published: (2025) -
From Intuition to Investigation: A Tool-Augmented Reasoning MLLM Framework for Generalizable Face Anti-Spoofing
by: Zhang, Haoyuan, et al.
Published: (2026) -
ALoRE: Efficient Visual Adaptation via Aggregating Low Rank Experts
by: Du, Sinan, et al.
Published: (2024) -
Robust Multimodal Semantic Segmentation with Balanced Modality Contributions
by: Tan, Jiaqi, et al.
Published: (2025) -
MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning
by: Gao, Tianhong, et al.
Published: (2025)