Closing the Modality Gap for Mixed Modality Search
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Binxu, Zhang, Yuhui, Wang, Xiaohan, Liang, Weixin, Schmidt, Ludwig, Yeung-Levy, Serena |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
by: Wang, Xiaohan, et al.
Published: (2024)
by: Wang, Xiaohan, et al.
Published: (2024)
Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data
by: Zhang, Yuhui, et al.
Published: (2024)
by: Zhang, Yuhui, et al.
Published: (2024)
Smart Multi-Modal Search: Contextual Sparse and Dense Embedding Integration in Adobe Express
by: Aroraa, Cherag, et al.
Published: (2024)
by: Aroraa, Cherag, et al.
Published: (2024)
Why are Visually-Grounded Language Models Bad at Image Classification?
by: Zhang, Yuhui, et al.
Published: (2024)
by: Zhang, Yuhui, et al.
Published: (2024)
It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap
by: Fahim, Abrar, et al.
Published: (2024)
by: Fahim, Abrar, et al.
Published: (2024)
UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities
by: Yeo, Woongyeong, et al.
Published: (2025)
by: Yeo, Woongyeong, et al.
Published: (2025)
TelcoAI: Advancing 3GPP Technical Specification Search through Agentic Multi-Modal Retrieval-Augmented Generation
by: Ghosh, Rahul, et al.
Published: (2025)
by: Ghosh, Rahul, et al.
Published: (2025)
Modality-Aware Integration with Large Language Models for Knowledge-based Visual Question Answering
by: Dong, Junnan, et al.
Published: (2024)
by: Dong, Junnan, et al.
Published: (2024)
Data or Language Supervision: What Makes CLIP Better than DINO?
by: Liu, Yiming, et al.
Published: (2025)
by: Liu, Yiming, et al.
Published: (2025)
Reversed in Time: A Novel Temporal-Emphasized Benchmark for Cross-Modal Video-Text Retrieval
by: Du, Yang, et al.
Published: (2024)
by: Du, Yang, et al.
Published: (2024)
Attribute-Aware Implicit Modality Alignment for Text Attribute Person Search
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines
by: Zhang, Zhixin, et al.
Published: (2024)
by: Zhang, Zhixin, et al.
Published: (2024)
PaperSearchQA: Learning to Search and Reason over Scientific Papers with RLVR
by: Burgess, James, et al.
Published: (2026)
by: Burgess, James, et al.
Published: (2026)
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
by: Li, Yongqi, et al.
Published: (2024)
by: Li, Yongqi, et al.
Published: (2024)
Pailitao-VL: Unified Embedding and Reranker for Real-Time Multi-Modal Industrial Search
by: Chen, Lei, et al.
Published: (2026)
by: Chen, Lei, et al.
Published: (2026)
Knowledge Graphs Meet Multi-Modal Learning: A Comprehensive Survey
by: Chen, Zhuo, et al.
Published: (2024)
by: Chen, Zhuo, et al.
Published: (2024)
Temporal Preference Optimization for Long-Form Video Understanding
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
EasyEdit: An Easy-to-use Knowledge Editing Framework for Large Language Models
by: Wang, Peng, et al.
Published: (2023)
by: Wang, Peng, et al.
Published: (2023)
MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines
by: Jiang, Dongzhi, et al.
Published: (2024)
by: Jiang, Dongzhi, et al.
Published: (2024)
MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding
by: Nie, Zhanheng, et al.
Published: (2025)
by: Nie, Zhanheng, et al.
Published: (2025)
OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal Retrieval
by: Yang, Wei, et al.
Published: (2025)
by: Yang, Wei, et al.
Published: (2025)
NegVQA: Can Vision Language Models Understand Negation?
by: Zhang, Yuhui, et al.
Published: (2025)
by: Zhang, Yuhui, et al.
Published: (2025)
Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity
by: Liang, Weixin, et al.
Published: (2025)
by: Liang, Weixin, et al.
Published: (2025)
Benchmarking Robustness of Contrastive Learning Models for Medical Image-Report Retrieval
by: Deanda, Demetrio, et al.
Published: (2025)
by: Deanda, Demetrio, et al.
Published: (2025)
Transductive Visual Programming: Evolving Tool Libraries from Experience for Spatial Reasoning
by: Wu, Shengguang, et al.
Published: (2025)
by: Wu, Shengguang, et al.
Published: (2025)
Compressible and Searchable: AI-native Multi-Modal Retrieval System with Learned Image Compression
by: Luo, Jixiang
Published: (2024)
by: Luo, Jixiang
Published: (2024)
Cross-Modal Retrieval with Cauchy-Schwarz Divergence
by: Zhang, Jiahao, et al.
Published: (2025)
by: Zhang, Jiahao, et al.
Published: (2025)
Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training
by: Wang, Mengru, et al.
Published: (2025)
by: Wang, Mengru, et al.
Published: (2025)
Multimedia-Aware Question Answering: A Review of Retrieval and Cross-Modal Reasoning Architectures
by: Raja, Rahul, et al.
Published: (2025)
by: Raja, Rahul, et al.
Published: (2025)
ModalChorus: Visual Probing and Alignment of Multi-modal Embeddings via Modal Fusion Map
by: Ye, Yilin, et al.
Published: (2024)
by: Ye, Yilin, et al.
Published: (2024)
Multi-Vector Index Compression in Any Modality
by: Qin, Hanxiang, et al.
Published: (2026)
by: Qin, Hanxiang, et al.
Published: (2026)
Efficient and High-Fidelity Omni Modality Retrieval
by: Huynh, Chuong, et al.
Published: (2026)
by: Huynh, Chuong, et al.
Published: (2026)
Chain of Evidence: Pixel-Level Visual Attribution for Iterative Retrieval-Augmented Generation
by: Liu, Peiyang, et al.
Published: (2026)
by: Liu, Peiyang, et al.
Published: (2026)
$M^3EL$: A Multi-task Multi-topic Dataset for Multi-modal Entity Linking
by: Wang, Fang, et al.
Published: (2024)
by: Wang, Fang, et al.
Published: (2024)
PC$^2$: Pseudo-Classification Based Pseudo-Captioning for Noisy Correspondence Learning in Cross-Modal Retrieval
by: Duan, Yue, et al.
Published: (2024)
by: Duan, Yue, et al.
Published: (2024)
Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation
by: Zhang, Yuhui, et al.
Published: (2025)
by: Zhang, Yuhui, et al.
Published: (2025)
Evaluating VisualRAG: Quantifying Cross-Modal Performance in Enterprise Document Understanding
by: Mannam, Varun, et al.
Published: (2025)
by: Mannam, Varun, et al.
Published: (2025)
FM2DS: Few-Shot Multimodal Multihop Data Synthesis with Knowledge Distillation for Question Answering
by: Abaskohi, Amirhossein, et al.
Published: (2024)
by: Abaskohi, Amirhossein, et al.
Published: (2024)
Retrieval-augmented Prompt Learning for Pre-trained Foundation Models
by: Chen, Xiang, et al.
Published: (2025)
by: Chen, Xiang, et al.
Published: (2025)
Evaluation of Embedding-Based and Generative Methods for LLM-Driven Document Classification: Opportunities and Challenges
by: Lu, Rong, et al.
Published: (2026)
by: Lu, Rong, et al.
Published: (2026)
Similar Items
-
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
by: Wang, Xiaohan, et al.
Published: (2024) -
Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data
by: Zhang, Yuhui, et al.
Published: (2024) -
Smart Multi-Modal Search: Contextual Sparse and Dense Embedding Integration in Adobe Express
by: Aroraa, Cherag, et al.
Published: (2024) -
Why are Visually-Grounded Language Models Bad at Image Classification?
by: Zhang, Yuhui, et al.
Published: (2024) -
It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap
by: Fahim, Abrar, et al.
Published: (2024)