MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Sheng-Chieh, Lee, Chankyu, Shoeybi, Mohammad, Lin, Jimmy, Catanzaro, Bryan, Ping, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models
von: Lee, Chankyu, et al.
Veröffentlicht: (2024)
von: Lee, Chankyu, et al.
Veröffentlicht: (2024)
MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction
von: Xiao, Zilin, et al.
Veröffentlicht: (2025)
von: Xiao, Zilin, et al.
Veröffentlicht: (2025)
E5-V: Universal Embeddings with Multimodal Large Language Models
von: Jiang, Ting, et al.
Veröffentlicht: (2024)
von: Jiang, Ting, et al.
Veröffentlicht: (2024)
Toward Automatic Relevance Judgment using Vision--Language Models for Image--Text Retrieval Evaluation
von: Yang, Jheng-Hong, et al.
Veröffentlicht: (2024)
von: Yang, Jheng-Hong, et al.
Veröffentlicht: (2024)
MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation
von: Hsiao, Chi-Hsiang, et al.
Veröffentlicht: (2025)
von: Hsiao, Chi-Hsiang, et al.
Veröffentlicht: (2025)
Windsock is Dancing: Adaptive Multimodal Retrieval-Augmented Generation
von: Zhao, Shu, et al.
Veröffentlicht: (2025)
von: Zhao, Shu, et al.
Veröffentlicht: (2025)
ChatQA: Surpassing GPT-4 on Conversational QA and RAG
von: Liu, Zihan, et al.
Veröffentlicht: (2024)
von: Liu, Zihan, et al.
Veröffentlicht: (2024)
VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval
von: Zhou, Junjie, et al.
Veröffentlicht: (2024)
von: Zhou, Junjie, et al.
Veröffentlicht: (2024)
Indexing Multimodal Language Models for Large-scale Image Retrieval
von: Tharwat, Bahey, et al.
Veröffentlicht: (2026)
von: Tharwat, Bahey, et al.
Veröffentlicht: (2026)
Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval
von: Wan, David, et al.
Veröffentlicht: (2025)
von: Wan, David, et al.
Veröffentlicht: (2025)
Seeing Through the MiRAGE: Evaluating Multimodal Retrieval Augmented Generation
von: Martin, Alexander, et al.
Veröffentlicht: (2025)
von: Martin, Alexander, et al.
Veröffentlicht: (2025)
M3DR: Towards Universal Multilingual Multimodal Document Retrieval
von: Kolavi, Adithya S, et al.
Veröffentlicht: (2025)
von: Kolavi, Adithya S, et al.
Veröffentlicht: (2025)
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
von: Kong, Fanheng, et al.
Veröffentlicht: (2025)
von: Kong, Fanheng, et al.
Veröffentlicht: (2025)
U-MARVEL: Unveiling Key Factors for Universal Multimodal Retrieval via Embedding Learning with MLLMs
von: Li, Xiaojie, et al.
Veröffentlicht: (2025)
von: Li, Xiaojie, et al.
Veröffentlicht: (2025)
Towards Universal Video Retrieval: Generalizing Video Embedding via Synthesized Multimodal Pyramid Curriculum
von: Guo, Zhuoning, et al.
Veröffentlicht: (2025)
von: Guo, Zhuoning, et al.
Veröffentlicht: (2025)
Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture
von: Zhang, Longxiang, et al.
Veröffentlicht: (2026)
von: Zhang, Longxiang, et al.
Veröffentlicht: (2026)
ReinPool: Reinforcement Learning Pooling Multi-Vector Embeddings for Retrieval System
von: Cha, Sungguk, et al.
Veröffentlicht: (2026)
von: Cha, Sungguk, et al.
Veröffentlicht: (2026)
Large Language Model Informed Patent Image Retrieval
von: Lo, Hao-Cheng, et al.
Veröffentlicht: (2024)
von: Lo, Hao-Cheng, et al.
Veröffentlicht: (2024)
Progressive Multimodal Reasoning via Active Retrieval
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
MMDocIR: Benchmarking Multimodal Retrieval for Long Documents
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning
von: Yeo, Woongyeong, et al.
Veröffentlicht: (2025)
von: Yeo, Woongyeong, et al.
Veröffentlicht: (2025)
X-Reflect: Cross-Reflection Prompting for Multimodal Recommendation
von: Lyu, Hanjia, et al.
Veröffentlicht: (2024)
von: Lyu, Hanjia, et al.
Veröffentlicht: (2024)
VLM2GeoVec: Toward Universal Multimodal Embeddings for Remote Sensing
von: Aimar, Emanuel Sánchez, et al.
Veröffentlicht: (2025)
von: Aimar, Emanuel Sánchez, et al.
Veröffentlicht: (2025)
Embedding-based Retrieval in Multimodal Content Moderation
von: Liang, Hanzhong, et al.
Veröffentlicht: (2025)
von: Liang, Hanzhong, et al.
Veröffentlicht: (2025)
InterLV-Search: Benchmarking Interleaved Multimodal Agentic Search
von: Hou, Bohan, et al.
Veröffentlicht: (2026)
von: Hou, Bohan, et al.
Veröffentlicht: (2026)
Accurate and Scalable Multimodal Pathology Retrieval via Attentive Vision-Language Alignment
von: Wang, Hongyi, et al.
Veröffentlicht: (2025)
von: Wang, Hongyi, et al.
Veröffentlicht: (2025)
LITTA: Late-Interaction and Test-Time Alignment for Visually-Grounded Multimodal Retrieval
von: Kim, Seonok
Veröffentlicht: (2026)
von: Kim, Seonok
Veröffentlicht: (2026)
Are We on the Right Way for Assessing Document Retrieval-Augmented Generation?
von: Shen, Wenxuan, et al.
Veröffentlicht: (2025)
von: Shen, Wenxuan, et al.
Veröffentlicht: (2025)
MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory
von: Guo, Minghao, et al.
Veröffentlicht: (2026)
von: Guo, Minghao, et al.
Veröffentlicht: (2026)
Supervised Fine-Tuning or Contrastive Learning? Towards Better Multimodal LLM Reranking
von: Dai, Ziqi, et al.
Veröffentlicht: (2025)
von: Dai, Ziqi, et al.
Veröffentlicht: (2025)
RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs
von: Yu, Yue, et al.
Veröffentlicht: (2024)
von: Yu, Yue, et al.
Veröffentlicht: (2024)
Prompt-Guided Attention Head Selection for Focus-Oriented Image Retrieval
von: Nozawa, Yuji, et al.
Veröffentlicht: (2025)
von: Nozawa, Yuji, et al.
Veröffentlicht: (2025)
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
MR$^2$-Bench: Going Beyond Matching to Reasoning in Multimodal Retrieval
von: Zhou, Junjie, et al.
Veröffentlicht: (2025)
von: Zhou, Junjie, et al.
Veröffentlicht: (2025)
Towards Natural Language-Based Document Image Retrieval: New Dataset and Benchmark
von: Guo, Hao, et al.
Veröffentlicht: (2025)
von: Guo, Hao, et al.
Veröffentlicht: (2025)
Compressing then Matching: An Efficient Pre-training Paradigm for Multimodal Embedding
von: Li, Da, et al.
Veröffentlicht: (2025)
von: Li, Da, et al.
Veröffentlicht: (2025)
XL-HeadTags: Leveraging Multimodal Retrieval Augmentation for the Multilingual Generation of News Headlines and Tags
von: Shohan, Faisal Tareque, et al.
Veröffentlicht: (2024)
von: Shohan, Faisal Tareque, et al.
Veröffentlicht: (2024)
CollEX -- A Multimodal Agentic RAG System Enabling Interactive Exploration of Scientific Collections
von: Schneider, Florian, et al.
Veröffentlicht: (2025)
von: Schneider, Florian, et al.
Veröffentlicht: (2025)
NVLM: Open Frontier-Class Multimodal LLMs
von: Dai, Wenliang, et al.
Veröffentlicht: (2024)
von: Dai, Wenliang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models
von: Lee, Chankyu, et al.
Veröffentlicht: (2024) -
MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction
von: Xiao, Zilin, et al.
Veröffentlicht: (2025) -
E5-V: Universal Embeddings with Multimodal Large Language Models
von: Jiang, Ting, et al.
Veröffentlicht: (2024) -
Toward Automatic Relevance Judgment using Vision--Language Models for Image--Text Retrieval Evaluation
von: Yang, Jheng-Hong, et al.
Veröffentlicht: (2024) -
MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation
von: Hsiao, Chi-Hsiang, et al.
Veröffentlicht: (2025)