U-MARVEL: Unveiling Key Factors for Universal Multimodal Retrieval via Embedding Learning with MLLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Xiaojie, Li, Chu, Chen, Shi-Zhe, Chen, Xi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
von: Kong, Fanheng, et al.
Veröffentlicht: (2025)
von: Kong, Fanheng, et al.
Veröffentlicht: (2025)
Adapting MLLMs for Nuanced Video Retrieval
von: Bagad, Piyush, et al.
Veröffentlicht: (2025)
von: Bagad, Piyush, et al.
Veröffentlicht: (2025)
Embedding-based Retrieval in Multimodal Content Moderation
von: Liang, Hanzhong, et al.
Veröffentlicht: (2025)
von: Liang, Hanzhong, et al.
Veröffentlicht: (2025)
Accurate and Scalable Multimodal Pathology Retrieval via Attentive Vision-Language Alignment
von: Wang, Hongyi, et al.
Veröffentlicht: (2025)
von: Wang, Hongyi, et al.
Veröffentlicht: (2025)
Towards Universal Video Retrieval: Generalizing Video Embedding via Synthesized Multimodal Pyramid Curriculum
von: Guo, Zhuoning, et al.
Veröffentlicht: (2025)
von: Guo, Zhuoning, et al.
Veröffentlicht: (2025)
MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction
von: Xiao, Zilin, et al.
Veröffentlicht: (2025)
von: Xiao, Zilin, et al.
Veröffentlicht: (2025)
BRIDGE: Multimodal-to-Text Retrieval via Reinforcement-Learned Query Alignment
von: Mounis, Mohamed Darwish, et al.
Veröffentlicht: (2026)
von: Mounis, Mohamed Darwish, et al.
Veröffentlicht: (2026)
UniNote: A Unified Embedding Model for Multimodal Representation and Ranking
von: Zhao, Jinghan, et al.
Veröffentlicht: (2026)
von: Zhao, Jinghan, et al.
Veröffentlicht: (2026)
VLM2GeoVec: Toward Universal Multimodal Embeddings for Remote Sensing
von: Aimar, Emanuel Sánchez, et al.
Veröffentlicht: (2025)
von: Aimar, Emanuel Sánchez, et al.
Veröffentlicht: (2025)
Multimodal Learned Sparse Retrieval with Probabilistic Expansion Control
von: Nguyen, Thong, et al.
Veröffentlicht: (2024)
von: Nguyen, Thong, et al.
Veröffentlicht: (2024)
Snap and Diagnose: An Advanced Multimodal Retrieval System for Identifying Plant Diseases in the Wild
von: Wei, Tianqi, et al.
Veröffentlicht: (2024)
von: Wei, Tianqi, et al.
Veröffentlicht: (2024)
MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs
von: Lin, Sheng-Chieh, et al.
Veröffentlicht: (2024)
von: Lin, Sheng-Chieh, et al.
Veröffentlicht: (2024)
VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval
von: Zhou, Junjie, et al.
Veröffentlicht: (2024)
von: Zhou, Junjie, et al.
Veröffentlicht: (2024)
Composed Multi-modal Retrieval: A Survey of Approaches and Applications
von: Zhang, Kun, et al.
Veröffentlicht: (2025)
von: Zhang, Kun, et al.
Veröffentlicht: (2025)
Compressing then Matching: An Efficient Pre-training Paradigm for Multimodal Embedding
von: Li, Da, et al.
Veröffentlicht: (2025)
von: Li, Da, et al.
Veröffentlicht: (2025)
Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval
von: Tu, Rong-Cheng, et al.
Veröffentlicht: (2025)
von: Tu, Rong-Cheng, et al.
Veröffentlicht: (2025)
Enhancing Subsequent Video Retrieval via Vision-Language Models (VLMs)
von: Duan, Yicheng, et al.
Veröffentlicht: (2025)
von: Duan, Yicheng, et al.
Veröffentlicht: (2025)
Beyond Global Similarity: Towards Fine-Grained, Multi-Condition Multimodal Retrieval
von: Lu, Xuan, et al.
Veröffentlicht: (2026)
von: Lu, Xuan, et al.
Veröffentlicht: (2026)
SAIL-Embedding Technical Report: Omni-modal Embedding Foundation Model
von: Lin, Lin, et al.
Veröffentlicht: (2025)
von: Lin, Lin, et al.
Veröffentlicht: (2025)
E5-V: Universal Embeddings with Multimodal Large Language Models
von: Jiang, Ting, et al.
Veröffentlicht: (2024)
von: Jiang, Ting, et al.
Veröffentlicht: (2024)
MR$^2$-Bench: Going Beyond Matching to Reasoning in Multimodal Retrieval
von: Zhou, Junjie, et al.
Veröffentlicht: (2025)
von: Zhou, Junjie, et al.
Veröffentlicht: (2025)
Unified Interactive Multimodal Moment Retrieval via Cascaded Embedding-Reranking and Temporal-Aware Score Fusion
von: Thanh, Toan Le Ngo, et al.
Veröffentlicht: (2025)
von: Thanh, Toan Le Ngo, et al.
Veröffentlicht: (2025)
A Resource-Efficient Training Framework for Remote Sensing Text--Image Retrieval
von: Zhang, Weihang, et al.
Veröffentlicht: (2025)
von: Zhang, Weihang, et al.
Veröffentlicht: (2025)
Understanding Long Videos via LLM-Powered Entity Relation Graphs
von: Chu, Meng, et al.
Veröffentlicht: (2025)
von: Chu, Meng, et al.
Veröffentlicht: (2025)
HLFormer: Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning
von: Li, Jun, et al.
Veröffentlicht: (2025)
von: Li, Jun, et al.
Veröffentlicht: (2025)
Any2Any: Incomplete Multimodal Retrieval with Conformal Prediction
von: Li, Po-han, et al.
Veröffentlicht: (2024)
von: Li, Po-han, et al.
Veröffentlicht: (2024)
Scale Up Composed Image Retrieval Learning via Modification Text Generation
von: Zhou, Yinan, et al.
Veröffentlicht: (2025)
von: Zhou, Yinan, et al.
Veröffentlicht: (2025)
Decoding Ancient Oracle Bone Script via Generative Dictionary Retrieval
von: Wu, Yin, et al.
Veröffentlicht: (2026)
von: Wu, Yin, et al.
Veröffentlicht: (2026)
LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts
von: Cai, Qifeng, et al.
Veröffentlicht: (2025)
von: Cai, Qifeng, et al.
Veröffentlicht: (2025)
ReinPool: Reinforcement Learning Pooling Multi-Vector Embeddings for Retrieval System
von: Cha, Sungguk, et al.
Veröffentlicht: (2026)
von: Cha, Sungguk, et al.
Veröffentlicht: (2026)
Active Learning via Classifier Impact and Greedy Selection for Interactive Image Retrieval
von: Bar, Leah, et al.
Veröffentlicht: (2024)
von: Bar, Leah, et al.
Veröffentlicht: (2024)
MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation
von: Hsiao, Chi-Hsiang, et al.
Veröffentlicht: (2025)
von: Hsiao, Chi-Hsiang, et al.
Veröffentlicht: (2025)
DeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Search
von: Narayan, Kartik, et al.
Veröffentlicht: (2025)
von: Narayan, Kartik, et al.
Veröffentlicht: (2025)
Entity Image and Mixed-Modal Image Retrieval Datasets
von: Blaga, Cristian-Ioan, et al.
Veröffentlicht: (2025)
von: Blaga, Cristian-Ioan, et al.
Veröffentlicht: (2025)
DeepImageSearch: Benchmarking Multimodal Agents for Context-Aware Image Retrieval in Visual Histories
von: Deng, Chenlong, et al.
Veröffentlicht: (2026)
von: Deng, Chenlong, et al.
Veröffentlicht: (2026)
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval
von: Huang, Jinghao, et al.
Veröffentlicht: (2025)
von: Huang, Jinghao, et al.
Veröffentlicht: (2025)
Chain-of-Thought Re-ranking for Image Retrieval Tasks
von: Wu, Shangrong, et al.
Veröffentlicht: (2025)
von: Wu, Shangrong, et al.
Veröffentlicht: (2025)
Modality and Task Adaptation for Enhanced Zero-shot Composed Image Retrieval
von: Li, Haiwen, et al.
Veröffentlicht: (2024)
von: Li, Haiwen, et al.
Veröffentlicht: (2024)
FashionStylist: An Expert Knowledge-enhanced Multimodal Dataset for Fashion Understanding
von: Feng, Kaidong, et al.
Veröffentlicht: (2026)
von: Feng, Kaidong, et al.
Veröffentlicht: (2026)
A Survey of Multimodal Composite Editing and Retrieval
von: Li, Suyan, et al.
Veröffentlicht: (2024)
von: Li, Suyan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
von: Kong, Fanheng, et al.
Veröffentlicht: (2025) -
Adapting MLLMs for Nuanced Video Retrieval
von: Bagad, Piyush, et al.
Veröffentlicht: (2025) -
Embedding-based Retrieval in Multimodal Content Moderation
von: Liang, Hanzhong, et al.
Veröffentlicht: (2025) -
Accurate and Scalable Multimodal Pathology Retrieval via Attentive Vision-Language Alignment
von: Wang, Hongyi, et al.
Veröffentlicht: (2025) -
Towards Universal Video Retrieval: Generalizing Video Embedding via Synthesized Multimodal Pyramid Curriculum
von: Guo, Zhuoning, et al.
Veröffentlicht: (2025)