Evo-Retriever: LLM-Guided Curriculum Evolution with Viewpoint-Pathway Collaboration for Multimodal Document Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Weiqing, Guo, Jinyue, Wang, Yaqi, Xiao, Haiyang, Zhang, Yuewei, Liu, Guohua, Wang, Hao Henry |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CurEvo: Curriculum-Guided Self-Evolution for Video Understanding
by: Zeng, Guiyi, et al.
Published: (2026)
by: Zeng, Guiyi, et al.
Published: (2026)
Multimodal LLM Enhanced Cross-lingual Cross-modal Retrieval
by: Wang, Yabing, et al.
Published: (2024)
by: Wang, Yabing, et al.
Published: (2024)
TRACE: Task-Adaptive Reasoning and Representation Learning for Universal Multimodal Retrieval
by: Hao, Xiangzhao, et al.
Published: (2026)
by: Hao, Xiangzhao, et al.
Published: (2026)
Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval
by: Zhu, Lanyun, et al.
Published: (2025)
by: Zhu, Lanyun, et al.
Published: (2025)
Reasoning Guided Embeddings: Leveraging MLLM Reasoning for Improved Multimodal Retrieval
by: Liu, Chunxu, et al.
Published: (2025)
by: Liu, Chunxu, et al.
Published: (2025)
AffectAgent: Collaborative Multi-Agent Reasoning for Retrieval-Augmented Multimodal Emotion Recognition
by: Wang, Zeheng, et al.
Published: (2026)
by: Wang, Zeheng, et al.
Published: (2026)
Utility-Oriented Visual Evidence Selection for Multimodal Retrieval-Augmented Generation
by: Luo, Weiqing, et al.
Published: (2026)
by: Luo, Weiqing, et al.
Published: (2026)
Deep Image Clustering Based on Curriculum Learning and Density Information
by: Zheng, Haiyang, et al.
Published: (2026)
by: Zheng, Haiyang, et al.
Published: (2026)
Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
by: Gao, Sensen, et al.
Published: (2025)
by: Gao, Sensen, et al.
Published: (2025)
V-Retrver: Evidence-Driven Agentic Reasoning for Universal Multimodal Retrieval
by: Chen, Dongyang, et al.
Published: (2026)
by: Chen, Dongyang, et al.
Published: (2026)
Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval
by: Sun, Hao, et al.
Published: (2026)
by: Sun, Hao, et al.
Published: (2026)
EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions
by: Vo, Dinh-Khoi, et al.
Published: (2025)
by: Vo, Dinh-Khoi, et al.
Published: (2025)
Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering
by: Dong, Kuicai, et al.
Published: (2025)
by: Dong, Kuicai, et al.
Published: (2025)
The Worse The Better: Content-Aware Viewpoint Generation Network for Projection-related Point Cloud Quality Assessment
by: Su, Zhiyong, et al.
Published: (2025)
by: Su, Zhiyong, et al.
Published: (2025)
SOLAR: Self-supervised Joint Learning for Symmetric Multimodal Retrieval
by: Yang, Wenjie, et al.
Published: (2026)
by: Yang, Wenjie, et al.
Published: (2026)
Multimodal Data Storage and Retrieval for Embodied AI: A Survey
by: Lu, Yihao, et al.
Published: (2025)
by: Lu, Yihao, et al.
Published: (2025)
STiTch: Semantic Transition and Transportation in Collaboration for Training-Free Zero-Shot Composed Image Retrieval
by: Li, Miaoge, et al.
Published: (2026)
by: Li, Miaoge, et al.
Published: (2026)
Hypergraph Convolutional Network based Weakly Supervised Point Cloud Semantic Segmentation with Scene-Level Annotations
by: Lu, Zhuheng, et al.
Published: (2022)
by: Lu, Zhuheng, et al.
Published: (2022)
Fine-grained Metrics for Point Cloud Semantic Segmentation
by: Lu, Zhuheng, et al.
Published: (2024)
by: Lu, Zhuheng, et al.
Published: (2024)
Generative Giants, Retrieval Weaklings: Why do Multimodal Large Language Models Fail at Multimodal Retrieval?
by: Feng, Hengyi, et al.
Published: (2025)
by: Feng, Hengyi, et al.
Published: (2025)
Combating Visual Neglect and Semantic Drift in Large Multimodal Models for Enhanced Cross-Modal Retrieval
by: Zhang, Guosheng, et al.
Published: (2026)
by: Zhang, Guosheng, et al.
Published: (2026)
Partial Scene Text Retrieval
by: Wang, Hao, et al.
Published: (2024)
by: Wang, Hao, et al.
Published: (2024)
UniFGVC: Universal Training-Free Few-Shot Fine-Grained Vision Classification via Attribute-Aware Multimodal Retrieval
by: Guo, Hongyu, et al.
Published: (2025)
by: Guo, Hongyu, et al.
Published: (2025)
EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models
by: Huang, Shiqi, et al.
Published: (2026)
by: Huang, Shiqi, et al.
Published: (2026)
DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding
by: Wu, Hao, et al.
Published: (2024)
by: Wu, Hao, et al.
Published: (2024)
Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models
by: Li, Aaron Branson Cigres, et al.
Published: (2026)
by: Li, Aaron Branson Cigres, et al.
Published: (2026)
Hierarchy-Guided Multimodal Representation Learning for Taxonomic Inference
by: Ahmed, Sk Miraj, et al.
Published: (2026)
by: Ahmed, Sk Miraj, et al.
Published: (2026)
Proactive Reasoning-with-Retrieval Framework for Medical Multimodal Large Language Models
by: Wang, Lehan, et al.
Published: (2025)
by: Wang, Lehan, et al.
Published: (2025)
RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented Generation
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
ReMatch: Boosting Representation through Matching for Multimodal Retrieval
by: Liu, Qianying, et al.
Published: (2025)
by: Liu, Qianying, et al.
Published: (2025)
Retrieval Augmented Image Harmonization
by: Wang, Haolin, et al.
Published: (2024)
by: Wang, Haolin, et al.
Published: (2024)
Multimodal Hypothetical Summary for Retrieval-based Multi-image Question Answering
by: Li, Peize, et al.
Published: (2024)
by: Li, Peize, et al.
Published: (2024)
MultiHaystack: Benchmarking Multimodal Retrieval and Reasoning over 40K Images, Videos, and Documents
by: Xu, Dannong, et al.
Published: (2026)
by: Xu, Dannong, et al.
Published: (2026)
EvoDriveVLA: Evolving Driving VLA Models via Collaborative Perception-Planning Distillation
by: Cao, Jiajun, et al.
Published: (2026)
by: Cao, Jiajun, et al.
Published: (2026)
RAGAR: Retrieval Augmented Personalized Image Generation Guided by Recommendation
by: Ling, Run, et al.
Published: (2025)
by: Ling, Run, et al.
Published: (2025)
Referring Expression Instance Retrieval and A Strong End-to-End Baseline
by: Hao, Xiangzhao, et al.
Published: (2025)
by: Hao, Xiangzhao, et al.
Published: (2025)
ReCALL: Recalibrating Capability Degradation for MLLM-based Composed Image Retrieval
by: Yang, Tianyu, et al.
Published: (2026)
by: Yang, Tianyu, et al.
Published: (2026)
IRGen: Generative Modeling for Image Retrieval
by: Zhang, Yidan, et al.
Published: (2023)
by: Zhang, Yidan, et al.
Published: (2023)
Enhanced Cross-modal 3D Retrieval via Tri-modal Reconstruction
by: Ren, Junlong, et al.
Published: (2025)
by: Ren, Junlong, et al.
Published: (2025)
OpenDPR: Open-Vocabulary Change Detection via Vision-Centric Diffusion-Guided Prototype Retrieval for Remote Sensing Imagery
by: Guo, Qi, et al.
Published: (2026)
by: Guo, Qi, et al.
Published: (2026)
Similar Items
-
CurEvo: Curriculum-Guided Self-Evolution for Video Understanding
by: Zeng, Guiyi, et al.
Published: (2026) -
Multimodal LLM Enhanced Cross-lingual Cross-modal Retrieval
by: Wang, Yabing, et al.
Published: (2024) -
TRACE: Task-Adaptive Reasoning and Representation Learning for Universal Multimodal Retrieval
by: Hao, Xiangzhao, et al.
Published: (2026) -
Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval
by: Zhu, Lanyun, et al.
Published: (2025) -
Reasoning Guided Embeddings: Leveraging MLLM Reasoning for Improved Multimodal Retrieval
by: Liu, Chunxu, et al.
Published: (2025)