Reasoning-Augmented Representations for Multimodal Retrieval
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Jianrui, Rajan, Anirudh Sundara, Han, Brandon, Lee, Soochahn, Ganguly, Sukanta, Lee, Yong Jae |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Your Embedding Model is SMARTer Than You Think
por: Zhang, Jianrui, et al.
Publicado: (2026)
por: Zhang, Jianrui, et al.
Publicado: (2026)
Open Multimodal Retrieval-Augmented Factual Image Generation
por: Tian, Yang, et al.
Publicado: (2025)
por: Tian, Yang, et al.
Publicado: (2025)
MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval
por: Ju, Yeong-Joon, et al.
Publicado: (2024)
por: Ju, Yeong-Joon, et al.
Publicado: (2024)
TabRAG: Improving Tabular Document Question Answering for Retrieval Augmented Generation via Structured Representations
por: Si, Jacob, et al.
Publicado: (2025)
por: Si, Jacob, et al.
Publicado: (2025)
MOON3.0: Reasoning-aware Multimodal Representation Learning for E-commerce Product Understanding
por: Wu, Junxian, et al.
Publicado: (2026)
por: Wu, Junxian, et al.
Publicado: (2026)
HV-Attack: Hierarchical Visual Attack for Multimodal Retrieval Augmented Generation
por: Luo, Linyin, et al.
Publicado: (2025)
por: Luo, Linyin, et al.
Publicado: (2025)
MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs
por: Lin, Sheng-Chieh, et al.
Publicado: (2024)
por: Lin, Sheng-Chieh, et al.
Publicado: (2024)
Utility-Oriented Visual Evidence Selection for Multimodal Retrieval-Augmented Generation
por: Luo, Weiqing, et al.
Publicado: (2026)
por: Luo, Weiqing, et al.
Publicado: (2026)
MOON Embedding: Multimodal Representation Learning for E-commerce Search Advertising
por: Fu, Chenghan, et al.
Publicado: (2025)
por: Fu, Chenghan, et al.
Publicado: (2025)
Progressive Multimodal Reasoning via Active Retrieval
por: Dong, Guanting, et al.
Publicado: (2024)
por: Dong, Guanting, et al.
Publicado: (2024)
Diffusion Augmented Retrieval: A Training-Free Approach to Interactive Text-to-Image Retrieval
por: Long, Zijun, et al.
Publicado: (2025)
por: Long, Zijun, et al.
Publicado: (2025)
MOON: Generative MLLM-based Multimodal Representation Learning for E-commerce Product Understanding
por: Zhang, Daoze, et al.
Publicado: (2025)
por: Zhang, Daoze, et al.
Publicado: (2025)
RAVEN: Multitask Retrieval Augmented Vision-Language Learning
por: Rao, Varun Nagaraj, et al.
Publicado: (2024)
por: Rao, Varun Nagaraj, et al.
Publicado: (2024)
MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding
por: Nie, Zhanheng, et al.
Publicado: (2025)
por: Nie, Zhanheng, et al.
Publicado: (2025)
MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation
por: Hsiao, Chi-Hsiang, et al.
Publicado: (2025)
por: Hsiao, Chi-Hsiang, et al.
Publicado: (2025)
Video Enriched Retrieval Augmented Generation Using Aligned Video Captions
por: Rosa, Kevin Dela
Publicado: (2024)
por: Rosa, Kevin Dela
Publicado: (2024)
VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos
por: Ren, Xubin, et al.
Publicado: (2025)
por: Ren, Xubin, et al.
Publicado: (2025)
CaLa: Complementary Association Learning for Augmenting Composed Image Retrieval
por: Jiang, Xintong, et al.
Publicado: (2024)
por: Jiang, Xintong, et al.
Publicado: (2024)
Smart Routing for Multimodal Video Retrieval: When to Search What
por: Rosa, Kevin Dela
Publicado: (2025)
por: Rosa, Kevin Dela
Publicado: (2025)
OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal Retrieval
por: Yang, Wei, et al.
Publicado: (2025)
por: Yang, Wei, et al.
Publicado: (2025)
Dreaming User Multimodal Representation Guided by The Platonic Representation Hypothesis for Micro-Video Recommendation
por: Lin, Chengzhi, et al.
Publicado: (2024)
por: Lin, Chengzhi, et al.
Publicado: (2024)
VideoRAG: Retrieval-Augmented Generation over Video Corpus
por: Jeong, Soyeong, et al.
Publicado: (2025)
por: Jeong, Soyeong, et al.
Publicado: (2025)
MMDocIR: Benchmarking Multimodal Retrieval for Long Documents
por: Dong, Kuicai, et al.
Publicado: (2025)
por: Dong, Kuicai, et al.
Publicado: (2025)
CRISP -- Clustering-Based Redundancy-Reduced Instance Sampling for Pathology Case Representation and Retrieval
por: Afzal, Zahra Rahimi, et al.
Publicado: (2026)
por: Afzal, Zahra Rahimi, et al.
Publicado: (2026)
ContextIQ: A Multimodal Expert-Based Video Retrieval System for Contextual Advertising
por: Chaubey, Ashutosh, et al.
Publicado: (2024)
por: Chaubey, Ashutosh, et al.
Publicado: (2024)
RE-TRIANGLE: Does TRIANGLE Enable Multimodal Alignment Beyond Cosine Similarity in Retrieval?
por: Ghosh, Arijit, et al.
Publicado: (2026)
por: Ghosh, Arijit, et al.
Publicado: (2026)
XL-HeadTags: Leveraging Multimodal Retrieval Augmentation for the Multilingual Generation of News Headlines and Tags
por: Shohan, Faisal Tareque, et al.
Publicado: (2024)
por: Shohan, Faisal Tareque, et al.
Publicado: (2024)
Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos
por: Zhang, Jianrui, et al.
Publicado: (2024)
por: Zhang, Jianrui, et al.
Publicado: (2024)
A Survey of Multimodal Composite Editing and Retrieval
por: Li, Suyan, et al.
Publicado: (2024)
por: Li, Suyan, et al.
Publicado: (2024)
VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings
por: Giahi, Ramin, et al.
Publicado: (2025)
por: Giahi, Ramin, et al.
Publicado: (2025)
Unified Interactive Multimodal Moment Retrieval via Cascaded Embedding-Reranking and Temporal-Aware Score Fusion
por: Thanh, Toan Le Ngo, et al.
Publicado: (2025)
por: Thanh, Toan Le Ngo, et al.
Publicado: (2025)
UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities
por: Yeo, Woongyeong, et al.
Publicado: (2025)
por: Yeo, Woongyeong, et al.
Publicado: (2025)
ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents
por: Wang, Qiuchen, et al.
Publicado: (2025)
por: Wang, Qiuchen, et al.
Publicado: (2025)
Read and Think: An Efficient Step-wise Multimodal Language Model for Document Understanding and Reasoning
por: Zhang, Jinxu
Publicado: (2024)
por: Zhang, Jinxu
Publicado: (2024)
Multimodal RAG Enhanced Visual Description
por: Jaiswal, Amit Kumar, et al.
Publicado: (2025)
por: Jaiswal, Amit Kumar, et al.
Publicado: (2025)
Stay-Positive: A Case for Ignoring Real Image Features in Fake Image Detection
por: Rajan, Anirudh Sundara, et al.
Publicado: (2025)
por: Rajan, Anirudh Sundara, et al.
Publicado: (2025)
Benchmarking Robustness of Contrastive Learning Models for Medical Image-Report Retrieval
por: Deanda, Demetrio, et al.
Publicado: (2025)
por: Deanda, Demetrio, et al.
Publicado: (2025)
HotelMatch-LLM: Joint Multi-Task Training of Small and Large Language Models for Efficient Multimodal Hotel Retrieval
por: Askari, Arian, et al.
Publicado: (2025)
por: Askari, Arian, et al.
Publicado: (2025)
Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation
por: Wen, Tiansheng, et al.
Publicado: (2025)
por: Wen, Tiansheng, et al.
Publicado: (2025)
GENIUS: A Generative Framework for Universal Multimodal Search
por: Kim, Sungyeon, et al.
Publicado: (2025)
por: Kim, Sungyeon, et al.
Publicado: (2025)
Ejemplares similares
-
Your Embedding Model is SMARTer Than You Think
por: Zhang, Jianrui, et al.
Publicado: (2026) -
Open Multimodal Retrieval-Augmented Factual Image Generation
por: Tian, Yang, et al.
Publicado: (2025) -
MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval
por: Ju, Yeong-Joon, et al.
Publicado: (2024) -
TabRAG: Improving Tabular Document Question Answering for Retrieval Augmented Generation via Structured Representations
por: Si, Jacob, et al.
Publicado: (2025) -
MOON3.0: Reasoning-aware Multimodal Representation Learning for E-commerce Product Understanding
por: Wu, Junxian, et al.
Publicado: (2026)