Unifying Multimodal Retrieval via Document Screenshot Embedding
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Xueguang, Lin, Sheng-Chieh, Li, Minghan, Chen, Wenhu, Lin, Jimmy |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Document Screenshot Retrievers are Vulnerable to Pixel Poisoning Attacks
by: Zhuang, Shengyao, et al.
Published: (2025)
by: Zhuang, Shengyao, et al.
Published: (2025)
VISA: Retrieval Augmented Generation with Visual Source Attribution
by: Ma, Xueguang, et al.
Published: (2024)
by: Ma, Xueguang, et al.
Published: (2024)
Tevatron 2.0: Unified Document Retrieval Toolkit across Scale, Language, and Modality
by: Ma, Xueguang, et al.
Published: (2025)
by: Ma, Xueguang, et al.
Published: (2025)
MAGMaR Shared Task System Description: Video Retrieval with OmniEmbed
by: Zhan, Jiaqi Samantha, et al.
Published: (2025)
by: Zhan, Jiaqi Samantha, et al.
Published: (2025)
PromptReps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document Retrieval
by: Zhuang, Shengyao, et al.
Published: (2024)
by: Zhuang, Shengyao, et al.
Published: (2024)
UniRAG: Universal Retrieval Augmentation for Large Vision Language Models
by: Sharifymoghaddam, Sahel, et al.
Published: (2024)
by: Sharifymoghaddam, Sahel, et al.
Published: (2024)
Rank-R1: Enhancing Reasoning in LLM-based Document Rerankers via Reinforcement Learning
by: Zhuang, Shengyao, et al.
Published: (2025)
by: Zhuang, Shengyao, et al.
Published: (2025)
MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs
by: Lin, Sheng-Chieh, et al.
Published: (2024)
by: Lin, Sheng-Chieh, et al.
Published: (2024)
Hard Negatives, Hard Lessons: Revisiting Training Data Quality for Robust Information Retrieval with LLMs
by: Thakur, Nandan, et al.
Published: (2025)
by: Thakur, Nandan, et al.
Published: (2025)
DRAMA: Diverse Augmentation from Large Language Models to Smaller Dense Retrievers
by: Ma, Xueguang, et al.
Published: (2025)
by: Ma, Xueguang, et al.
Published: (2025)
ORBIT: Scalable and Verifiable Data Generation for Search Agents on a Tight Budget
by: Thakur, Nandan, et al.
Published: (2026)
by: Thakur, Nandan, et al.
Published: (2026)
Efficient Long-Document Reranking via Block-Level Embeddings and Top-k Interaction Refinement
by: Li, Minghan, et al.
Published: (2025)
by: Li, Minghan, et al.
Published: (2025)
LACONIC: Dense-Level Effectiveness for Scalable Sparse Retrieval via a Two-Phase Training Curriculum
by: Xu, Zhichao, et al.
Published: (2026)
by: Xu, Zhichao, et al.
Published: (2026)
Operational Advice for Dense and Sparse Retrievers: HNSW, Flat, or Inverted Indexes?
by: Lin, Jimmy
Published: (2024)
by: Lin, Jimmy
Published: (2024)
Domain Adaptation for Dense Retrieval and Conversational Dense Retrieval through Self-Supervision by Meticulous Pseudo-Relevance Labeling
by: Li, Minghan, et al.
Published: (2024)
by: Li, Minghan, et al.
Published: (2024)
Illusions of Relevance: Arbitrary Content Injection Attacks Deceive Retrievers, Rerankers, and LLM Judges
by: Tamber, Manveer Singh, et al.
Published: (2025)
by: Tamber, Manveer Singh, et al.
Published: (2025)
Synergistic Approach for Simultaneous Optimization of Monolingual, Cross-lingual, and Multilingual Information Retrieval
by: Elmahdy, Adel, et al.
Published: (2024)
by: Elmahdy, Adel, et al.
Published: (2024)
A Survey of Long-Document Retrieval in the PLM and LLM Era
by: Li, Minghan, et al.
Published: (2025)
by: Li, Minghan, et al.
Published: (2025)
Can't Hide Behind the API: Stealing Black-Box Commercial Embedding Models
by: Tamber, Manveer Singh, et al.
Published: (2024)
by: Tamber, Manveer Singh, et al.
Published: (2024)
KAP: MLLM-assisted OCR Text Enhancement for Hybrid Retrieval in Chinese Non-Narrative Documents
by: Hsu, Hsin-Ling, et al.
Published: (2025)
by: Hsu, Hsin-Ling, et al.
Published: (2025)
Categorizing Social Media Screenshots for Identifying Author Misattribution
by: Farris, Ashlyn M., et al.
Published: (2024)
by: Farris, Ashlyn M., et al.
Published: (2024)
InteraRec: Screenshot Based Recommendations Using Multimodal Large Language Models
by: Karra, Saketh Reddy, et al.
Published: (2024)
by: Karra, Saketh Reddy, et al.
Published: (2024)
Unified Multimodal and Multilingual Retrieval via Multi-Task Learning with NLU Integration
by: Zhang, Xinyuan, et al.
Published: (2026)
by: Zhang, Xinyuan, et al.
Published: (2026)
Retrieval-Feedback-Driven Distillation and Preference Alignment for Efficient LLM-based Query Expansion
by: Li, Minghan, et al.
Published: (2026)
by: Li, Minghan, et al.
Published: (2026)
AnnoRetrieve: Efficient Structured Retrieval for Unstructured Document Analysis
by: Lin, Teng, et al.
Published: (2026)
by: Lin, Teng, et al.
Published: (2026)
EviRerank: Adaptive Evidence Construction for Long-Document LLM Reranking
by: Li, Minghan, et al.
Published: (2024)
by: Li, Minghan, et al.
Published: (2024)
Study on LLMs for Promptagator-Style Dense Retriever Training
by: Gwon, Daniel, et al.
Published: (2025)
by: Gwon, Daniel, et al.
Published: (2025)
Unified Multimodal Interleaved Document Representation for Retrieval
by: Lee, Jaewoo, et al.
Published: (2024)
by: Lee, Jaewoo, et al.
Published: (2024)
Teaching Dense Retrieval Models to Specialize with Listwise Distillation and LLM Data Augmentation
by: Tamber, Manveer Singh, et al.
Published: (2025)
by: Tamber, Manveer Singh, et al.
Published: (2025)
ReasonEmbed: Enhanced Text Embeddings for Reasoning-Intensive Document Retrieval
by: Chen, Jianlyu, et al.
Published: (2025)
by: Chen, Jianlyu, et al.
Published: (2025)
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval
by: Hu, Ruofan, et al.
Published: (2025)
by: Hu, Ruofan, et al.
Published: (2025)
Conventional Contrastive Learning Often Falls Short: Improving Dense Retrieval with Cross-Encoder Listwise Distillation and Synthetic Data
by: Tamber, Manveer Singh, et al.
Published: (2025)
by: Tamber, Manveer Singh, et al.
Published: (2025)
Unifying Adversarial Robustness and Training Across Text Scoring Models
by: Tamber, Manveer Singh, et al.
Published: (2026)
by: Tamber, Manveer Singh, et al.
Published: (2026)
Revisiting Feedback Models for HyDE
by: Jedidi, Nour, et al.
Published: (2025)
by: Jedidi, Nour, et al.
Published: (2025)
Rerank Before You Reason: Analyzing Reranking Tradeoffs through Effective Token Cost in Deep Search Agents
by: Sharifymoghaddam, Sahel, et al.
Published: (2026)
by: Sharifymoghaddam, Sahel, et al.
Published: (2026)
A Unified Model and Document Representation for On-Device Retrieval-Augmented Generation
by: Killingback, Julian, et al.
Published: (2026)
by: Killingback, Julian, et al.
Published: (2026)
Synergizing Implicit and Explicit User Interests: A Multi-Embedding Retrieval Framework at Pinterest
by: Fan, Zhibo, et al.
Published: (2025)
by: Fan, Zhibo, et al.
Published: (2025)
Overview of the TREC 2025 Retrieval Augmented Generation (RAG) Track
by: Upadhyay, Shivani, et al.
Published: (2026)
by: Upadhyay, Shivani, et al.
Published: (2026)
Systematic Evaluation of Neural Retrieval Models on the Touché 2020 Argument Retrieval Subset of BEIR
by: Thakur, Nandan, et al.
Published: (2024)
by: Thakur, Nandan, et al.
Published: (2024)
DocMMIR: A Framework for Document Multi-modal Information Retrieval
by: Li, Zirui, et al.
Published: (2025)
by: Li, Zirui, et al.
Published: (2025)
Similar Items
-
Document Screenshot Retrievers are Vulnerable to Pixel Poisoning Attacks
by: Zhuang, Shengyao, et al.
Published: (2025) -
VISA: Retrieval Augmented Generation with Visual Source Attribution
by: Ma, Xueguang, et al.
Published: (2024) -
Tevatron 2.0: Unified Document Retrieval Toolkit across Scale, Language, and Modality
by: Ma, Xueguang, et al.
Published: (2025) -
MAGMaR Shared Task System Description: Video Retrieval with OmniEmbed
by: Zhan, Jiaqi Samantha, et al.
Published: (2025) -
PromptReps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document Retrieval
by: Zhuang, Shengyao, et al.
Published: (2024)