Attention Grounded Enhancement for Visual Document Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Cui, Wanqing, Huang, Wei, Guo, Yazhi, Hu, Yibo, Jin, Meiguang, Ma, Junfeng, Bi, Keping |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Do LLM-Generated Texts Impact Term-Based Retrieval Models?
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
Visual Late Chunking: An Empirical Study of Contextual Chunking for Efficient Visual Document Retrieval
by: Yan, Yibo, et al.
Published: (2026)
by: Yan, Yibo, et al.
Published: (2026)
Injecting External Knowledge into the Reasoning Process Enhances Retrieval-Augmented Generation
by: Tang, Minghao, et al.
Published: (2025)
by: Tang, Minghao, et al.
Published: (2025)
Sculpting the Vector Space: Towards Efficient Multi-Vector Visual Document Retrieval via Prune-then-Merge Framework
by: Yan, Yibo, et al.
Published: (2026)
by: Yan, Yibo, et al.
Published: (2026)
Understanding Parametric Knowledge Injection in Retrieval-Augmented Generation
by: Tang, Minghao, et al.
Published: (2025)
by: Tang, Minghao, et al.
Published: (2025)
Beyond the Grid: Layout-Informed Multi-Vector Retrieval with Parsed Visual Document Representations
by: Yan, Yibo, et al.
Published: (2026)
by: Yan, Yibo, et al.
Published: (2026)
A Comparative Study of Specialized LLMs as Dense Retrievers
by: Zhang, Hengran, et al.
Published: (2025)
by: Zhang, Hengran, et al.
Published: (2025)
Structural Anchor Pruning: Training-Free Multi-Vector Compression for Visual Document Retrieval
by: Liu, Zhuchenyang, et al.
Published: (2026)
by: Liu, Zhuchenyang, et al.
Published: (2026)
Unlocking Multimodal Document Intelligence: From Current Triumphs to Future Frontiers of Visual Document Retrieval
by: Yan, Yibo, et al.
Published: (2026)
by: Yan, Yibo, et al.
Published: (2026)
DocPruner: A Storage-Efficient Framework for Multi-Vector Visual Document Retrieval via Adaptive Patch-Level Embedding Pruning
by: Yan, Yibo, et al.
Published: (2025)
by: Yan, Yibo, et al.
Published: (2025)
Beyond Relevance: Utility-Centric Retrieval in the LLM Era
by: Zhang, Hengran, et al.
Published: (2026)
by: Zhang, Hengran, et al.
Published: (2026)
Utility-Focused LLM Annotation for Retrieval and Retrieval-Augmented Generation
by: Zhang, Hengran, et al.
Published: (2025)
by: Zhang, Hengran, et al.
Published: (2025)
Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering
by: Dong, Kuicai, et al.
Published: (2025)
by: Dong, Kuicai, et al.
Published: (2025)
LLM-Specific Utility: A New Perspective for Retrieval-Augmented Generation
by: Zhang, Hengran, et al.
Published: (2025)
by: Zhang, Hengran, et al.
Published: (2025)
Chain of Evidence: Pixel-Level Visual Attribution for Iterative Retrieval-Augmented Generation
by: Liu, Peiyang, et al.
Published: (2026)
by: Liu, Peiyang, et al.
Published: (2026)
Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search Engines
by: Long, Xinwei, et al.
Published: (2025)
by: Long, Xinwei, et al.
Published: (2025)
LITTA: Late-Interaction and Test-Time Alignment for Visually-Grounded Multimodal Retrieval
by: Kim, Seonok
Published: (2026)
by: Kim, Seonok
Published: (2026)
Towards Natural Language-Based Document Image Retrieval: New Dataset and Benchmark
by: Guo, Hao, et al.
Published: (2025)
by: Guo, Hao, et al.
Published: (2025)
Unleashing the Power of LLMs in Dense Retrieval with Query Likelihood Modeling
by: Zhang, Hengran, et al.
Published: (2025)
by: Zhang, Hengran, et al.
Published: (2025)
Visual Lifelog Retrieval through Captioning-Enhanced Interpretation
by: Shih, Yu-Fei, et al.
Published: (2025)
by: Shih, Yu-Fei, et al.
Published: (2025)
VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
by: Tanaka, Ryota, et al.
Published: (2025)
by: Tanaka, Ryota, et al.
Published: (2025)
Are We on the Right Way for Assessing Document Retrieval-Augmented Generation?
by: Shen, Wenxuan, et al.
Published: (2025)
by: Shen, Wenxuan, et al.
Published: (2025)
ColPali: Efficient Document Retrieval with Vision Language Models
by: Faysse, Manuel, et al.
Published: (2024)
by: Faysse, Manuel, et al.
Published: (2024)
VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval
by: Zhou, Junjie, et al.
Published: (2024)
by: Zhou, Junjie, et al.
Published: (2024)
LiveThinking: Enabling Real-Time Efficient Reasoning for AI-Powered Livestreaming via Reinforcement Learning
by: Sun, Yuhan, et al.
Published: (2025)
by: Sun, Yuhan, et al.
Published: (2025)
Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings
by: Ma, Yubo, et al.
Published: (2025)
by: Ma, Yubo, et al.
Published: (2025)
Roles of MLLMs in Visually Rich Document Retrieval for RAG: A Survey
by: Zhang, Xiantao
Published: (2025)
by: Zhang, Xiantao
Published: (2025)
Index Light, Reason Deep: Deferred Visual Ingestion for Visual-Dense Document Question Answering
by: Xu, Tao
Published: (2026)
by: Xu, Tao
Published: (2026)
ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents
by: Wang, Qiuchen, et al.
Published: (2025)
by: Wang, Qiuchen, et al.
Published: (2025)
One Pic is All it Takes: Poisoning Visual Document Retrieval Augmented Generation with a Single Image
by: Shereen, Ezzeldin, et al.
Published: (2025)
by: Shereen, Ezzeldin, et al.
Published: (2025)
VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents
by: Yu, Shi, et al.
Published: (2024)
by: Yu, Shi, et al.
Published: (2024)
An Iterative Utility Judgment Framework Inspired by Philosophical Relevance via LLMs
by: Zhang, Hengran, et al.
Published: (2024)
by: Zhang, Hengran, et al.
Published: (2024)
Distilling a Small Utility-Based Passage Selector to Enhance Retrieval-Augmented Generation
by: Zhang, Hengran, et al.
Published: (2025)
by: Zhang, Hengran, et al.
Published: (2025)
Analysis of Plan-based Retrieval for Grounded Text Generation
by: Godbole, Ameya, et al.
Published: (2024)
by: Godbole, Ameya, et al.
Published: (2024)
CoRank: LLM-Based Compact Reranking with Document Features for Scientific Retrieval
by: Tian, Runchu, et al.
Published: (2025)
by: Tian, Runchu, et al.
Published: (2025)
Seeing Through the MiRAGE: Evaluating Multimodal Retrieval Augmented Generation
by: Martin, Alexander, et al.
Published: (2025)
by: Martin, Alexander, et al.
Published: (2025)
Utility-Oriented Visual Evidence Selection for Multimodal Retrieval-Augmented Generation
by: Luo, Weiqing, et al.
Published: (2026)
by: Luo, Weiqing, et al.
Published: (2026)
MMDocIR: Benchmarking Multimodal Retrieval for Long Documents
by: Dong, Kuicai, et al.
Published: (2025)
by: Dong, Kuicai, et al.
Published: (2025)
MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction
by: Xiao, Zilin, et al.
Published: (2025)
by: Xiao, Zilin, et al.
Published: (2025)
Bagging-Based Model Merging for Robust General Text Embeddings
by: Zhang, Hengran, et al.
Published: (2026)
by: Zhang, Hengran, et al.
Published: (2026)
Similar Items
-
How Do LLM-Generated Texts Impact Term-Based Retrieval Models?
by: Huang, Wei, et al.
Published: (2025) -
Visual Late Chunking: An Empirical Study of Contextual Chunking for Efficient Visual Document Retrieval
by: Yan, Yibo, et al.
Published: (2026) -
Injecting External Knowledge into the Reasoning Process Enhances Retrieval-Augmented Generation
by: Tang, Minghao, et al.
Published: (2025) -
Sculpting the Vector Space: Towards Efficient Multi-Vector Visual Document Retrieval via Prune-then-Merge Framework
by: Yan, Yibo, et al.
Published: (2026) -
Understanding Parametric Knowledge Injection in Retrieval-Augmented Generation
by: Tang, Minghao, et al.
Published: (2025)