DocPruner: A Storage-Efficient Framework for Multi-Vector Visual Document Retrieval via Adaptive Patch-Level Embedding Pruning
Fuente:
arXiv
Saved in:
| Main Authors: | Yan, Yibo, Xu, Guangwei, Zou, Xin, Liu, Shuliang, Kwok, James, Hu, Xuming |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sculpting the Vector Space: Towards Efficient Multi-Vector Visual Document Retrieval via Prune-then-Merge Framework
by: Yan, Yibo, et al.
Published: (2026)
by: Yan, Yibo, et al.
Published: (2026)
Beyond the Grid: Layout-Informed Multi-Vector Retrieval with Parsed Visual Document Representations
by: Yan, Yibo, et al.
Published: (2026)
by: Yan, Yibo, et al.
Published: (2026)
Visual Late Chunking: An Empirical Study of Contextual Chunking for Efficient Visual Document Retrieval
by: Yan, Yibo, et al.
Published: (2026)
by: Yan, Yibo, et al.
Published: (2026)
Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings
by: Ma, Yubo, et al.
Published: (2025)
by: Ma, Yubo, et al.
Published: (2025)
Unlocking Multimodal Document Intelligence: From Current Triumphs to Future Frontiers of Visual Document Retrieval
by: Yan, Yibo, et al.
Published: (2026)
by: Yan, Yibo, et al.
Published: (2026)
Hierarchical Patch Compression for ColPali: Efficient Multi-Vector Document Retrieval with Dynamic Pruning and Quantization
by: Bach, Duong
Published: (2025)
by: Bach, Duong
Published: (2025)
TreeHop: Generate and Filter Next Query Embeddings Efficiently for Multi-hop Question Answering
by: Li, Zhonghao, et al.
Published: (2025)
by: Li, Zhonghao, et al.
Published: (2025)
DocMMIR: A Framework for Document Multi-modal Information Retrieval
by: Li, Zirui, et al.
Published: (2025)
by: Li, Zirui, et al.
Published: (2025)
LLM-Augmented Retrieval: Enhancing Retrieval Models Through Language Models and Doc-Level Embedding
by: Wu, Mingrui, et al.
Published: (2024)
by: Wu, Mingrui, et al.
Published: (2024)
Structural Anchor Pruning: Training-Free Multi-Vector Compression for Visual Document Retrieval
by: Liu, Zhuchenyang, et al.
Published: (2026)
by: Liu, Zhuchenyang, et al.
Published: (2026)
DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive Benchmark
by: Hu, Ruofan, et al.
Published: (2026)
by: Hu, Ruofan, et al.
Published: (2026)
SpatCode: Rotary-based Unified Encoding Framework for Efficient Spatiotemporal Vector Retrieval
by: Hu, Bingde, et al.
Published: (2026)
by: Hu, Bingde, et al.
Published: (2026)
DocReLM: Mastering Document Retrieval with Language Model
by: Wei, Gengchen, et al.
Published: (2024)
by: Wei, Gengchen, et al.
Published: (2024)
Attention Grounded Enhancement for Visual Document Retrieval
by: Cui, Wanqing, et al.
Published: (2025)
by: Cui, Wanqing, et al.
Published: (2025)
Poly-Vector Retrieval: Reference and Content Embeddings for Legal Documents
by: Lima, João Alberto de Oliveira
Published: (2025)
by: Lima, João Alberto de Oliveira
Published: (2025)
Nemotron ColEmbed V2: Top-Performing Late Interaction Embedding Models for Visual Document Retrieval
by: Moreira, Gabriel de Souza P., et al.
Published: (2026)
by: Moreira, Gabriel de Souza P., et al.
Published: (2026)
A Survey on Proactive Defense Strategies Against Misinformation in Large Language Models
by: Liu, Shuliang, et al.
Published: (2025)
by: Liu, Shuliang, et al.
Published: (2025)
Efficient Multi-Vector Dense Retrieval Using Bit Vectors
by: Nardini, Franco Maria, et al.
Published: (2024)
by: Nardini, Franco Maria, et al.
Published: (2024)
Hybrid-Vector Retrieval for Visually Rich Documents: Combining Single-Vector Efficiency and Multi-Vector Accuracy
by: Kim, Juyeon, et al.
Published: (2025)
by: Kim, Juyeon, et al.
Published: (2025)
WARP: An Efficient Engine for Multi-Vector Retrieval
by: Scheerer, Jan Luca, et al.
Published: (2025)
by: Scheerer, Jan Luca, et al.
Published: (2025)
VectorSearch: Enhancing Document Retrieval with Semantic Embeddings and Optimized Search
by: Monir, Solmaz Seyed, et al.
Published: (2024)
by: Monir, Solmaz Seyed, et al.
Published: (2024)
Efficient Recommendation with Millions of Items by Dynamic Pruning of Sub-Item Embeddings
by: Petrov, Aleksandr V., et al.
Published: (2025)
by: Petrov, Aleksandr V., et al.
Published: (2025)
Semantic Certainty Assessment in Vector Retrieval Systems: A Novel Framework for Embedding Quality Evaluation
by: Du, Y.
Published: (2025)
by: Du, Y.
Published: (2025)
Unifying Multimodal Retrieval via Document Screenshot Embedding
by: Ma, Xueguang, et al.
Published: (2024)
by: Ma, Xueguang, et al.
Published: (2024)
Efficient Long-Document Reranking via Block-Level Embeddings and Top-k Interaction Refinement
by: Li, Minghan, et al.
Published: (2025)
by: Li, Minghan, et al.
Published: (2025)
Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
by: Qin, Jialong, et al.
Published: (2025)
by: Qin, Jialong, et al.
Published: (2025)
DocQAC: Adaptive Trie-Guided Decoding for Effective In-Document Query Auto-Completion
by: Mehta, Rahul, et al.
Published: (2026)
by: Mehta, Rahul, et al.
Published: (2026)
CRISP: Clustering Multi-Vector Representations for Denoising and Pruning
by: Veneroso, João, et al.
Published: (2025)
by: Veneroso, João, et al.
Published: (2025)
Efficient Constant-Space Multi-Vector Retrieval
by: MacAvaney, Sean, et al.
Published: (2025)
by: MacAvaney, Sean, et al.
Published: (2025)
AnnoRetrieve: Efficient Structured Retrieval for Unstructured Document Analysis
by: Lin, Teng, et al.
Published: (2026)
by: Lin, Teng, et al.
Published: (2026)
ReasonEmbed: Enhanced Text Embeddings for Reasoning-Intensive Document Retrieval
by: Chen, Jianlyu, et al.
Published: (2025)
by: Chen, Jianlyu, et al.
Published: (2025)
PruneRAG: Confidence-Guided Query Decomposition Trees for Efficient Retrieval-Augmented Generation
by: Jiao, Shuguang, et al.
Published: (2026)
by: Jiao, Shuguang, et al.
Published: (2026)
Synergizing Implicit and Explicit User Interests: A Multi-Embedding Retrieval Framework at Pinterest
by: Fan, Zhibo, et al.
Published: (2025)
by: Fan, Zhibo, et al.
Published: (2025)
ModernVBERT: Towards Smaller Visual Document Retrievers
by: Teiletche, Paul, et al.
Published: (2025)
by: Teiletche, Paul, et al.
Published: (2025)
DiVA-DocRE: A Discriminative and Voice-Aware Paradigm for Document-Level Relation Extraction
by: Wu, Yiheng, et al.
Published: (2024)
by: Wu, Yiheng, et al.
Published: (2024)
Transform Before You Query: A Privacy-Preserving Approach for Vector Retrieval with Embedding Space Alignment
by: He, Ruiqi, et al.
Published: (2025)
by: He, Ruiqi, et al.
Published: (2025)
KBest: Efficient Vector Search on Kunpeng CPU
by: Ma, Kaihao, et al.
Published: (2025)
by: Ma, Kaihao, et al.
Published: (2025)
A Multi-Granularity Retrieval Framework for Visually-Rich Documents
by: Xu, Mingjun, et al.
Published: (2025)
by: Xu, Mingjun, et al.
Published: (2025)
ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained Alignment
by: Yang, Hao, et al.
Published: (2026)
by: Yang, Hao, et al.
Published: (2026)
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval
by: Hu, Ruofan, et al.
Published: (2025)
by: Hu, Ruofan, et al.
Published: (2025)
Similar Items
-
Sculpting the Vector Space: Towards Efficient Multi-Vector Visual Document Retrieval via Prune-then-Merge Framework
by: Yan, Yibo, et al.
Published: (2026) -
Beyond the Grid: Layout-Informed Multi-Vector Retrieval with Parsed Visual Document Representations
by: Yan, Yibo, et al.
Published: (2026) -
Visual Late Chunking: An Empirical Study of Contextual Chunking for Efficient Visual Document Retrieval
by: Yan, Yibo, et al.
Published: (2026) -
Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings
by: Ma, Yubo, et al.
Published: (2025) -
Unlocking Multimodal Document Intelligence: From Current Triumphs to Future Frontiers of Visual Document Retrieval
by: Yan, Yibo, et al.
Published: (2026)