Visual Late Chunking: An Empirical Study of Contextual Chunking for Efficient Visual Document Retrieval
Fuente:
arXiv
Guardado en:
| Autores principales: | Yan, Yibo, Ou, Mingdong, Cao, Yi, Huo, Jiahao, Zou, Xin, Liu, Shuliang, Kwok, James, Hu, Xuming |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Sculpting the Vector Space: Towards Efficient Multi-Vector Visual Document Retrieval via Prune-then-Merge Framework
por: Yan, Yibo, et al.
Publicado: (2026)
por: Yan, Yibo, et al.
Publicado: (2026)
Beyond the Grid: Layout-Informed Multi-Vector Retrieval with Parsed Visual Document Representations
por: Yan, Yibo, et al.
Publicado: (2026)
por: Yan, Yibo, et al.
Publicado: (2026)
DocPruner: A Storage-Efficient Framework for Multi-Vector Visual Document Retrieval via Adaptive Patch-Level Embedding Pruning
por: Yan, Yibo, et al.
Publicado: (2025)
por: Yan, Yibo, et al.
Publicado: (2025)
Unlocking Multimodal Document Intelligence: From Current Triumphs to Future Frontiers of Visual Document Retrieval
por: Yan, Yibo, et al.
Publicado: (2026)
por: Yan, Yibo, et al.
Publicado: (2026)
Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models
por: Günther, Michael, et al.
Publicado: (2024)
por: Günther, Michael, et al.
Publicado: (2024)
SmartChunk Retrieval: Query-Aware Chunk Compression with Planning for Efficient Document RAG
por: Zhang, Xuechen, et al.
Publicado: (2025)
por: Zhang, Xuechen, et al.
Publicado: (2025)
Beyond Chunking: Discourse-Aware Hierarchical Retrieval for Long Document Question Answering
por: Chen, Huiyao, et al.
Publicado: (2025)
por: Chen, Huiyao, et al.
Publicado: (2025)
Chunking, Retrieval, and Re-ranking: An Empirical Evaluation of RAG Architectures for Policy Document Question Answering
por: Maharjan, Anuj, et al.
Publicado: (2026)
por: Maharjan, Anuj, et al.
Publicado: (2026)
Cross-Document Topic-Aligned Chunking for Retrieval-Augmented Generation
por: Stankovic, Mile
Publicado: (2025)
por: Stankovic, Mile
Publicado: (2025)
Query-Adaptive Semantic Chunking for Retrieval-Augmented Generation: A Dynamic Strategy with Contextual Window Expansion
por: Rastogi, Mudit
Publicado: (2026)
por: Rastogi, Mudit
Publicado: (2026)
Adaptive Chunking: Optimizing Chunking-Method Selection for RAG
por: Júnior, Paulo Roberto de Moura, et al.
Publicado: (2026)
por: Júnior, Paulo Roberto de Moura, et al.
Publicado: (2026)
Structure-Aware Chunking for Tabular Data in Retrieval-Augmented Generation
por: Guttal, Pooja, et al.
Publicado: (2026)
por: Guttal, Pooja, et al.
Publicado: (2026)
Empirical Evaluation of PDF Parsing and Chunking for Financial Question Answering with RAG
por: Bachyr, Omar El, et al.
Publicado: (2026)
por: Bachyr, Omar El, et al.
Publicado: (2026)
The Chronicles of RAG: The Retriever, the Chunk and the Generator
por: Finardi, Paulo, et al.
Publicado: (2024)
por: Finardi, Paulo, et al.
Publicado: (2024)
Grounding Language Model with Chunking-Free In-Context Retrieval
por: Qian, Hongjin, et al.
Publicado: (2024)
por: Qian, Hongjin, et al.
Publicado: (2024)
Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings
por: Ma, Yubo, et al.
Publicado: (2025)
por: Ma, Yubo, et al.
Publicado: (2025)
Is Semantic Chunking Worth the Computational Cost?
por: Qu, Renyi, et al.
Publicado: (2024)
por: Qu, Renyi, et al.
Publicado: (2024)
ChunkNorris: A High-Performance and Low-Energy Approach to PDF Parsing and Chunking
por: Ciancone, Mathieu, et al.
Publicado: (2025)
por: Ciancone, Mathieu, et al.
Publicado: (2025)
Attention Grounded Enhancement for Visual Document Retrieval
por: Cui, Wanqing, et al.
Publicado: (2025)
por: Cui, Wanqing, et al.
Publicado: (2025)
Beyond Chunk-Then-Embed: A Comprehensive Taxonomy and Evaluation of Document Chunking Strategies for Information Retrieval
por: Zhou, Yongjie, et al.
Publicado: (2026)
por: Zhou, Yongjie, et al.
Publicado: (2026)
CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding
por: Huo, Jiahao, et al.
Publicado: (2026)
por: Huo, Jiahao, et al.
Publicado: (2026)
Reconstructing Context: Evaluating Advanced Chunking Strategies for Retrieval-Augmented Generation
por: Merola, Carlo, et al.
Publicado: (2025)
por: Merola, Carlo, et al.
Publicado: (2025)
TreeHop: Generate and Filter Next Query Embeddings Efficiently for Multi-hop Question Answering
por: Li, Zhonghao, et al.
Publicado: (2025)
por: Li, Zhonghao, et al.
Publicado: (2025)
Toward General Semantic Chunking: A Discriminative Framework for Ultra-Long Documents
por: Wu, Kaifeng, et al.
Publicado: (2025)
por: Wu, Kaifeng, et al.
Publicado: (2025)
A Systematic Analysis of Chunking Strategies for Reliable Question Answering
por: Bennani, Sofia, et al.
Publicado: (2026)
por: Bennani, Sofia, et al.
Publicado: (2026)
Graph-Aware Late Chunking for Retrieval-Augmented Generation in Biomedical Literature
por: Mortezaagha, Pouria, et al.
Publicado: (2026)
por: Mortezaagha, Pouria, et al.
Publicado: (2026)
S2 Chunking: A Hybrid Framework for Document Segmentation Through Integrated Spatial and Semantic Analysis
por: Verma, Prashant
Publicado: (2025)
por: Verma, Prashant
Publicado: (2025)
Rethinking Chunk Size For Long-Document Retrieval: A Multi-Dataset Analysis
por: Bhat, Sinchana Ramakanth, et al.
Publicado: (2025)
por: Bhat, Sinchana Ramakanth, et al.
Publicado: (2025)
cAST: Enhancing Code Retrieval-Augmented Generation with Structural Chunking via Abstract Syntax Tree
por: Zhang, Yilin, et al.
Publicado: (2025)
por: Zhang, Yilin, et al.
Publicado: (2025)
A Picture is Worth a Thousand Words? An Empirical Study of Aggregation Strategies for Visual Financial Document Retrieval
por: Lim, Ho Hung, et al.
Publicado: (2026)
por: Lim, Ho Hung, et al.
Publicado: (2026)
Evaluating Chunking Strategies For Retrieval-Augmented Generation in Oil and Gas Enterprise Documents
por: Taiwo, Samuel, et al.
Publicado: (2026)
por: Taiwo, Samuel, et al.
Publicado: (2026)
A Survey on Proactive Defense Strategies Against Misinformation in Large Language Models
por: Liu, Shuliang, et al.
Publicado: (2025)
por: Liu, Shuliang, et al.
Publicado: (2025)
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval
por: Wan, David, et al.
Publicado: (2025)
por: Wan, David, et al.
Publicado: (2025)
Roles of MLLMs in Visually Rich Document Retrieval for RAG: A Survey
por: Zhang, Xiantao
Publicado: (2025)
por: Zhang, Xiantao
Publicado: (2025)
Structural Anchor Pruning: Training-Free Multi-Vector Compression for Visual Document Retrieval
por: Liu, Zhuchenyang, et al.
Publicado: (2026)
por: Liu, Zhuchenyang, et al.
Publicado: (2026)
Identity-Decoupled Anonymization for Visual Evidence in Multi-modal Retrieval-Augmented Generation
por: Cheng, Zehua, et al.
Publicado: (2026)
por: Cheng, Zehua, et al.
Publicado: (2026)
Agentic Chunking and Bayesian De-chunking of AI Generated Fuzzy Cognitive Maps: A Model of the Thucydides Trap
por: Panda, Akash Kumar, et al.
Publicado: (2026)
por: Panda, Akash Kumar, et al.
Publicado: (2026)
Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval
por: Reddy, Arun, et al.
Publicado: (2025)
por: Reddy, Arun, et al.
Publicado: (2025)
Web Retrieval-Aware Chunking (W-RAC) for Efficient and Cost-Effective Retrieval-Augmented Generation Systems
por: Allu, Uday, et al.
Publicado: (2026)
por: Allu, Uday, et al.
Publicado: (2026)
LITTA: Late-Interaction and Test-Time Alignment for Visually-Grounded Multimodal Retrieval
por: Kim, Seonok
Publicado: (2026)
por: Kim, Seonok
Publicado: (2026)
Ejemplares similares
-
Sculpting the Vector Space: Towards Efficient Multi-Vector Visual Document Retrieval via Prune-then-Merge Framework
por: Yan, Yibo, et al.
Publicado: (2026) -
Beyond the Grid: Layout-Informed Multi-Vector Retrieval with Parsed Visual Document Representations
por: Yan, Yibo, et al.
Publicado: (2026) -
DocPruner: A Storage-Efficient Framework for Multi-Vector Visual Document Retrieval via Adaptive Patch-Level Embedding Pruning
por: Yan, Yibo, et al.
Publicado: (2025) -
Unlocking Multimodal Document Intelligence: From Current Triumphs to Future Frontiers of Visual Document Retrieval
por: Yan, Yibo, et al.
Publicado: (2026) -
Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models
por: Günther, Michael, et al.
Publicado: (2024)