Beyond Chunk-Then-Embed: A Comprehensive Taxonomy and Evaluation of Document Chunking Strategies for Information Retrieval
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Yongjie, Wang, Shuai, Koopman, Bevan, Zuccon, Guido |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
2D Matryoshka Training for Information Retrieval
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
Team IELAB at TREC Clinical Trial Track 2023: Enhancing Clinical Trial Retrieval with Neural Rankers and Large Language Models
von: Zhuang, Shengyao, et al.
Veröffentlicht: (2024)
von: Zhuang, Shengyao, et al.
Veröffentlicht: (2024)
Starbucks-v2: Improved Training for 2D Matryoshka Embeddings
von: Zhuang, Shengyao, et al.
Veröffentlicht: (2024)
von: Zhuang, Shengyao, et al.
Veröffentlicht: (2024)
AutoBool: An Reinforcement-Learning trained LLM for Effective Automated Boolean Query Generation for Systematic Reviews
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
A Reproducibility Study of Goldilocks: Just-Right Tuning of BERT for TAR
von: Mao, Xinyu, et al.
Veröffentlicht: (2024)
von: Mao, Xinyu, et al.
Veröffentlicht: (2024)
Dense Retrieval with Continuous Explicit Feedback for Systematic Review Screening Prioritisation
von: Mao, Xinyu, et al.
Veröffentlicht: (2024)
von: Mao, Xinyu, et al.
Veröffentlicht: (2024)
Does Vec2Text Pose a New Corpus Poisoning Threat?
von: Zhuang, Shengyao, et al.
Veröffentlicht: (2024)
von: Zhuang, Shengyao, et al.
Veröffentlicht: (2024)
PromptReps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document Retrieval
von: Zhuang, Shengyao, et al.
Veröffentlicht: (2024)
von: Zhuang, Shengyao, et al.
Veröffentlicht: (2024)
Can It Reach the Generator? Investigating the Survival of Prompt-Injection Attacks in Realistic RAG Settings
von: Yin, Yu, et al.
Veröffentlicht: (2026)
von: Yin, Yu, et al.
Veröffentlicht: (2026)
ReSLLM: Large Language Models are Strong Resource Selectors for Federated Search
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
Reassessing Large Language Model Boolean Query Generation for Systematic Reviews
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
Pseudo Relevance Feedback is Enough to Close the Gap Between Small and Large Dense Retrieval Models
von: Li, Hang, et al.
Veröffentlicht: (2025)
von: Li, Hang, et al.
Veröffentlicht: (2025)
Document Screenshot Retrievers are Vulnerable to Pixel Poisoning Attacks
von: Zhuang, Shengyao, et al.
Veröffentlicht: (2025)
von: Zhuang, Shengyao, et al.
Veröffentlicht: (2025)
Understanding and Mitigating the Threat of Vec2Text to Dense Retrieval Systems
von: Zhuang, Shengyao, et al.
Veröffentlicht: (2024)
von: Zhuang, Shengyao, et al.
Veröffentlicht: (2024)
TPRF: A Transformer-based Pseudo-Relevance Feedback Model for Efficient and Effective Retrieval
von: Li, Hang, et al.
Veröffentlicht: (2024)
von: Li, Hang, et al.
Veröffentlicht: (2024)
VISA: Retrieval Augmented Generation with Visual Source Attribution
von: Ma, Xueguang, et al.
Veröffentlicht: (2024)
von: Ma, Xueguang, et al.
Veröffentlicht: (2024)
A Setwise Approach for Effective and Highly Efficient Zero-shot Ranking with Large Language Models
von: Zhuang, Shengyao, et al.
Veröffentlicht: (2023)
von: Zhuang, Shengyao, et al.
Veröffentlicht: (2023)
LLM-VPRF: Large Language Model Based Vector Pseudo Relevance Feedback
von: Li, Hang, et al.
Veröffentlicht: (2025)
von: Li, Hang, et al.
Veröffentlicht: (2025)
Rank-R1: Enhancing Reasoning in LLM-based Document Rerankers via Reinforcement Learning
von: Zhuang, Shengyao, et al.
Veröffentlicht: (2025)
von: Zhuang, Shengyao, et al.
Veröffentlicht: (2025)
Whole-Pool Setwise Reranking with Long-Context Language Models
von: Li, Hang, et al.
Veröffentlicht: (2026)
von: Li, Hang, et al.
Veröffentlicht: (2026)
Zero-shot Generative Large Language Models for Systematic Review Screening Automation
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
Evaluating Chunking Strategies For Retrieval-Augmented Generation in Oil and Gas Enterprise Documents
von: Taiwo, Samuel, et al.
Veröffentlicht: (2026)
von: Taiwo, Samuel, et al.
Veröffentlicht: (2026)
Rethinking Chunk Size For Long-Document Retrieval: A Multi-Dataset Analysis
von: Bhat, Sinchana Ramakanth, et al.
Veröffentlicht: (2025)
von: Bhat, Sinchana Ramakanth, et al.
Veröffentlicht: (2025)
On the impact of retrieved content representations in RAG Pipelines
von: Ross, Jonathan J, et al.
Veröffentlicht: (2026)
von: Ross, Jonathan J, et al.
Veröffentlicht: (2026)
Beyond Chunking: Discourse-Aware Hierarchical Retrieval for Long Document Question Answering
von: Chen, Huiyao, et al.
Veröffentlicht: (2025)
von: Chen, Huiyao, et al.
Veröffentlicht: (2025)
Visual Late Chunking: An Empirical Study of Contextual Chunking for Efficient Visual Document Retrieval
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
von: Yan, Yibo, et al.
Veröffentlicht: (2026)
SmartChunk Retrieval: Query-Aware Chunk Compression with Planning for Efficient Document RAG
von: Zhang, Xuechen, et al.
Veröffentlicht: (2025)
von: Zhang, Xuechen, et al.
Veröffentlicht: (2025)
Pre-training vs. Fine-tuning: A Reproducibility Study on Dense Retrieval Knowledge Acquisition
von: Yao, Zheng, et al.
Veröffentlicht: (2025)
von: Yao, Zheng, et al.
Veröffentlicht: (2025)
Reconstructing Context: Evaluating Advanced Chunking Strategies for Retrieval-Augmented Generation
von: Merola, Carlo, et al.
Veröffentlicht: (2025)
von: Merola, Carlo, et al.
Veröffentlicht: (2025)
FeB4RAG: Evaluating Federated Search in the Context of Retrieval Augmented Generation
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models
von: Günther, Michael, et al.
Veröffentlicht: (2024)
von: Günther, Michael, et al.
Veröffentlicht: (2024)
Beyond Chunks and Graphs: Retrieval-Augmented Generation through Triplet-Driven Thinking
von: Gong, Shengbo, et al.
Veröffentlicht: (2025)
von: Gong, Shengbo, et al.
Veröffentlicht: (2025)
Cross-Document Topic-Aligned Chunking for Retrieval-Augmented Generation
von: Stankovic, Mile
Veröffentlicht: (2025)
von: Stankovic, Mile
Veröffentlicht: (2025)
Grounding Language Model with Chunking-Free In-Context Retrieval
von: Qian, Hongjin, et al.
Veröffentlicht: (2024)
von: Qian, Hongjin, et al.
Veröffentlicht: (2024)
Adaptive Chunking: Optimizing Chunking-Method Selection for RAG
von: Júnior, Paulo Roberto de Moura, et al.
Veröffentlicht: (2026)
von: Júnior, Paulo Roberto de Moura, et al.
Veröffentlicht: (2026)
Rank-DistiLLM: Closing the Effectiveness Gap Between Cross-Encoders and LLMs for Passage Re-Ranking
von: Schlatt, Ferdinand, et al.
Veröffentlicht: (2024)
von: Schlatt, Ferdinand, et al.
Veröffentlicht: (2024)
Set-Encoder: Permutation-Invariant Inter-Passage Attention for Listwise Passage Re-Ranking with Cross-Encoders
von: Schlatt, Ferdinand, et al.
Veröffentlicht: (2024)
von: Schlatt, Ferdinand, et al.
Veröffentlicht: (2024)
Chunking, Retrieval, and Re-ranking: An Empirical Evaluation of RAG Architectures for Policy Document Question Answering
von: Maharjan, Anuj, et al.
Veröffentlicht: (2026)
von: Maharjan, Anuj, et al.
Veröffentlicht: (2026)
A Systematic Analysis of Chunking Strategies for Reliable Question Answering
von: Bennani, Sofia, et al.
Veröffentlicht: (2026)
von: Bennani, Sofia, et al.
Veröffentlicht: (2026)
Structure-Aware Chunking for Tabular Data in Retrieval-Augmented Generation
von: Guttal, Pooja, et al.
Veröffentlicht: (2026)
von: Guttal, Pooja, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
2D Matryoshka Training for Information Retrieval
von: Wang, Shuai, et al.
Veröffentlicht: (2024) -
Team IELAB at TREC Clinical Trial Track 2023: Enhancing Clinical Trial Retrieval with Neural Rankers and Large Language Models
von: Zhuang, Shengyao, et al.
Veröffentlicht: (2024) -
Starbucks-v2: Improved Training for 2D Matryoshka Embeddings
von: Zhuang, Shengyao, et al.
Veröffentlicht: (2024) -
AutoBool: An Reinforcement-Learning trained LLM for Effective Automated Boolean Query Generation for Systematic Reviews
von: Wang, Shuai, et al.
Veröffentlicht: (2025) -
A Reproducibility Study of Goldilocks: Just-Right Tuning of BERT for TAR
von: Mao, Xinyu, et al.
Veröffentlicht: (2024)