Toward General Semantic Chunking: A Discriminative Framework for Ultra-Long Documents
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Kaifeng, Wu, Junyan, Liu, Qiang, Zhang, Jiarui, Xu, Wen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Chunking: Discourse-Aware Hierarchical Retrieval for Long Document Question Answering
by: Chen, Huiyao, et al.
Published: (2025)
by: Chen, Huiyao, et al.
Published: (2025)
Cross-Document Topic-Aligned Chunking for Retrieval-Augmented Generation
by: Stankovic, Mile
Published: (2025)
by: Stankovic, Mile
Published: (2025)
Adaptive Chunking: Optimizing Chunking-Method Selection for RAG
by: Júnior, Paulo Roberto de Moura, et al.
Published: (2026)
by: Júnior, Paulo Roberto de Moura, et al.
Published: (2026)
DiVA-DocRE: A Discriminative and Voice-Aware Paradigm for Document-Level Relation Extraction
by: Wu, Yiheng, et al.
Published: (2024)
by: Wu, Yiheng, et al.
Published: (2024)
SmartChunk Retrieval: Query-Aware Chunk Compression with Planning for Efficient Document RAG
by: Zhang, Xuechen, et al.
Published: (2025)
by: Zhang, Xuechen, et al.
Published: (2025)
ChunkNorris: A High-Performance and Low-Energy Approach to PDF Parsing and Chunking
by: Ciancone, Mathieu, et al.
Published: (2025)
by: Ciancone, Mathieu, et al.
Published: (2025)
Chunking, Retrieval, and Re-ranking: An Empirical Evaluation of RAG Architectures for Policy Document Question Answering
by: Maharjan, Anuj, et al.
Published: (2026)
by: Maharjan, Anuj, et al.
Published: (2026)
Reconstructing Context: Evaluating Advanced Chunking Strategies for Retrieval-Augmented Generation
by: Merola, Carlo, et al.
Published: (2025)
by: Merola, Carlo, et al.
Published: (2025)
Grounding Language Model with Chunking-Free In-Context Retrieval
by: Qian, Hongjin, et al.
Published: (2024)
by: Qian, Hongjin, et al.
Published: (2024)
Advancing Event Causality Identification via Heuristic Semantic Dependency Inquiry Network
by: Li, Haoran, et al.
Published: (2024)
by: Li, Haoran, et al.
Published: (2024)
The Chronicles of RAG: The Retriever, the Chunk and the Generator
by: Finardi, Paulo, et al.
Published: (2024)
by: Finardi, Paulo, et al.
Published: (2024)
cAST: Enhancing Code Retrieval-Augmented Generation with Structural Chunking via Abstract Syntax Tree
by: Zhang, Yilin, et al.
Published: (2025)
by: Zhang, Yilin, et al.
Published: (2025)
MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question Answering
by: Wu, Hui, et al.
Published: (2026)
by: Wu, Hui, et al.
Published: (2026)
JuDGE: Benchmarking Judgment Document Generation for Chinese Legal System
by: Su, Weihang, et al.
Published: (2025)
by: Su, Weihang, et al.
Published: (2025)
S-Path-RAG: Semantic-Aware Shortest-Path Retrieval Augmented Generation for Multi-Hop Knowledge Graph Question Answering
by: Fu, Rong, et al.
Published: (2026)
by: Fu, Rong, et al.
Published: (2026)
Enhancing Judgment Document Generation via Agentic Legal Information Collection and Rubric-Guided Optimization
by: Su, Weihang, et al.
Published: (2026)
by: Su, Weihang, et al.
Published: (2026)
Phonetically-Augmented Discriminative Rescoring for Voice Search Error Correction
by: Van Gysel, Christophe, et al.
Published: (2025)
by: Van Gysel, Christophe, et al.
Published: (2025)
SEFD: Semantic-Enhanced Framework for Detecting LLM-Generated Text
by: He, Weiqing, et al.
Published: (2024)
by: He, Weiqing, et al.
Published: (2024)
Neural Retrievers are Biased Towards LLM-Generated Content
by: Dai, Sunhao, et al.
Published: (2023)
by: Dai, Sunhao, et al.
Published: (2023)
Stealthy Attack on Large Language Model based Recommendation
by: Zhang, Jinghao, et al.
Published: (2024)
by: Zhang, Jinghao, et al.
Published: (2024)
GOLFer: Smaller LM-Generated Documents Hallucination Filter & Combiner for Query Expansion in Information Retrieval
by: Liu, Lingyuan, et al.
Published: (2025)
by: Liu, Lingyuan, et al.
Published: (2025)
HierSearch: A Hierarchical Enterprise Deep Search Framework Integrating Local and Web Searches
by: Tan, Jiejun, et al.
Published: (2025)
by: Tan, Jiejun, et al.
Published: (2025)
Clinical Document Metadata Extraction: A Scoping Review
by: Miller, Kurt, et al.
Published: (2025)
by: Miller, Kurt, et al.
Published: (2025)
Structured Attention Matters to Multimodal LLMs in Document Understanding
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
Docs2KG: Unified Knowledge Graph Construction from Heterogeneous Documents Assisted by Large Language Models
by: Sun, Qiang, et al.
Published: (2024)
by: Sun, Qiang, et al.
Published: (2024)
Tuning LLMs by RAG Principles: Towards LLM-native Memory
by: Wei, Jiale, et al.
Published: (2025)
by: Wei, Jiale, et al.
Published: (2025)
Event GDR: Event-Centric Generative Document Retrieval
by: Guan, Yong, et al.
Published: (2024)
by: Guan, Yong, et al.
Published: (2024)
Position: Text Embeddings Should Capture Implicit Semantics, Not Just Surface Meaning
by: Sun, Yiqun, et al.
Published: (2025)
by: Sun, Yiqun, et al.
Published: (2025)
Doc2SAR: A Synergistic Framework for High-Fidelity Extraction of Structure-Activity Relationships from Scientific Documents
by: Zhuang, Jiaxi, et al.
Published: (2025)
by: Zhuang, Jiaxi, et al.
Published: (2025)
Perplexity Trap: PLM-Based Retrievers Overrate Low Perplexity Documents
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
HybridRAG: A Practical LLM-based ChatBot Framework based on Pre-Generated Q&A over Raw Unstructured Documents
by: Kim, Sungmoon, et al.
Published: (2025)
by: Kim, Sungmoon, et al.
Published: (2025)
Agentic Chunking and Bayesian De-chunking of AI Generated Fuzzy Cognitive Maps: A Model of the Thucydides Trap
by: Panda, Akash Kumar, et al.
Published: (2026)
by: Panda, Akash Kumar, et al.
Published: (2026)
SurGE: A Benchmark and Evaluation Framework for Scientific Survey Generation
by: Su, Weihang, et al.
Published: (2025)
by: Su, Weihang, et al.
Published: (2025)
Reasoning-Enhanced Self-Training for Long-Form Personalized Text Generation
by: Salemi, Alireza, et al.
Published: (2025)
by: Salemi, Alireza, et al.
Published: (2025)
Semantic Convergence: Harmonizing Recommender Systems via Two-Stage Alignment and Behavioral Semantic Tokenization
by: Li, Guanghan, et al.
Published: (2024)
by: Li, Guanghan, et al.
Published: (2024)
Benchmarking Complex Multimodal Document Processing Pipelines: A Unified Evaluation Framework for Enterprise AI
by: Singh, Saurabh K., et al.
Published: (2026)
by: Singh, Saurabh K., et al.
Published: (2026)
Think Before Recommend: Unleashing the Latent Reasoning Power for Sequential Recommendation
by: Tang, Jiakai, et al.
Published: (2025)
by: Tang, Jiakai, et al.
Published: (2025)
SCARV: Structure-Constrained Aggregation for Stable Sample Ranking in Redundant NLP Datasets
by: Zheng, Xu, et al.
Published: (2026)
by: Zheng, Xu, et al.
Published: (2026)
GenCRF: Generative Clustering and Reformulation Framework for Enhanced Intent-Driven Information Retrieval
by: Seo, Wonduk, et al.
Published: (2024)
by: Seo, Wonduk, et al.
Published: (2024)
LongKey: Keyphrase Extraction for Long Documents
by: Alves, Jeovane Honorio, et al.
Published: (2024)
by: Alves, Jeovane Honorio, et al.
Published: (2024)
Similar Items
-
Beyond Chunking: Discourse-Aware Hierarchical Retrieval for Long Document Question Answering
by: Chen, Huiyao, et al.
Published: (2025) -
Cross-Document Topic-Aligned Chunking for Retrieval-Augmented Generation
by: Stankovic, Mile
Published: (2025) -
Adaptive Chunking: Optimizing Chunking-Method Selection for RAG
by: Júnior, Paulo Roberto de Moura, et al.
Published: (2026) -
DiVA-DocRE: A Discriminative and Voice-Aware Paradigm for Document-Level Relation Extraction
by: Wu, Yiheng, et al.
Published: (2024) -
SmartChunk Retrieval: Query-Aware Chunk Compression with Planning for Efficient Document RAG
by: Zhang, Xuechen, et al.
Published: (2025)