A Survey of Long-Document Retrieval in the PLM and LLM Era
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Minghan, Luo, Miyang, Lv, Tianrui, Zhang, Yishuai, Zhao, Siqi, Nie, Ercong, Zhou, Guodong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Automatic In-Domain Exemplar Construction and LLM-Based Refinement of Multi-LLM Expansions for Query Expansion
by: Li, Minghan, et al.
Published: (2026)
by: Li, Minghan, et al.
Published: (2026)
GenState-AI: State-Aware Dataset for Text-to-Video Retrieval on AI-Generated Videos
by: Li, Minghan, et al.
Published: (2026)
by: Li, Minghan, et al.
Published: (2026)
Query Expansion in the Age of Pre-trained and Large Language Models: A Comprehensive Survey
by: Li, Minghan, et al.
Published: (2025)
by: Li, Minghan, et al.
Published: (2025)
GLIER: Generative Legal Inference and Evidence Ranking for Legal Case Retrieval
by: Li, Minghan, et al.
Published: (2026)
by: Li, Minghan, et al.
Published: (2026)
EviRerank: Adaptive Evidence Construction for Long-Document LLM Reranking
by: Li, Minghan, et al.
Published: (2024)
by: Li, Minghan, et al.
Published: (2024)
Retrieval-Feedback-Driven Distillation and Preference Alignment for Efficient LLM-based Query Expansion
by: Li, Minghan, et al.
Published: (2026)
by: Li, Minghan, et al.
Published: (2026)
Efficient Long-Document Reranking via Block-Level Embeddings and Top-k Interaction Refinement
by: Li, Minghan, et al.
Published: (2025)
by: Li, Minghan, et al.
Published: (2025)
S2G-RAG: Structured Sufficiency and Gap Judging for Iterative Retrieval-Augmented QA
by: Li, Minghan, et al.
Published: (2026)
by: Li, Minghan, et al.
Published: (2026)
Perplexity Trap: PLM-Based Retrievers Overrate Low Perplexity Documents
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
Domain Adaptation for Dense Retrieval and Conversational Dense Retrieval through Self-Supervision by Meticulous Pseudo-Relevance Labeling
by: Li, Minghan, et al.
Published: (2024)
by: Li, Minghan, et al.
Published: (2024)
Unifying Multimodal Retrieval via Document Screenshot Embedding
by: Ma, Xueguang, et al.
Published: (2024)
by: Ma, Xueguang, et al.
Published: (2024)
Semantics-Aware Denoising: A PLM-Guided Sample Reweighting Strategy for Robust Recommendation
by: Yang, Xikai, et al.
Published: (2026)
by: Yang, Xikai, et al.
Published: (2026)
A Survey on Open Dataset Search in the LLM Era: Retrospectives and Perspectives
by: Li, Pengyue, et al.
Published: (2025)
by: Li, Pengyue, et al.
Published: (2025)
EraRAG: Efficient and Incremental Retrieval Augmented Generation for Growing Corpora
by: Zhang, Fangyuan, et al.
Published: (2025)
by: Zhang, Fangyuan, et al.
Published: (2025)
Length-Induced Embedding Collapse in PLM-based Models
by: Zhou, Yuqi, et al.
Published: (2024)
by: Zhou, Yuqi, et al.
Published: (2024)
AnnoRetrieve: Efficient Structured Retrieval for Unstructured Document Analysis
by: Lin, Teng, et al.
Published: (2026)
by: Lin, Teng, et al.
Published: (2026)
On the Reproducibility of Learned Sparse Retrieval Adaptations for Long Documents
by: Lionis, Emmanouil Georgios, et al.
Published: (2025)
by: Lionis, Emmanouil Georgios, et al.
Published: (2025)
AdversarialCoT: Single-Document Retrieval Poisoning for LLM Reasoning
by: Song, Hongru, et al.
Published: (2026)
by: Song, Hongru, et al.
Published: (2026)
A Survey of Generative Search and Recommendation in the Era of Large Language Models
by: Li, Yongqi, et al.
Published: (2024)
by: Li, Yongqi, et al.
Published: (2024)
Conversational Search: From Fundamentals to Frontiers in the LLM Era
by: Mo, Fengran, et al.
Published: (2025)
by: Mo, Fengran, et al.
Published: (2025)
Personalized Generation In Large Model Era: A Survey
by: Xu, Yiyan, et al.
Published: (2025)
by: Xu, Yiyan, et al.
Published: (2025)
AttentionRetriever: Attention Layers are Secretly Long Document Retrievers
by: Fu, David Jiahao, et al.
Published: (2026)
by: Fu, David Jiahao, et al.
Published: (2026)
DiffuGR: Generative Document Retrieval with Diffusion Language Models
by: Zhao, Xinpeng, et al.
Published: (2025)
by: Zhao, Xinpeng, et al.
Published: (2025)
Roles of MLLMs in Visually Rich Document Retrieval for RAG: A Survey
by: Zhang, Xiantao
Published: (2025)
by: Zhang, Xiantao
Published: (2025)
Rethinking Chunk Size For Long-Document Retrieval: A Multi-Dataset Analysis
by: Bhat, Sinchana Ramakanth, et al.
Published: (2025)
by: Bhat, Sinchana Ramakanth, et al.
Published: (2025)
GraphRAG-IRL: Personalized Recommendation with Graph-Grounded Inverse Reinforcement Learning and LLM Re-ranking
by: Liang, Siqi, et al.
Published: (2026)
by: Liang, Siqi, et al.
Published: (2026)
DocMMIR: A Framework for Document Multi-modal Information Retrieval
by: Li, Zirui, et al.
Published: (2025)
by: Li, Zirui, et al.
Published: (2025)
Personalize Before Retrieve: LLM-based Personalized Query Expansion for User-Centric Retrieval
by: Zhang, Yingyi, et al.
Published: (2025)
by: Zhang, Yingyi, et al.
Published: (2025)
Retrieval-in-the-Chain: Bootstrapping Large Language Models for Generative Retrieval
by: Zhang, Yingchen, et al.
Published: (2025)
by: Zhang, Yingchen, et al.
Published: (2025)
Multivector Reranking in the Era of Strong First-Stage Retrievers
by: Martinico, Silvio, et al.
Published: (2026)
by: Martinico, Silvio, et al.
Published: (2026)
UniFAR: A Unified Facet-Aware Retrieval Framework for Scientific Documents
by: Dou, Zheng, et al.
Published: (2026)
by: Dou, Zheng, et al.
Published: (2026)
Towards Next-Generation LLM-based Recommender Systems: A Survey and Beyond
by: Wang, Qi, et al.
Published: (2024)
by: Wang, Qi, et al.
Published: (2024)
Multi-Objective Recommendation in the Era of Generative AI: A Survey of Recent Progress and Future Prospects
by: Hong, Zihan, et al.
Published: (2025)
by: Hong, Zihan, et al.
Published: (2025)
A Survey of Model Architectures in Information Retrieval
by: Xu, Zhichao, et al.
Published: (2025)
by: Xu, Zhichao, et al.
Published: (2025)
PairSem: LLM-Guided Pairwise Semantic Matching for Scientific Document Retrieval
by: Kweon, Wonbin, et al.
Published: (2025)
by: Kweon, Wonbin, et al.
Published: (2025)
SyNeg: LLM-Driven Synthetic Hard-Negatives for Dense Retrieval
by: Li, Xiaopeng, et al.
Published: (2024)
by: Li, Xiaopeng, et al.
Published: (2024)
Synthetic Data Powers Product Retrieval for Long-tail Knowledge-Intensive Queries in E-commerce Search
by: Ling, Gui, et al.
Published: (2026)
by: Ling, Gui, et al.
Published: (2026)
A Survey on Deep Text Hashing: Efficient Semantic Text Retrieval with Binary Representation
by: He, Liyang, et al.
Published: (2025)
by: He, Liyang, et al.
Published: (2025)
SEAL: Structure and Element Aware Learning to Improve Long Structured Document Retrieval
by: Huang, Xinhao, et al.
Published: (2025)
by: Huang, Xinhao, et al.
Published: (2025)
Cocktail: A Comprehensive Information Retrieval Benchmark with LLM-Generated Documents Integration
by: Dai, Sunhao, et al.
Published: (2024)
by: Dai, Sunhao, et al.
Published: (2024)
Similar Items
-
Automatic In-Domain Exemplar Construction and LLM-Based Refinement of Multi-LLM Expansions for Query Expansion
by: Li, Minghan, et al.
Published: (2026) -
GenState-AI: State-Aware Dataset for Text-to-Video Retrieval on AI-Generated Videos
by: Li, Minghan, et al.
Published: (2026) -
Query Expansion in the Age of Pre-trained and Large Language Models: A Comprehensive Survey
by: Li, Minghan, et al.
Published: (2025) -
GLIER: Generative Legal Inference and Evidence Ranking for Legal Case Retrieval
by: Li, Minghan, et al.
Published: (2026) -
EviRerank: Adaptive Evidence Construction for Long-Document LLM Reranking
by: Li, Minghan, et al.
Published: (2024)