Pre-training vs. Fine-tuning: A Reproducibility Study on Dense Retrieval Knowledge Acquisition
Fuente:
arXiv
Saved in:
| Main Authors: | Yao, Zheng, Wang, Shuai, Zuccon, Guido |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding and Mitigating the Threat of Vec2Text to Dense Retrieval Systems
by: Zhuang, Shengyao, et al.
Published: (2024)
by: Zhuang, Shengyao, et al.
Published: (2024)
2D Matryoshka Training for Information Retrieval
by: Wang, Shuai, et al.
Published: (2024)
by: Wang, Shuai, et al.
Published: (2024)
FeB4RAG: Evaluating Federated Search in the Context of Retrieval Augmented Generation
by: Wang, Shuai, et al.
Published: (2024)
by: Wang, Shuai, et al.
Published: (2024)
Large Language Models for Stemming: Promises, Pitfalls and Failures
by: Wang, Shuai, et al.
Published: (2024)
by: Wang, Shuai, et al.
Published: (2024)
Reassessing Large Language Model Boolean Query Generation for Systematic Reviews
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
An Investigation of Prompt Variations for Zero-shot LLM-based Rankers
by: Sun, Shuoqi, et al.
Published: (2024)
by: Sun, Shuoqi, et al.
Published: (2024)
Drop your Decoder: Pre-training with Bag-of-Word Prediction for Dense Passage Retrieval
by: Ma, Guangyuan, et al.
Published: (2024)
by: Ma, Guangyuan, et al.
Published: (2024)
Reproducing HotFlip for Corpus Poisoning Attacks in Dense Retrieval
by: Li, Yongkang, et al.
Published: (2025)
by: Li, Yongkang, et al.
Published: (2025)
Zero-shot Generative Large Language Models for Systematic Review Screening Automation
by: Wang, Shuai, et al.
Published: (2024)
by: Wang, Shuai, et al.
Published: (2024)
Evaluating Generative Ad Hoc Information Retrieval
by: Gienapp, Lukas, et al.
Published: (2023)
by: Gienapp, Lukas, et al.
Published: (2023)
DenseReviewer: A Screening Prioritisation Tool for Systematic Review based on Dense Retrieval
by: Mao, Xinyu, et al.
Published: (2025)
by: Mao, Xinyu, et al.
Published: (2025)
Inferential Question Answering
by: Mozafari, Jamshid, et al.
Published: (2026)
by: Mozafari, Jamshid, et al.
Published: (2026)
A Reproducibility Study of Goldilocks: Just-Right Tuning of BERT for TAR
by: Mao, Xinyu, et al.
Published: (2024)
by: Mao, Xinyu, et al.
Published: (2024)
Leveraging LLMs for Unsupervised Dense Retriever Ranking
by: Khramtsova, Ekaterina, et al.
Published: (2024)
by: Khramtsova, Ekaterina, et al.
Published: (2024)
RbFT: Robust Fine-tuning for Retrieval-Augmented Generation against Retrieval Defects
by: Tu, Yiteng, et al.
Published: (2025)
by: Tu, Yiteng, et al.
Published: (2025)
Dense Passage Retrieval: Is it Retrieving?
by: Reichman, Benjamin, et al.
Published: (2024)
by: Reichman, Benjamin, et al.
Published: (2024)
Rank-R1: Enhancing Reasoning in LLM-based Document Rerankers via Reinforcement Learning
by: Zhuang, Shengyao, et al.
Published: (2025)
by: Zhuang, Shengyao, et al.
Published: (2025)
Embark on DenseQuest: A System for Selecting the Best Dense Retriever for a Custom Collection
by: Khramtsova, Ekaterina, et al.
Published: (2024)
by: Khramtsova, Ekaterina, et al.
Published: (2024)
Dense Retrieval with Continuous Explicit Feedback for Systematic Review Screening Prioritisation
by: Mao, Xinyu, et al.
Published: (2024)
by: Mao, Xinyu, et al.
Published: (2024)
Study on LLMs for Promptagator-Style Dense Retriever Training
by: Gwon, Daniel, et al.
Published: (2025)
by: Gwon, Daniel, et al.
Published: (2025)
Knowledge Compression via Question Generation: Enhancing Multihop Document Retrieval without Fine-tuning
by: Eponon, Anvi Alex, et al.
Published: (2025)
by: Eponon, Anvi Alex, et al.
Published: (2025)
Cohort Retrieval using Dense Passage Retrieval
by: Jadhav, Pranav
Published: (2025)
by: Jadhav, Pranav
Published: (2025)
Evaluating LLM-based Approaches to Legal Citation Prediction: Domain-specific Pre-training, Fine-tuning, or RAG? A Benchmark and an Australian Law Case Study
by: Han, Jiuzhou, et al.
Published: (2024)
by: Han, Jiuzhou, et al.
Published: (2024)
Beyond Chunk-Then-Embed: A Comprehensive Taxonomy and Evaluation of Document Chunking Strategies for Information Retrieval
by: Zhou, Yongjie, et al.
Published: (2026)
by: Zhou, Yongjie, et al.
Published: (2026)
Reproducing Complex Set-Compositional Information Retrieval
by: Degenhart, Vincent, et al.
Published: (2026)
by: Degenhart, Vincent, et al.
Published: (2026)
Scaling Laws For Dense Retrieval
by: Fang, Yan, et al.
Published: (2024)
by: Fang, Yan, et al.
Published: (2024)
ChatRetriever: Adapting Large Language Models for Generalized and Robust Conversational Dense Retrieval
by: Mao, Kelong, et al.
Published: (2024)
by: Mao, Kelong, et al.
Published: (2024)
Unlearning for Federated Online Learning to Rank: A Reproducibility Study
by: Tao, Yiling, et al.
Published: (2025)
by: Tao, Yiling, et al.
Published: (2025)
AutoBool: An Reinforcement-Learning trained LLM for Effective Automated Boolean Query Generation for Systematic Reviews
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
A Reproducibility Study of PLAID
by: MacAvaney, Sean, et al.
Published: (2024)
by: MacAvaney, Sean, et al.
Published: (2024)
Towards Context-Robust LLMs: A Gated Representation Fine-tuning Approach
by: Zeng, Shenglai, et al.
Published: (2025)
by: Zeng, Shenglai, et al.
Published: (2025)
Generative Dense Retrieval: Memory Can Be a Burden
by: Yuan, Peiwen, et al.
Published: (2024)
by: Yuan, Peiwen, et al.
Published: (2024)
History-Aware Conversational Dense Retrieval
by: Mo, Fengran, et al.
Published: (2024)
by: Mo, Fengran, et al.
Published: (2024)
ColBERT-Zero: To Pre-train Or Not To Pre-train ColBERT models
by: Chaffin, Antoine, et al.
Published: (2026)
by: Chaffin, Antoine, et al.
Published: (2026)
QuCo-RAG: Quantifying Uncertainty from the Pre-training Corpus for Dynamic Retrieval-Augmented Generation
by: Min, Dehai, et al.
Published: (2025)
by: Min, Dehai, et al.
Published: (2025)
DELTA: Pre-train a Discriminative Encoder for Legal Case Retrieval via Structural Word Alignment
by: Li, Haitao, et al.
Published: (2024)
by: Li, Haitao, et al.
Published: (2024)
Improve Dense Passage Retrieval with Entailment Tuning
by: Dai, Lu, et al.
Published: (2024)
by: Dai, Lu, et al.
Published: (2024)
Pseudo Relevance Feedback is Enough to Close the Gap Between Small and Large Dense Retrieval Models
by: Li, Hang, et al.
Published: (2025)
by: Li, Hang, et al.
Published: (2025)
A Representation Sharpening Framework for Zero Shot Dense Retrieval
by: Ashok, Dhananjay, et al.
Published: (2025)
by: Ashok, Dhananjay, et al.
Published: (2025)
Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval
by: Sinha, Aarush
Published: (2025)
by: Sinha, Aarush
Published: (2025)
Similar Items
-
Understanding and Mitigating the Threat of Vec2Text to Dense Retrieval Systems
by: Zhuang, Shengyao, et al.
Published: (2024) -
2D Matryoshka Training for Information Retrieval
by: Wang, Shuai, et al.
Published: (2024) -
FeB4RAG: Evaluating Federated Search in the Context of Retrieval Augmented Generation
by: Wang, Shuai, et al.
Published: (2024) -
Large Language Models for Stemming: Promises, Pitfalls and Failures
by: Wang, Shuai, et al.
Published: (2024) -
Reassessing Large Language Model Boolean Query Generation for Systematic Reviews
by: Wang, Shuai, et al.
Published: (2025)