Simple Domain Adaptation for Sparse Retrievers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vast, Mathias, Zong, Yuxuan, Van Cooten, Basile, Piwowarski, Benjamin, Soulier, Laure
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913416875081728
author Vast, Mathias
Zong, Yuxuan
Van Cooten, Basile
Piwowarski, Benjamin
Soulier, Laure
author_facet Vast, Mathias
Zong, Yuxuan
Van Cooten, Basile
Piwowarski, Benjamin
Soulier, Laure
contents In Information Retrieval, and more generally in Natural Language Processing, adapting models to specific domains is conducted through fine-tuning. Despite the successes achieved by this method and its versatility, the need for human-curated and labeled data makes it impractical to transfer to new tasks, domains, and/or languages when training data doesn't exist. Using the model without training (zero-shot) is another option that however suffers an effectiveness cost, especially in the case of first-stage retrievers. Numerous research directions have emerged to tackle these issues, most of them in the context of adapting to a task or a language. However, the literature is scarcer for domain (or topic) adaptation. In this paper, we address this issue of cross-topic discrepancy for a sparse first-stage retriever by transposing a method initially designed for language adaptation. By leveraging pre-training on the target data to learn domain-specific knowledge, this technique alleviates the need for annotated data and expands the scope of domain adaptation. Despite their relatively good generalization ability, we show that even sparse retrievers can benefit from our simple domain adaptation method.
format Preprint
id arxiv_https___arxiv_org_abs_2401_11509
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Simple Domain Adaptation for Sparse Retrievers
Vast, Mathias
Zong, Yuxuan
Van Cooten, Basile
Piwowarski, Benjamin
Soulier, Laure
Information Retrieval
In Information Retrieval, and more generally in Natural Language Processing, adapting models to specific domains is conducted through fine-tuning. Despite the successes achieved by this method and its versatility, the need for human-curated and labeled data makes it impractical to transfer to new tasks, domains, and/or languages when training data doesn't exist. Using the model without training (zero-shot) is another option that however suffers an effectiveness cost, especially in the case of first-stage retrievers. Numerous research directions have emerged to tackle these issues, most of them in the context of adapting to a task or a language. However, the literature is scarcer for domain (or topic) adaptation. In this paper, we address this issue of cross-topic discrepancy for a sparse first-stage retriever by transposing a method initially designed for language adaptation. By leveraging pre-training on the target data to learn domain-specific knowledge, this technique alleviates the need for annotated data and expands the scope of domain adaptation. Despite their relatively good generalization ability, we show that even sparse retrievers can benefit from our simple domain adaptation method.
title Simple Domain Adaptation for Sparse Retrievers
topic Information Retrieval
url https://arxiv.org/abs/2401.11509