Advancing continual lifelong learning in neural information retrieval: definition, dataset, framework, and empirical evaluation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hou, Jingrui, Cosma, Georgina, Finke, Axel
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916294617464832
author Hou, Jingrui
Cosma, Georgina
Finke, Axel
author_facet Hou, Jingrui
Cosma, Georgina
Finke, Axel
contents Continual learning refers to the capability of a machine learning model to learn and adapt to new information, without compromising its performance on previously learned tasks. Although several studies have investigated continual learning methods for information retrieval tasks, a well-defined task formulation is still lacking, and it is unclear how typical learning strategies perform in this context. To address this challenge, a systematic task formulation of continual neural information retrieval is presented, along with a multiple-topic dataset that simulates continuous information retrieval. A comprehensive continual neural information retrieval framework consisting of typical retrieval models and continual learning strategies is then proposed. Empirical evaluations illustrate that the proposed framework can successfully prevent catastrophic forgetting in neural information retrieval and enhance performance on previously learned tasks. The results indicate that embedding-based retrieval models experience a decline in their continual learning performance as the topic shift distance and dataset volume of new tasks increase. In contrast, pretraining-based models do not show any such correlation. Adopting suitable learning strategies can mitigate the effects of topic shift and data augmentation.
format Preprint
id arxiv_https___arxiv_org_abs_2308_08378
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Advancing continual lifelong learning in neural information retrieval: definition, dataset, framework, and empirical evaluation
Hou, Jingrui
Cosma, Georgina
Finke, Axel
Information Retrieval
Computation and Language
Continual learning refers to the capability of a machine learning model to learn and adapt to new information, without compromising its performance on previously learned tasks. Although several studies have investigated continual learning methods for information retrieval tasks, a well-defined task formulation is still lacking, and it is unclear how typical learning strategies perform in this context. To address this challenge, a systematic task formulation of continual neural information retrieval is presented, along with a multiple-topic dataset that simulates continuous information retrieval. A comprehensive continual neural information retrieval framework consisting of typical retrieval models and continual learning strategies is then proposed. Empirical evaluations illustrate that the proposed framework can successfully prevent catastrophic forgetting in neural information retrieval and enhance performance on previously learned tasks. The results indicate that embedding-based retrieval models experience a decline in their continual learning performance as the topic shift distance and dataset volume of new tasks increase. In contrast, pretraining-based models do not show any such correlation. Adopting suitable learning strategies can mitigate the effects of topic shift and data augmentation.
title Advancing continual lifelong learning in neural information retrieval: definition, dataset, framework, and empirical evaluation
topic Information Retrieval
Computation and Language
url https://arxiv.org/abs/2308.08378