Learning Effective Representations for Retrieval Using Self-Distillation with Adaptive Relevance Margins

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gienapp, Lukas, Deckers, Niklas, Potthast, Martin, Scells, Harrisen
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912443053113344
author Gienapp, Lukas
Deckers, Niklas
Potthast, Martin
Scells, Harrisen
author_facet Gienapp, Lukas
Deckers, Niklas
Potthast, Martin
Scells, Harrisen
contents Representation-based retrieval models, so-called bi-encoders, estimate the relevance of a document to a query by calculating the similarity of their respective embeddings. Current state-of-the-art bi-encoders are trained using an expensive training regime involving knowledge distillation from a teacher model and batch-sampling. Instead of relying on a teacher model, we contribute a novel parameter-free loss function for self-supervision that exploits the pre-trained language modeling capabilities of the encoder model as a training signal, eliminating the need for batch sampling by performing implicit hard negative mining. We investigate the capabilities of our proposed approach through extensive experiments, demonstrating that self-distillation can match the effectiveness of teacher distillation using only 13.5% of the data, while offering a speedup in training time between 3x and 15x compared to parametrized losses. All code and data is made openly available.
format Preprint
id arxiv_https___arxiv_org_abs_2407_21515
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning Effective Representations for Retrieval Using Self-Distillation with Adaptive Relevance Margins
Gienapp, Lukas
Deckers, Niklas
Potthast, Martin
Scells, Harrisen
Information Retrieval
Representation-based retrieval models, so-called bi-encoders, estimate the relevance of a document to a query by calculating the similarity of their respective embeddings. Current state-of-the-art bi-encoders are trained using an expensive training regime involving knowledge distillation from a teacher model and batch-sampling. Instead of relying on a teacher model, we contribute a novel parameter-free loss function for self-supervision that exploits the pre-trained language modeling capabilities of the encoder model as a training signal, eliminating the need for batch sampling by performing implicit hard negative mining. We investigate the capabilities of our proposed approach through extensive experiments, demonstrating that self-distillation can match the effectiveness of teacher distillation using only 13.5% of the data, while offering a speedup in training time between 3x and 15x compared to parametrized losses. All code and data is made openly available.
title Learning Effective Representations for Retrieval Using Self-Distillation with Adaptive Relevance Margins
topic Information Retrieval
url https://arxiv.org/abs/2407.21515