Learning Effective Representations for Retrieval Using Self-Distillation with Adaptive Relevance Margins
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912443053113344 |
|---|---|
| author | Gienapp, Lukas Deckers, Niklas Potthast, Martin Scells, Harrisen |
| author_facet | Gienapp, Lukas Deckers, Niklas Potthast, Martin Scells, Harrisen |
| contents | Representation-based retrieval models, so-called bi-encoders, estimate the relevance of a document to a query by calculating the similarity of their respective embeddings. Current state-of-the-art bi-encoders are trained using an expensive training regime involving knowledge distillation from a teacher model and batch-sampling. Instead of relying on a teacher model, we contribute a novel parameter-free loss function for self-supervision that exploits the pre-trained language modeling capabilities of the encoder model as a training signal, eliminating the need for batch sampling by performing implicit hard negative mining. We investigate the capabilities of our proposed approach through extensive experiments, demonstrating that self-distillation can match the effectiveness of teacher distillation using only 13.5% of the data, while offering a speedup in training time between 3x and 15x compared to parametrized losses. All code and data is made openly available. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_21515 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Learning Effective Representations for Retrieval Using Self-Distillation with Adaptive Relevance Margins Gienapp, Lukas Deckers, Niklas Potthast, Martin Scells, Harrisen Information Retrieval Representation-based retrieval models, so-called bi-encoders, estimate the relevance of a document to a query by calculating the similarity of their respective embeddings. Current state-of-the-art bi-encoders are trained using an expensive training regime involving knowledge distillation from a teacher model and batch-sampling. Instead of relying on a teacher model, we contribute a novel parameter-free loss function for self-supervision that exploits the pre-trained language modeling capabilities of the encoder model as a training signal, eliminating the need for batch sampling by performing implicit hard negative mining. We investigate the capabilities of our proposed approach through extensive experiments, demonstrating that self-distillation can match the effectiveness of teacher distillation using only 13.5% of the data, while offering a speedup in training time between 3x and 15x compared to parametrized losses. All code and data is made openly available. |
| title | Learning Effective Representations for Retrieval Using Self-Distillation with Adaptive Relevance Margins |
| topic | Information Retrieval |
| url | https://arxiv.org/abs/2407.21515 |