Intermediate Distillation: Data-Efficient Distillation from Black-Box LLMs for Information Retrieval

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Li, Zizhong, Zhang, Haopeng, Zhang, Jiawei
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910491981381632
author Li, Zizhong
Zhang, Haopeng
Zhang, Jiawei
author_facet Li, Zizhong
Zhang, Haopeng
Zhang, Jiawei
contents Recent research has explored distilling knowledge from large language models (LLMs) to optimize retriever models, especially within the retrieval-augmented generation (RAG) framework. However, most existing training methods rely on extracting supervision signals from LLMs' weights or their output probabilities, which is not only resource-intensive but also incompatible with black-box LLMs. In this paper, we introduce \textit{Intermediate Distillation}, a data-efficient knowledge distillation training scheme that treats LLMs as black boxes and distills their knowledge via an innovative LLM-ranker-retriever pipeline, solely using LLMs' ranking generation as the supervision signal. Extensive experiments demonstrate that our proposed method can significantly improve the performance of retriever models with only 1,000 training instances. Moreover, our distilled retriever model significantly boosts performance in question-answering tasks within the RAG framework, demonstrating the potential of LLMs to economically and effectively train smaller models.
format Preprint
id arxiv_https___arxiv_org_abs_2406_12169
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Intermediate Distillation: Data-Efficient Distillation from Black-Box LLMs for Information Retrieval
Li, Zizhong
Zhang, Haopeng
Zhang, Jiawei
Information Retrieval
Recent research has explored distilling knowledge from large language models (LLMs) to optimize retriever models, especially within the retrieval-augmented generation (RAG) framework. However, most existing training methods rely on extracting supervision signals from LLMs' weights or their output probabilities, which is not only resource-intensive but also incompatible with black-box LLMs. In this paper, we introduce \textit{Intermediate Distillation}, a data-efficient knowledge distillation training scheme that treats LLMs as black boxes and distills their knowledge via an innovative LLM-ranker-retriever pipeline, solely using LLMs' ranking generation as the supervision signal. Extensive experiments demonstrate that our proposed method can significantly improve the performance of retriever models with only 1,000 training instances. Moreover, our distilled retriever model significantly boosts performance in question-answering tasks within the RAG framework, demonstrating the potential of LLMs to economically and effectively train smaller models.
title Intermediate Distillation: Data-Efficient Distillation from Black-Box LLMs for Information Retrieval
topic Information Retrieval
url https://arxiv.org/abs/2406.12169