Unveiling the Magic: Investigating Attention Distillation in Retrieval-augmented Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zizhong, Zhang, Haopeng, Zhang, Jiawei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911779408314368
author Li, Zizhong
Zhang, Haopeng
Zhang, Jiawei
author_facet Li, Zizhong
Zhang, Haopeng
Zhang, Jiawei
contents Retrieval-augmented generation framework can address the limitations of large language models by enabling real-time knowledge updates for more accurate answers. An efficient way in the training phase of retrieval-augmented models is attention distillation, which uses attention scores as a supervision signal instead of manually annotated query-document pairs. Despite its growing popularity, the detailed mechanisms behind the success of attention distillation remain unexplored, particularly the specific patterns it leverages to benefit training. In this paper, we address this gap by conducting a comprehensive review of attention distillation workflow and identifying key factors influencing the learning quality of retrieval-augmented language models. We further propose indicators for optimizing models' training methods and avoiding ineffective training.
format Preprint
id arxiv_https___arxiv_org_abs_2402_11794
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unveiling the Magic: Investigating Attention Distillation in Retrieval-augmented Generation
Li, Zizhong
Zhang, Haopeng
Zhang, Jiawei
Computation and Language
Information Retrieval
Retrieval-augmented generation framework can address the limitations of large language models by enabling real-time knowledge updates for more accurate answers. An efficient way in the training phase of retrieval-augmented models is attention distillation, which uses attention scores as a supervision signal instead of manually annotated query-document pairs. Despite its growing popularity, the detailed mechanisms behind the success of attention distillation remain unexplored, particularly the specific patterns it leverages to benefit training. In this paper, we address this gap by conducting a comprehensive review of attention distillation workflow and identifying key factors influencing the learning quality of retrieval-augmented language models. We further propose indicators for optimizing models' training methods and avoiding ineffective training.
title Unveiling the Magic: Investigating Attention Distillation in Retrieval-augmented Generation
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2402.11794