Similarity is Not All You Need: Endowing Retrieval Augmented Generation with Multi Layered Thoughts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gan, Chunjing, Yang, Dan, Hu, Binbin, Zhang, Hanxiao, Li, Siyuan, Liu, Ziqi, Shen, Yue, Ju, Lin, Zhang, Zhiqiang, Gu, Jinjie, Liang, Lei, Zhou, Jun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917679275704320
author Gan, Chunjing
Yang, Dan
Hu, Binbin
Zhang, Hanxiao
Li, Siyuan
Liu, Ziqi
Shen, Yue
Ju, Lin
Zhang, Zhiqiang
Gu, Jinjie
Liang, Lei
Zhou, Jun
author_facet Gan, Chunjing
Yang, Dan
Hu, Binbin
Zhang, Hanxiao
Li, Siyuan
Liu, Ziqi
Shen, Yue
Ju, Lin
Zhang, Zhiqiang
Gu, Jinjie
Liang, Lei
Zhou, Jun
contents In recent years, large language models (LLMs) have made remarkable achievements in various domains. However, the untimeliness and cost of knowledge updates coupled with hallucination issues of LLMs have curtailed their applications in knowledge intensive tasks, where retrieval augmented generation (RAG) can be of help. Nevertheless, existing retrieval augmented models typically use similarity as a bridge between queries and documents and follow a retrieve then read procedure. In this work, we argue that similarity is not always the panacea and totally relying on similarity would sometimes degrade the performance of retrieval augmented generation. To this end, we propose MetRag, a Multi layEred Thoughts enhanced Retrieval Augmented Generation framework. To begin with, beyond existing similarity oriented thought, we embrace a small scale utility model that draws supervision from an LLM for utility oriented thought and further come up with a smarter model by comprehensively combining the similarity and utility oriented thoughts. Furthermore, given the fact that the retrieved document set tends to be huge and using them in isolation makes it difficult to capture the commonalities and characteristics among them, we propose to make an LLM as a task adaptive summarizer to endow retrieval augmented generation with compactness-oriented thought. Finally, with multi layered thoughts from the precedent stages, an LLM is called for knowledge augmented generation. Extensive experiments on knowledge-intensive tasks have demonstrated the superiority of MetRag.
format Preprint
id arxiv_https___arxiv_org_abs_2405_19893
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Similarity is Not All You Need: Endowing Retrieval Augmented Generation with Multi Layered Thoughts
Gan, Chunjing
Yang, Dan
Hu, Binbin
Zhang, Hanxiao
Li, Siyuan
Liu, Ziqi
Shen, Yue
Ju, Lin
Zhang, Zhiqiang
Gu, Jinjie
Liang, Lei
Zhou, Jun
Machine Learning
Artificial Intelligence
Computation and Language
In recent years, large language models (LLMs) have made remarkable achievements in various domains. However, the untimeliness and cost of knowledge updates coupled with hallucination issues of LLMs have curtailed their applications in knowledge intensive tasks, where retrieval augmented generation (RAG) can be of help. Nevertheless, existing retrieval augmented models typically use similarity as a bridge between queries and documents and follow a retrieve then read procedure. In this work, we argue that similarity is not always the panacea and totally relying on similarity would sometimes degrade the performance of retrieval augmented generation. To this end, we propose MetRag, a Multi layEred Thoughts enhanced Retrieval Augmented Generation framework. To begin with, beyond existing similarity oriented thought, we embrace a small scale utility model that draws supervision from an LLM for utility oriented thought and further come up with a smarter model by comprehensively combining the similarity and utility oriented thoughts. Furthermore, given the fact that the retrieved document set tends to be huge and using them in isolation makes it difficult to capture the commonalities and characteristics among them, we propose to make an LLM as a task adaptive summarizer to endow retrieval augmented generation with compactness-oriented thought. Finally, with multi layered thoughts from the precedent stages, an LLM is called for knowledge augmented generation. Extensive experiments on knowledge-intensive tasks have demonstrated the superiority of MetRag.
title Similarity is Not All You Need: Endowing Retrieval Augmented Generation with Multi Layered Thoughts
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2405.19893