Enregistré dans:
Détails bibliographiques
Auteurs principaux: Sarthi, Parth, Abdullah, Salman, Tuli, Aditi, Khanna, Shubh, Goldie, Anna, Manning, Christopher D.
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:https://arxiv.org/abs/2401.18059
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916111338962944
author Sarthi, Parth
Abdullah, Salman
Tuli, Aditi
Khanna, Shubh
Goldie, Anna
Manning, Christopher D.
author_facet Sarthi, Parth
Abdullah, Salman
Tuli, Aditi
Khanna, Shubh
Goldie, Anna
Manning, Christopher D.
contents Retrieval-augmented language models can better adapt to changes in world state and incorporate long-tail knowledge. However, most existing methods retrieve only short contiguous chunks from a retrieval corpus, limiting holistic understanding of the overall document context. We introduce the novel approach of recursively embedding, clustering, and summarizing chunks of text, constructing a tree with differing levels of summarization from the bottom up. At inference time, our RAPTOR model retrieves from this tree, integrating information across lengthy documents at different levels of abstraction. Controlled experiments show that retrieval with recursive summaries offers significant improvements over traditional retrieval-augmented LMs on several tasks. On question-answering tasks that involve complex, multi-step reasoning, we show state-of-the-art results; for example, by coupling RAPTOR retrieval with the use of GPT-4, we can improve the best performance on the QuALITY benchmark by 20% in absolute accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2401_18059
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval
Sarthi, Parth
Abdullah, Salman
Tuli, Aditi
Khanna, Shubh
Goldie, Anna
Manning, Christopher D.
Computation and Language
Machine Learning
Retrieval-augmented language models can better adapt to changes in world state and incorporate long-tail knowledge. However, most existing methods retrieve only short contiguous chunks from a retrieval corpus, limiting holistic understanding of the overall document context. We introduce the novel approach of recursively embedding, clustering, and summarizing chunks of text, constructing a tree with differing levels of summarization from the bottom up. At inference time, our RAPTOR model retrieves from this tree, integrating information across lengthy documents at different levels of abstraction. Controlled experiments show that retrieval with recursive summaries offers significant improvements over traditional retrieval-augmented LMs on several tasks. On question-answering tasks that involve complex, multi-step reasoning, we show state-of-the-art results; for example, by coupling RAPTOR retrieval with the use of GPT-4, we can improve the best performance on the QuALITY benchmark by 20% in absolute accuracy.
title RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2401.18059