Superposition Prompting: Improving and Accelerating Retrieval-Augmented Generation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Merth, Thomas, Fu, Qichen, Rastegari, Mohammad, Najibi, Mahyar
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913437688266752
author Merth, Thomas
Fu, Qichen
Rastegari, Mohammad
Najibi, Mahyar
author_facet Merth, Thomas
Fu, Qichen
Rastegari, Mohammad
Najibi, Mahyar
contents Despite the successes of large language models (LLMs), they exhibit significant drawbacks, particularly when processing long contexts. Their inference cost scales quadratically with respect to sequence length, making it expensive for deployment in some real-world text processing applications, such as retrieval-augmented generation (RAG). Additionally, LLMs also exhibit the "distraction phenomenon", where irrelevant context in the prompt degrades output quality. To address these drawbacks, we propose a novel RAG prompting methodology, *superposition prompting*, which can be directly applied to pre-trained transformer-based LLMs *without the need for fine-tuning*. At a high level, superposition prompting allows the LLM to process input documents in parallel *prompt paths*, discarding paths once they are deemed irrelevant. We demonstrate the capability of our method to simultaneously enhance time efficiency across a variety of question-answering benchmarks using multiple pre-trained LLMs. Furthermore, our technique significantly improves accuracy when the retrieved context is large relative the context the model was trained on. For example, our approach facilitates a 93x reduction in compute time while *improving* accuracy by 43% on the NaturalQuestions-Open dataset with the MPT-7B instruction-tuned model over naive RAG.
format Preprint
id arxiv_https___arxiv_org_abs_2404_06910
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Superposition Prompting: Improving and Accelerating Retrieval-Augmented Generation
Merth, Thomas
Fu, Qichen
Rastegari, Mohammad
Najibi, Mahyar
Computation and Language
Artificial Intelligence
Machine Learning
Despite the successes of large language models (LLMs), they exhibit significant drawbacks, particularly when processing long contexts. Their inference cost scales quadratically with respect to sequence length, making it expensive for deployment in some real-world text processing applications, such as retrieval-augmented generation (RAG). Additionally, LLMs also exhibit the "distraction phenomenon", where irrelevant context in the prompt degrades output quality. To address these drawbacks, we propose a novel RAG prompting methodology, *superposition prompting*, which can be directly applied to pre-trained transformer-based LLMs *without the need for fine-tuning*. At a high level, superposition prompting allows the LLM to process input documents in parallel *prompt paths*, discarding paths once they are deemed irrelevant. We demonstrate the capability of our method to simultaneously enhance time efficiency across a variety of question-answering benchmarks using multiple pre-trained LLMs. Furthermore, our technique significantly improves accuracy when the retrieved context is large relative the context the model was trained on. For example, our approach facilitates a 93x reduction in compute time while *improving* accuracy by 43% on the NaturalQuestions-Open dataset with the MPT-7B instruction-tuned model over naive RAG.
title Superposition Prompting: Improving and Accelerating Retrieval-Augmented Generation
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2404.06910