RAGO: Systematic Performance Optimization for Retrieval-Augmented Generation Serving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Wenqi, Subramanian, Suvinay, Graves, Cat, Alonso, Gustavo, Yazdanbakhsh, Amir, Dadu, Vidushi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913749260042240
author Jiang, Wenqi
Subramanian, Suvinay
Graves, Cat
Alonso, Gustavo
Yazdanbakhsh, Amir
Dadu, Vidushi
author_facet Jiang, Wenqi
Subramanian, Suvinay
Graves, Cat
Alonso, Gustavo
Yazdanbakhsh, Amir
Dadu, Vidushi
contents Retrieval-augmented generation (RAG), which combines large language models (LLMs) with retrievals from external knowledge databases, is emerging as a popular approach for reliable LLM serving. However, efficient RAG serving remains an open challenge due to the rapid emergence of many RAG variants and the substantial differences in workload characteristics across them. In this paper, we make three fundamental contributions to advancing RAG serving. First, we introduce RAGSchema, a structured abstraction that captures the wide range of RAG algorithms, serving as a foundation for performance optimization. Second, we analyze several representative RAG workloads with distinct RAGSchema, revealing significant performance variability across these workloads. Third, to address this variability and meet diverse performance requirements, we propose RAGO (Retrieval-Augmented Generation Optimizer), a system optimization framework for efficient RAG serving. Our evaluation shows that RAGO achieves up to a 2x increase in QPS per chip and a 55% reduction in time-to-first-token latency compared to RAG systems built on LLM-system extensions.
format Preprint
id arxiv_https___arxiv_org_abs_2503_14649
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RAGO: Systematic Performance Optimization for Retrieval-Augmented Generation Serving
Jiang, Wenqi
Subramanian, Suvinay
Graves, Cat
Alonso, Gustavo
Yazdanbakhsh, Amir
Dadu, Vidushi
Information Retrieval
Artificial Intelligence
Computation and Language
Distributed, Parallel, and Cluster Computing
C.1; C.4; H.3
Retrieval-augmented generation (RAG), which combines large language models (LLMs) with retrievals from external knowledge databases, is emerging as a popular approach for reliable LLM serving. However, efficient RAG serving remains an open challenge due to the rapid emergence of many RAG variants and the substantial differences in workload characteristics across them. In this paper, we make three fundamental contributions to advancing RAG serving. First, we introduce RAGSchema, a structured abstraction that captures the wide range of RAG algorithms, serving as a foundation for performance optimization. Second, we analyze several representative RAG workloads with distinct RAGSchema, revealing significant performance variability across these workloads. Third, to address this variability and meet diverse performance requirements, we propose RAGO (Retrieval-Augmented Generation Optimizer), a system optimization framework for efficient RAG serving. Our evaluation shows that RAGO achieves up to a 2x increase in QPS per chip and a 55% reduction in time-to-first-token latency compared to RAG systems built on LLM-system extensions.
title RAGO: Systematic Performance Optimization for Retrieval-Augmented Generation Serving
topic Information Retrieval
Artificial Intelligence
Computation and Language
Distributed, Parallel, and Cluster Computing
C.1; C.4; H.3
url https://arxiv.org/abs/2503.14649