RAGDoll: Efficient Offloading-based Online RAG System on a Single GPU

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Weiping, Liao, Ningyi, Luo, Siqiang, Liu, Junfeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909587999817728
author Yu, Weiping
Liao, Ningyi
Luo, Siqiang
Liu, Junfeng
author_facet Yu, Weiping
Liao, Ningyi
Luo, Siqiang
Liu, Junfeng
contents Retrieval-Augmented Generation (RAG) enhances large language model (LLM) generation quality by incorporating relevant external knowledge. However, deploying RAG on consumer-grade platforms is challenging due to limited memory and the increasing scale of both models and knowledge bases. In this work, we introduce RAGDoll, a resource-efficient, self-adaptive RAG serving system integrated with LLMs, specifically designed for resource-constrained platforms. RAGDoll exploits the insight that RAG retrieval and LLM generation impose different computational and memory demands, which in a traditional serial workflow result in substantial idle times and poor resource utilization. Based on this insight, RAGDoll decouples retrieval and generation into parallel pipelines, incorporating joint memory placement and dynamic batch scheduling strategies to optimize resource usage across diverse hardware devices and workloads. Extensive experiments demonstrate that RAGDoll adapts effectively to various hardware configurations and LLM scales, achieving up to 3.6 times speedup in average latency compared to serial RAG systems based on vLLM.
format Preprint
id arxiv_https___arxiv_org_abs_2504_15302
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RAGDoll: Efficient Offloading-based Online RAG System on a Single GPU
Yu, Weiping
Liao, Ningyi
Luo, Siqiang
Liu, Junfeng
Distributed, Parallel, and Cluster Computing
Operating Systems
Retrieval-Augmented Generation (RAG) enhances large language model (LLM) generation quality by incorporating relevant external knowledge. However, deploying RAG on consumer-grade platforms is challenging due to limited memory and the increasing scale of both models and knowledge bases. In this work, we introduce RAGDoll, a resource-efficient, self-adaptive RAG serving system integrated with LLMs, specifically designed for resource-constrained platforms. RAGDoll exploits the insight that RAG retrieval and LLM generation impose different computational and memory demands, which in a traditional serial workflow result in substantial idle times and poor resource utilization. Based on this insight, RAGDoll decouples retrieval and generation into parallel pipelines, incorporating joint memory placement and dynamic batch scheduling strategies to optimize resource usage across diverse hardware devices and workloads. Extensive experiments demonstrate that RAGDoll adapts effectively to various hardware configurations and LLM scales, achieving up to 3.6 times speedup in average latency compared to serial RAG systems based on vLLM.
title RAGDoll: Efficient Offloading-based Online RAG System on a Single GPU
topic Distributed, Parallel, and Cluster Computing
Operating Systems
url https://arxiv.org/abs/2504.15302