RAGO: Systematic Performance Optimization for Retrieval-Augmented Generation Serving
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Wenqi, Subramanian, Suvinay, Graves, Cat, Alonso, Gustavo, Yazdanbakhsh, Amir, Dadu, Vidushi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
C-FedRAG: A Confidential Federated Retrieval-Augmented Generation System
by: Addison, Parker, et al.
Published: (2024)
by: Addison, Parker, et al.
Published: (2024)
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
by: He, Jiaao, et al.
Published: (2024)
by: He, Jiaao, et al.
Published: (2024)
Efficient Federated Search for Retrieval-Augmented Generation using Lightweight Routing
by: Dhasade, Akash, et al.
Published: (2025)
by: Dhasade, Akash, et al.
Published: (2025)
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving
by: Wu, Hanjiang, et al.
Published: (2026)
by: Wu, Hanjiang, et al.
Published: (2026)
TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication
by: Wang, Zongwu, et al.
Published: (2024)
by: Wang, Zongwu, et al.
Published: (2024)
Learning to Keep a Promise: Scaling Language Model Decoding Parallelism with Learned Asynchronous Decoding
by: Jin, Tian, et al.
Published: (2025)
by: Jin, Tian, et al.
Published: (2025)
Efficient Distributed Retrieval-Augmented Generation for Enhancing Language Model Performance
by: Liu, Shangyu, et al.
Published: (2025)
by: Liu, Shangyu, et al.
Published: (2025)
One Pool, Two Caches: Adaptive HBM Partitioning for Accelerating Generative Recommender Serving
by: Yu, Wenjun, et al.
Published: (2026)
by: Yu, Wenjun, et al.
Published: (2026)
Federated Cross-Domain Click-Through Rate Prediction With Large Language Model Augmentation
by: Qin, Jiangcheng, et al.
Published: (2025)
by: Qin, Jiangcheng, et al.
Published: (2025)
ElasticRec: A Microservice-based Model Serving Architecture Enabling Elastic Resource Scaling for Recommendation Models
by: Choi, Yujeong, et al.
Published: (2024)
by: Choi, Yujeong, et al.
Published: (2024)
Limitless FaaS: Overcoming serverless functions execution time limits with invoke driven architecture and memory checkpoints
by: Andraca, Rodrigo Landa, et al.
Published: (2024)
by: Andraca, Rodrigo Landa, et al.
Published: (2024)
Demystifying NCCL: An In-depth Analysis of GPU Communication Protocols and Algorithms
by: Hu, Zhiyi, et al.
Published: (2025)
by: Hu, Zhiyi, et al.
Published: (2025)
Efficient and Reuseable Cloud Configuration Search Using Discovery Spaces
by: Johnston, Michael, et al.
Published: (2025)
by: Johnston, Michael, et al.
Published: (2025)
Chat3GPP: An Open-Source Retrieval-Augmented Generation Framework for 3GPP Documents
by: Huang, Long, et al.
Published: (2025)
by: Huang, Long, et al.
Published: (2025)
COSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems
by: Raju, Aditi, et al.
Published: (2025)
by: Raju, Aditi, et al.
Published: (2025)
From Data to Decisions: The Transformational Power of Machine Learning in Business Recommendations
by: Gangadharan, Kapilya, et al.
Published: (2024)
by: Gangadharan, Kapilya, et al.
Published: (2024)
Characterizing the Dilemma of Performance and Index Size in Billion-Scale Vector Search and Breaking It with Second-Tier Memory
by: Cheng, Rongxin, et al.
Published: (2024)
by: Cheng, Rongxin, et al.
Published: (2024)
A Model-agnostic Strategy to Mitigate Embedding Degradation in Personalized Federated Recommendation
by: Shen, Jiakui, et al.
Published: (2025)
by: Shen, Jiakui, et al.
Published: (2025)
Exact Nearest-Neighbor Search on Energy-Efficient FPGA Devices
by: Dazzi, Patrizio, et al.
Published: (2025)
by: Dazzi, Patrizio, et al.
Published: (2025)
A Big Data Architecture for Early Identification and Categorization of Dark Web Sites
by: Pastor-Galindo, Javier, et al.
Published: (2024)
by: Pastor-Galindo, Javier, et al.
Published: (2024)
An OPC UA-based industrial Big Data architecture
by: Hirsch, Eduard, et al.
Published: (2023)
by: Hirsch, Eduard, et al.
Published: (2023)
Accelerating Retrieval-Augmented Generation
by: Quinn, Derrick, et al.
Published: (2024)
by: Quinn, Derrick, et al.
Published: (2024)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
by: Jaiswal, Shashwat, et al.
Published: (2025)
by: Jaiswal, Shashwat, et al.
Published: (2025)
GPU-accelerated Multi-relational Parallel Graph Retrieval for Web-scale Recommendations
by: Guo, Zhuoning, et al.
Published: (2025)
by: Guo, Zhuoning, et al.
Published: (2025)
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
by: Jiang, Youhe, et al.
Published: (2026)
by: Jiang, Youhe, et al.
Published: (2026)
DeepServe: Serverless Large Language Model Serving at Scale
by: Hu, Junhao, et al.
Published: (2025)
by: Hu, Junhao, et al.
Published: (2025)
FPGA-Accelerated Lock Management and Transaction Processing: Architecture, Optimization, and Design Space Exploration
by: Zhu, Shien, et al.
Published: (2026)
by: Zhu, Shien, et al.
Published: (2026)
ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
by: Jiang, Youhe, et al.
Published: (2025)
by: Jiang, Youhe, et al.
Published: (2025)
TridentServe: A Stage-level Serving System for Diffusion Pipelines
by: Xia, Yifei, et al.
Published: (2025)
by: Xia, Yifei, et al.
Published: (2025)
Robust Implementation of Retrieval-Augmented Generation on Edge-based Computing-in-Memory Architectures
by: Qin, Ruiyang, et al.
Published: (2024)
by: Qin, Ruiyang, et al.
Published: (2024)
Distributed Retrieval-Augmented Generation
by: Xu, Chenhao, et al.
Published: (2025)
by: Xu, Chenhao, et al.
Published: (2025)
LLaMCAT: Optimizing Large Language Model Inference with Cache Arbitration and Throttling
by: Zhou, Zhongchun, et al.
Published: (2025)
by: Zhou, Zhongchun, et al.
Published: (2025)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
by: Hu, Cunchen, et al.
Published: (2024)
by: Hu, Cunchen, et al.
Published: (2024)
Performance Evaluation of LLMs in Automated RDF Knowledge Graph Generation
by: Martin, Ioana Ramona, et al.
Published: (2026)
by: Martin, Ioana Ramona, et al.
Published: (2026)
eScope: A Fine-Grained Power Prediction Mechanism for Mobile Applications
by: Mukherjee, Dipayan, et al.
Published: (2024)
by: Mukherjee, Dipayan, et al.
Published: (2024)
SkyNomad: On Using Multi-Region Spot Instances to Minimize AI Batch Job Cost
by: Li, Zhifei, et al.
Published: (2026)
by: Li, Zhifei, et al.
Published: (2026)
Distributed Recoverable Sketches (Extended Version)
by: Cohen, Diana, et al.
Published: (2025)
by: Cohen, Diana, et al.
Published: (2025)
Intersections of Web3 and AI -- View in 2024
by: Hyland-Wood, David, et al.
Published: (2024)
by: Hyland-Wood, David, et al.
Published: (2024)
Artifact Evaluation for Distributed Systems: Current Practices and Beyond
by: Sedghpour, Mohammad Reza Saleh, et al.
Published: (2024)
by: Sedghpour, Mohammad Reza Saleh, et al.
Published: (2024)
Joint Training on AMD and NVIDIA GPUs
by: Hu, Jon, et al.
Published: (2026)
by: Hu, Jon, et al.
Published: (2026)
Similar Items
-
C-FedRAG: A Confidential Federated Retrieval-Augmented Generation System
by: Addison, Parker, et al.
Published: (2024) -
FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
by: He, Jiaao, et al.
Published: (2024) -
Efficient Federated Search for Retrieval-Augmented Generation using Lightweight Routing
by: Dhasade, Akash, et al.
Published: (2025) -
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving
by: Wu, Hanjiang, et al.
Published: (2026) -
TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication
by: Wang, Zongwu, et al.
Published: (2024)