Efficient Distributed Retrieval-Augmented Generation for Enhancing Language Model Performance
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Shangyu, Zheng, Zhenzhe, Huang, Xiaoyao, Wu, Fan, Chen, Guihai, Wu, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
C-FedRAG: A Confidential Federated Retrieval-Augmented Generation System
by: Addison, Parker, et al.
Published: (2024)
by: Addison, Parker, et al.
Published: (2024)
Efficient Federated Search for Retrieval-Augmented Generation using Lightweight Routing
by: Dhasade, Akash, et al.
Published: (2025)
by: Dhasade, Akash, et al.
Published: (2025)
Federated Cross-Domain Click-Through Rate Prediction With Large Language Model Augmentation
by: Qin, Jiangcheng, et al.
Published: (2025)
by: Qin, Jiangcheng, et al.
Published: (2025)
Forward Once for All: Structural Parameterized Adaptation for Efficient Cloud-coordinated On-device Recommendation
by: Fu, Kairui, et al.
Published: (2025)
by: Fu, Kairui, et al.
Published: (2025)
Robust Implementation of Retrieval-Augmented Generation on Edge-based Computing-in-Memory Architectures
by: Qin, Ruiyang, et al.
Published: (2024)
by: Qin, Ruiyang, et al.
Published: (2024)
Exact Nearest-Neighbor Search on Energy-Efficient FPGA Devices
by: Dazzi, Patrizio, et al.
Published: (2025)
by: Dazzi, Patrizio, et al.
Published: (2025)
GPU-accelerated Multi-relational Parallel Graph Retrieval for Web-scale Recommendations
by: Guo, Zhuoning, et al.
Published: (2025)
by: Guo, Zhuoning, et al.
Published: (2025)
Chat3GPP: An Open-Source Retrieval-Augmented Generation Framework for 3GPP Documents
by: Huang, Long, et al.
Published: (2025)
by: Huang, Long, et al.
Published: (2025)
MIRAGE: Runtime Scheduling for Multi-Vector Image Retrieval with Hierarchical Decomposition
by: Li, Maoliang, et al.
Published: (2025)
by: Li, Maoliang, et al.
Published: (2025)
A Model-agnostic Strategy to Mitigate Embedding Degradation in Personalized Federated Recommendation
by: Shen, Jiakui, et al.
Published: (2025)
by: Shen, Jiakui, et al.
Published: (2025)
Characterizing the Dilemma of Performance and Index Size in Billion-Scale Vector Search and Breaking It with Second-Tier Memory
by: Cheng, Rongxin, et al.
Published: (2024)
by: Cheng, Rongxin, et al.
Published: (2024)
Enhancing Cloud-Based Large Language Model Processing with Elasticsearch and Transformer Models
by: Ni, Chunhe, et al.
Published: (2024)
by: Ni, Chunhe, et al.
Published: (2024)
Performance Evaluation of LLMs in Automated RDF Knowledge Graph Generation
by: Martin, Ioana Ramona, et al.
Published: (2026)
by: Martin, Ioana Ramona, et al.
Published: (2026)
A Big Data Architecture for Early Identification and Categorization of Dark Web Sites
by: Pastor-Galindo, Javier, et al.
Published: (2024)
by: Pastor-Galindo, Javier, et al.
Published: (2024)
An OPC UA-based industrial Big Data architecture
by: Hirsch, Eduard, et al.
Published: (2023)
by: Hirsch, Eduard, et al.
Published: (2023)
Accelerating Retrieval-Augmented Generation
by: Quinn, Derrick, et al.
Published: (2024)
by: Quinn, Derrick, et al.
Published: (2024)
Disaggregated Multi-Tower: Topology-aware Modeling Technique for Efficient Large-Scale Recommendation
by: Luo, Liang, et al.
Published: (2024)
by: Luo, Liang, et al.
Published: (2024)
RecIS: Sparse to Dense, A Unified Training Framework for Recommendation Models
by: Zong, Hua, et al.
Published: (2025)
by: Zong, Hua, et al.
Published: (2025)
Curator: Efficient Indexing for Multi-Tenant Vector Databases
by: Jin, Yicheng, et al.
Published: (2024)
by: Jin, Yicheng, et al.
Published: (2024)
SaberLDA: Sparsity-Aware Learning of Topic Models on GPUs
by: Li, Kaiwei, et al.
Published: (2016)
by: Li, Kaiwei, et al.
Published: (2016)
G-Meta: Distributed Meta Learning in GPU Clusters for Large-Scale Recommender Systems
by: Xiao, Youshao, et al.
Published: (2024)
by: Xiao, Youshao, et al.
Published: (2024)
Intelligent Model Update Strategy for Sequential Recommendation
by: Lv, Zheqi, et al.
Published: (2023)
by: Lv, Zheqi, et al.
Published: (2023)
SIVF: GPU-Resident IVF Index for Streaming Vector Search
by: Zhao, Dongfang
Published: (2026)
by: Zhao, Dongfang
Published: (2026)
A Context-Aware Knowledge Graph Platform for Stream Processing in Industrial IoT
by: Sciarroni, Monica Marconi, et al.
Published: (2026)
by: Sciarroni, Monica Marconi, et al.
Published: (2026)
PiPNN: Ultra-Scalable Graph-Based Nearest Neighbor Indexing
by: Rubel, Tobias, et al.
Published: (2026)
by: Rubel, Tobias, et al.
Published: (2026)
Data Dams: A Novel Framework for Regulating and Managing Data Flow in Large-Scale Systems
by: Bouke, Mohamed Aly, et al.
Published: (2025)
by: Bouke, Mohamed Aly, et al.
Published: (2025)
Distributed Retrieval-Augmented Generation
by: Xu, Chenhao, et al.
Published: (2025)
by: Xu, Chenhao, et al.
Published: (2025)
Collaboration of Large Language Models and Small Recommendation Models for Device-Cloud Recommendation
by: Lv, Zheqi, et al.
Published: (2025)
by: Lv, Zheqi, et al.
Published: (2025)
One Pool, Two Caches: Adaptive HBM Partitioning for Accelerating Generative Recommender Serving
by: Yu, Wenjun, et al.
Published: (2026)
by: Yu, Wenjun, et al.
Published: (2026)
ElasticRec: A Microservice-based Model Serving Architecture Enabling Elastic Resource Scaling for Recommendation Models
by: Choi, Yujeong, et al.
Published: (2024)
by: Choi, Yujeong, et al.
Published: (2024)
DIET: Customized Slimming for Incompatible Networks in Sequential Recommendation
by: Fu, Kairui, et al.
Published: (2024)
by: Fu, Kairui, et al.
Published: (2024)
FreeScale: Distributed Training for Sequence Recommendation Models with Minimal Scaling Cost
by: Feng, Chenhao, et al.
Published: (2026)
by: Feng, Chenhao, et al.
Published: (2026)
DUET: A Tuning-Free Device-Cloud Collaborative Parameters Generation Framework for Efficient Device Model Generalization
by: Lv, Zheqi, et al.
Published: (2022)
by: Lv, Zheqi, et al.
Published: (2022)
Stalactite: Toolbox for Fast Prototyping of Vertical Federated Learning Systems
by: Zakharova, Anastasiia, et al.
Published: (2024)
by: Zakharova, Anastasiia, et al.
Published: (2024)
From Data to Decisions: The Transformational Power of Machine Learning in Business Recommendations
by: Gangadharan, Kapilya, et al.
Published: (2024)
by: Gangadharan, Kapilya, et al.
Published: (2024)
Far From Sight, Far From Mind: Inverse Distance Weighting for Graph Federated Recommendation
by: Khouas, Aymen Rayane, et al.
Published: (2025)
by: Khouas, Aymen Rayane, et al.
Published: (2025)
A Large-Scale Web Search Dataset for Federated Online Learning to Rank
by: Gregoriadis, Marcel, et al.
Published: (2025)
by: Gregoriadis, Marcel, et al.
Published: (2025)
Alto: Orchestrating Distributed Compound AI Systems with Nested Ancestry
by: Raghavan, Deepti, et al.
Published: (2024)
by: Raghavan, Deepti, et al.
Published: (2024)
ERCache: An Efficient and Reliable Caching Framework for Large-Scale User Representations in Meta's Ads System
by: Zhou, Fang, et al.
Published: (2024)
by: Zhou, Fang, et al.
Published: (2024)
RAGO: Systematic Performance Optimization for Retrieval-Augmented Generation Serving
by: Jiang, Wenqi, et al.
Published: (2025)
by: Jiang, Wenqi, et al.
Published: (2025)
Similar Items
-
C-FedRAG: A Confidential Federated Retrieval-Augmented Generation System
by: Addison, Parker, et al.
Published: (2024) -
Efficient Federated Search for Retrieval-Augmented Generation using Lightweight Routing
by: Dhasade, Akash, et al.
Published: (2025) -
Federated Cross-Domain Click-Through Rate Prediction With Large Language Model Augmentation
by: Qin, Jiangcheng, et al.
Published: (2025) -
Forward Once for All: Structural Parameterized Adaptation for Efficient Cloud-coordinated On-device Recommendation
by: Fu, Kairui, et al.
Published: (2025) -
Robust Implementation of Retrieval-Augmented Generation on Edge-based Computing-in-Memory Architectures
by: Qin, Ruiyang, et al.
Published: (2024)