One Pool, Two Caches: Adaptive HBM Partitioning for Accelerating Generative Recommender Serving
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Wenjun, Han, Shuguang, Zhou, Amelie Chi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ElasticRec: A Microservice-based Model Serving Architecture Enabling Elastic Resource Scaling for Recommendation Models
by: Choi, Yujeong, et al.
Published: (2024)
by: Choi, Yujeong, et al.
Published: (2024)
RecIS: Sparse to Dense, A Unified Training Framework for Recommendation Models
by: Zong, Hua, et al.
Published: (2025)
by: Zong, Hua, et al.
Published: (2025)
G-Meta: Distributed Meta Learning in GPU Clusters for Large-Scale Recommender Systems
by: Xiao, Youshao, et al.
Published: (2024)
by: Xiao, Youshao, et al.
Published: (2024)
From Data to Decisions: The Transformational Power of Machine Learning in Business Recommendations
by: Gangadharan, Kapilya, et al.
Published: (2024)
by: Gangadharan, Kapilya, et al.
Published: (2024)
GPU-accelerated Multi-relational Parallel Graph Retrieval for Web-scale Recommendations
by: Guo, Zhuoning, et al.
Published: (2025)
by: Guo, Zhuoning, et al.
Published: (2025)
Disaggregated Multi-Tower: Topology-aware Modeling Technique for Efficient Large-Scale Recommendation
by: Luo, Liang, et al.
Published: (2024)
by: Luo, Liang, et al.
Published: (2024)
Far From Sight, Far From Mind: Inverse Distance Weighting for Graph Federated Recommendation
by: Khouas, Aymen Rayane, et al.
Published: (2025)
by: Khouas, Aymen Rayane, et al.
Published: (2025)
Near-Zero-Overhead Freshness for Recommendation Systems via Inference-Side Model Updates
by: Yu, Wenjun, et al.
Published: (2025)
by: Yu, Wenjun, et al.
Published: (2025)
Efficient Federated Search for Retrieval-Augmented Generation using Lightweight Routing
by: Dhasade, Akash, et al.
Published: (2025)
by: Dhasade, Akash, et al.
Published: (2025)
Curator: Efficient Indexing for Multi-Tenant Vector Databases
by: Jin, Yicheng, et al.
Published: (2024)
by: Jin, Yicheng, et al.
Published: (2024)
ERCache: An Efficient and Reliable Caching Framework for Large-Scale User Representations in Meta's Ads System
by: Zhou, Fang, et al.
Published: (2024)
by: Zhou, Fang, et al.
Published: (2024)
Relay Buffer Independent Communication over Pooled HBM for Efficient MoE Inference on Ascend
by: Hu, Tianlun, et al.
Published: (2026)
by: Hu, Tianlun, et al.
Published: (2026)
RcLLM: Accelerating Generative Recommendation via Beyond-Prefix KV Caching
by: Zhao, Zhan, et al.
Published: (2026)
by: Zhao, Zhan, et al.
Published: (2026)
FedGrAINS: Personalized SubGraph Federated Learning with Adaptive Neighbor Sampling
by: Ceyani, Emir, et al.
Published: (2025)
by: Ceyani, Emir, et al.
Published: (2025)
SaberLDA: Sparsity-Aware Learning of Topic Models on GPUs
by: Li, Kaiwei, et al.
Published: (2016)
by: Li, Kaiwei, et al.
Published: (2016)
Stalactite: Toolbox for Fast Prototyping of Vertical Federated Learning Systems
by: Zakharova, Anastasiia, et al.
Published: (2024)
by: Zakharova, Anastasiia, et al.
Published: (2024)
FedFlex: Federated Learning for Diverse Netflix Recommendations
by: Lankester, Sven, et al.
Published: (2025)
by: Lankester, Sven, et al.
Published: (2025)
FreeScale: Distributed Training for Sequence Recommendation Models with Minimal Scaling Cost
by: Feng, Chenhao, et al.
Published: (2026)
by: Feng, Chenhao, et al.
Published: (2026)
PIFS-Rec: Process-In-Fabric-Switch for Large-Scale Recommendation System Inferences
by: Huo, Pingyi, et al.
Published: (2024)
by: Huo, Pingyi, et al.
Published: (2024)
Towards Efficient Communication and Secure Federated Recommendation System via Low-rank Training
by: Nguyen, Ngoc-Hieu, et al.
Published: (2024)
by: Nguyen, Ngoc-Hieu, et al.
Published: (2024)
FedPDD: A Privacy-preserving Double Distillation Framework for Cross-silo Federated Recommendation
by: Wan, Sheng, et al.
Published: (2023)
by: Wan, Sheng, et al.
Published: (2023)
AdaParse: An Adaptive Parallel PDF Parsing and Resource Scaling Engine
by: Siebenschuh, Carlo, et al.
Published: (2025)
by: Siebenschuh, Carlo, et al.
Published: (2025)
SocFedGPT: Federated GPT-based Adaptive Content Filtering System Leveraging User Interactions in Social Networks
by: Puppala, Sai, et al.
Published: (2024)
by: Puppala, Sai, et al.
Published: (2024)
A Model-agnostic Strategy to Mitigate Embedding Degradation in Personalized Federated Recommendation
by: Shen, Jiakui, et al.
Published: (2025)
by: Shen, Jiakui, et al.
Published: (2025)
Efficient Distributed Retrieval-Augmented Generation for Enhancing Language Model Performance
by: Liu, Shangyu, et al.
Published: (2025)
by: Liu, Shangyu, et al.
Published: (2025)
Robust Implementation of Retrieval-Augmented Generation on Edge-based Computing-in-Memory Architectures
by: Qin, Ruiyang, et al.
Published: (2024)
by: Qin, Ruiyang, et al.
Published: (2024)
TokenLake: A Unified Segment-level Prefix Cache Pool for Fine-grained Elastic Long-Context LLM Serving
by: Wu, Bingyang, et al.
Published: (2025)
by: Wu, Bingyang, et al.
Published: (2025)
Robust DNN Partitioning and Resource Allocation Under Uncertain Inference Time
by: Nan, Zhaojun, et al.
Published: (2025)
by: Nan, Zhaojun, et al.
Published: (2025)
C-FedRAG: A Confidential Federated Retrieval-Augmented Generation System
by: Addison, Parker, et al.
Published: (2024)
by: Addison, Parker, et al.
Published: (2024)
FLASH: Federated Learning-Based LLMs for Advanced Query Processing in Social Networks through RAG
by: Puppala, Sai, et al.
Published: (2024)
by: Puppala, Sai, et al.
Published: (2024)
Feature Noise Resilient for QoS Prediction with Probabilistic Deep Supervision
by: Wang, Ziliang, et al.
Published: (2023)
by: Wang, Ziliang, et al.
Published: (2023)
Superior Parallel Big Data Clustering through Competitive Stochastic Sample Size Optimization in Big-means
by: Mussabayev, Rustam, et al.
Published: (2024)
by: Mussabayev, Rustam, et al.
Published: (2024)
MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
by: Hu, Cunchen, et al.
Published: (2024)
by: Hu, Cunchen, et al.
Published: (2024)
Intelligent Model Update Strategy for Sequential Recommendation
by: Lv, Zheqi, et al.
Published: (2023)
by: Lv, Zheqi, et al.
Published: (2023)
ShadowServe: Interference-Free KV Cache Fetching for Distributed Prefix Caching
by: Xiang, Xingyu, et al.
Published: (2025)
by: Xiang, Xingyu, et al.
Published: (2025)
Closing the Generalization Gap in Parameter-efficient Federated Edge Learning
by: Du, Xinnong, et al.
Published: (2025)
by: Du, Xinnong, et al.
Published: (2025)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
by: Su, Zhaoyuan, et al.
Published: (2025)
by: Su, Zhaoyuan, et al.
Published: (2025)
Exact Nearest-Neighbor Search on Energy-Efficient FPGA Devices
by: Dazzi, Patrizio, et al.
Published: (2025)
by: Dazzi, Patrizio, et al.
Published: (2025)
A Big Data Architecture for Early Identification and Categorization of Dark Web Sites
by: Pastor-Galindo, Javier, et al.
Published: (2024)
by: Pastor-Galindo, Javier, et al.
Published: (2024)
An OPC UA-based industrial Big Data architecture
by: Hirsch, Eduard, et al.
Published: (2023)
by: Hirsch, Eduard, et al.
Published: (2023)
Similar Items
-
ElasticRec: A Microservice-based Model Serving Architecture Enabling Elastic Resource Scaling for Recommendation Models
by: Choi, Yujeong, et al.
Published: (2024) -
RecIS: Sparse to Dense, A Unified Training Framework for Recommendation Models
by: Zong, Hua, et al.
Published: (2025) -
G-Meta: Distributed Meta Learning in GPU Clusters for Large-Scale Recommender Systems
by: Xiao, Youshao, et al.
Published: (2024) -
From Data to Decisions: The Transformational Power of Machine Learning in Business Recommendations
by: Gangadharan, Kapilya, et al.
Published: (2024) -
GPU-accelerated Multi-relational Parallel Graph Retrieval for Web-scale Recommendations
by: Guo, Zhuoning, et al.
Published: (2025)