ElasticRec: A Microservice-based Model Serving Architecture Enabling Elastic Resource Scaling for Recommendation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Choi, Yujeong, Kim, Jiin, Rhu, Minsoo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers
by: Yeo, Gwangoo, et al.
Published: (2024)
by: Yeo, Gwangoo, et al.
Published: (2024)
RecIS: Sparse to Dense, A Unified Training Framework for Recommendation Models
by: Zong, Hua, et al.
Published: (2025)
by: Zong, Hua, et al.
Published: (2025)
PIFS-Rec: Process-In-Fabric-Switch for Large-Scale Recommendation System Inferences
by: Huo, Pingyi, et al.
Published: (2024)
by: Huo, Pingyi, et al.
Published: (2024)
One Pool, Two Caches: Adaptive HBM Partitioning for Accelerating Generative Recommender Serving
by: Yu, Wenjun, et al.
Published: (2026)
by: Yu, Wenjun, et al.
Published: (2026)
Disaggregated Multi-Tower: Topology-aware Modeling Technique for Efficient Large-Scale Recommendation
by: Luo, Liang, et al.
Published: (2024)
by: Luo, Liang, et al.
Published: (2024)
G-Meta: Distributed Meta Learning in GPU Clusters for Large-Scale Recommender Systems
by: Xiao, Youshao, et al.
Published: (2024)
by: Xiao, Youshao, et al.
Published: (2024)
FreeScale: Distributed Training for Sequence Recommendation Models with Minimal Scaling Cost
by: Feng, Chenhao, et al.
Published: (2026)
by: Feng, Chenhao, et al.
Published: (2026)
From Data to Decisions: The Transformational Power of Machine Learning in Business Recommendations
by: Gangadharan, Kapilya, et al.
Published: (2024)
by: Gangadharan, Kapilya, et al.
Published: (2024)
GPU-accelerated Multi-relational Parallel Graph Retrieval for Web-scale Recommendations
by: Guo, Zhuoning, et al.
Published: (2025)
by: Guo, Zhuoning, et al.
Published: (2025)
ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism
by: Liu, Zedong, et al.
Published: (2025)
by: Liu, Zedong, et al.
Published: (2025)
Far From Sight, Far From Mind: Inverse Distance Weighting for Graph Federated Recommendation
by: Khouas, Aymen Rayane, et al.
Published: (2025)
by: Khouas, Aymen Rayane, et al.
Published: (2025)
SaberLDA: Sparsity-Aware Learning of Topic Models on GPUs
by: Li, Kaiwei, et al.
Published: (2016)
by: Li, Kaiwei, et al.
Published: (2016)
LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
by: Wu, Bingyang, et al.
Published: (2024)
by: Wu, Bingyang, et al.
Published: (2024)
AdaParse: An Adaptive Parallel PDF Parsing and Resource Scaling Engine
by: Siebenschuh, Carlo, et al.
Published: (2025)
by: Siebenschuh, Carlo, et al.
Published: (2025)
Enabling Elastic Model Serving with MultiWorld
by: Lee, Myungjin, et al.
Published: (2024)
by: Lee, Myungjin, et al.
Published: (2024)
Robust Implementation of Retrieval-Augmented Generation on Edge-based Computing-in-Memory Architectures
by: Qin, Ruiyang, et al.
Published: (2024)
by: Qin, Ruiyang, et al.
Published: (2024)
A Model-agnostic Strategy to Mitigate Embedding Degradation in Personalized Federated Recommendation
by: Shen, Jiakui, et al.
Published: (2025)
by: Shen, Jiakui, et al.
Published: (2025)
A Document-based Knowledge Discovery with Microservices Architecture
by: Gidey, Habtom Kahsay, et al.
Published: (2024)
by: Gidey, Habtom Kahsay, et al.
Published: (2024)
Stalactite: Toolbox for Fast Prototyping of Vertical Federated Learning Systems
by: Zakharova, Anastasiia, et al.
Published: (2024)
by: Zakharova, Anastasiia, et al.
Published: (2024)
Efficient Federated Search for Retrieval-Augmented Generation using Lightweight Routing
by: Dhasade, Akash, et al.
Published: (2025)
by: Dhasade, Akash, et al.
Published: (2025)
FedFlex: Federated Learning for Diverse Netflix Recommendations
by: Lankester, Sven, et al.
Published: (2025)
by: Lankester, Sven, et al.
Published: (2025)
Towards Efficient Communication and Secure Federated Recommendation System via Low-rank Training
by: Nguyen, Ngoc-Hieu, et al.
Published: (2024)
by: Nguyen, Ngoc-Hieu, et al.
Published: (2024)
FedPDD: A Privacy-preserving Double Distillation Framework for Cross-silo Federated Recommendation
by: Wan, Sheng, et al.
Published: (2023)
by: Wan, Sheng, et al.
Published: (2023)
Efficient Distributed Retrieval-Augmented Generation for Enhancing Language Model Performance
by: Liu, Shangyu, et al.
Published: (2025)
by: Liu, Shangyu, et al.
Published: (2025)
ERCache: An Efficient and Reliable Caching Framework for Large-Scale User Representations in Meta's Ads System
by: Zhou, Fang, et al.
Published: (2024)
by: Zhou, Fang, et al.
Published: (2024)
HyGen: Efficient LLM Serving via Elastic Online-Offline Request Co-location
by: Sun, Ting, et al.
Published: (2025)
by: Sun, Ting, et al.
Published: (2025)
StatuScale: Status-aware and Elastic Scaling Strategy for Microservice Applications
by: Wen, Linfeng, et al.
Published: (2024)
by: Wen, Linfeng, et al.
Published: (2024)
Lazarus: Resilient and Elastic Training of Mixture-of-Experts Models
by: Wu, Yongji, et al.
Published: (2024)
by: Wu, Yongji, et al.
Published: (2024)
Curator: Efficient Indexing for Multi-Tenant Vector Databases
by: Jin, Yicheng, et al.
Published: (2024)
by: Jin, Yicheng, et al.
Published: (2024)
A Big Data Architecture for Early Identification and Categorization of Dark Web Sites
by: Pastor-Galindo, Javier, et al.
Published: (2024)
by: Pastor-Galindo, Javier, et al.
Published: (2024)
Intelligent Model Update Strategy for Sequential Recommendation
by: Lv, Zheqi, et al.
Published: (2023)
by: Lv, Zheqi, et al.
Published: (2023)
SocFedGPT: Federated GPT-based Adaptive Content Filtering System Leveraging User Interactions in Social Networks
by: Puppala, Sai, et al.
Published: (2024)
by: Puppala, Sai, et al.
Published: (2024)
Dilu: Enabling GPU Resourcing-on-Demand for Serverless DL Serving via Introspective Elasticity
by: Lv, Cunchi, et al.
Published: (2025)
by: Lv, Cunchi, et al.
Published: (2025)
TokenLake: A Unified Segment-level Prefix Cache Pool for Fine-grained Elastic Long-Context LLM Serving
by: Wu, Bingyang, et al.
Published: (2025)
by: Wu, Bingyang, et al.
Published: (2025)
An OPC UA-based industrial Big Data architecture
by: Hirsch, Eduard, et al.
Published: (2023)
by: Hirsch, Eduard, et al.
Published: (2023)
ELMoE-3D: Leveraging Intrinsic Elasticity of MoE for Hybrid-Bonding-Enabled Self-Speculative Decoding in On-Premises Serving
by: Choi, Yuseon, et al.
Published: (2026)
by: Choi, Yuseon, et al.
Published: (2026)
FLASH: Federated Learning-Based LLMs for Advanced Query Processing in Social Networks through RAG
by: Puppala, Sai, et al.
Published: (2024)
by: Puppala, Sai, et al.
Published: (2024)
Feature Noise Resilient for QoS Prediction with Probabilistic Deep Supervision
by: Wang, Ziliang, et al.
Published: (2023)
by: Wang, Ziliang, et al.
Published: (2023)
FedGrAINS: Personalized SubGraph Federated Learning with Adaptive Neighbor Sampling
by: Ceyani, Emir, et al.
Published: (2025)
by: Ceyani, Emir, et al.
Published: (2025)
Superior Parallel Big Data Clustering through Competitive Stochastic Sample Size Optimization in Big-means
by: Mussabayev, Rustam, et al.
Published: (2024)
by: Mussabayev, Rustam, et al.
Published: (2024)
Similar Items
-
PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers
by: Yeo, Gwangoo, et al.
Published: (2024) -
RecIS: Sparse to Dense, A Unified Training Framework for Recommendation Models
by: Zong, Hua, et al.
Published: (2025) -
PIFS-Rec: Process-In-Fabric-Switch for Large-Scale Recommendation System Inferences
by: Huo, Pingyi, et al.
Published: (2024) -
One Pool, Two Caches: Adaptive HBM Partitioning for Accelerating Generative Recommender Serving
by: Yu, Wenjun, et al.
Published: (2026) -
Disaggregated Multi-Tower: Topology-aware Modeling Technique for Efficient Large-Scale Recommendation
by: Luo, Liang, et al.
Published: (2024)