Near-Zero-Overhead Freshness for Recommendation Systems via Inference-Side Model Updates
Fuente:
arXiv
Salvato in:
| Autori principali: | Yu, Wenjun, Chen, Sitian, Chen, Cheng, Zhou, Amelie Chi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
One Pool, Two Caches: Adaptive HBM Partitioning for Accelerating Generative Recommender Serving
di: Yu, Wenjun, et al.
Pubblicazione: (2026)
di: Yu, Wenjun, et al.
Pubblicazione: (2026)
Faster Distributed Inference-Only Recommender Systems via Bounded Lag Synchronous Collectives
di: Dichev, Kiril, et al.
Pubblicazione: (2025)
di: Dichev, Kiril, et al.
Pubblicazione: (2025)
TaxBreak: Unmasking the Hidden Costs of LLM Inference Through Overhead Decomposition
di: Vellaisamy, Prabhu, et al.
Pubblicazione: (2026)
di: Vellaisamy, Prabhu, et al.
Pubblicazione: (2026)
RcLLM: Accelerating Generative Recommendation via Beyond-Prefix KV Caching
di: Zhao, Zhan, et al.
Pubblicazione: (2026)
di: Zhao, Zhan, et al.
Pubblicazione: (2026)
Lion Cub: Minimizing Communication Overhead in Distributed Lion
di: Ishikawa, Satoki, et al.
Pubblicazione: (2024)
di: Ishikawa, Satoki, et al.
Pubblicazione: (2024)
KaMPIng: Flexible and (Near) Zero-Overhead C++ Bindings for MPI
di: Uhl, Tim Niklas, et al.
Pubblicazione: (2024)
di: Uhl, Tim Niklas, et al.
Pubblicazione: (2024)
Enabling Disaggregated Multi-Stage MLLM Inference via GPU-Internal Scheduling and Resource Sharing
di: Zhao, Lingxiao, et al.
Pubblicazione: (2025)
di: Zhao, Lingxiao, et al.
Pubblicazione: (2025)
Checkmate: Zero-Overhead Model Checkpointing via Network Gradient Replication
di: Bhardwaj, Ankit, et al.
Pubblicazione: (2025)
di: Bhardwaj, Ankit, et al.
Pubblicazione: (2025)
Adaptive Compression in Federated Learning via Side Information
di: Isik, Berivan, et al.
Pubblicazione: (2023)
di: Isik, Berivan, et al.
Pubblicazione: (2023)
Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training and Inference
di: Luo, Shuqing, et al.
Pubblicazione: (2025)
di: Luo, Shuqing, et al.
Pubblicazione: (2025)
Harli: SLO-Aware Co-location of LLM Inference and PEFT-based Finetuning on Model-as-a-Service Platforms
di: Xu, Ao, et al.
Pubblicazione: (2025)
di: Xu, Ao, et al.
Pubblicazione: (2025)
FedCAda: Adaptive Client-Side Optimization for Accelerated and Stable Federated Learning
di: Zhou, Liuzhi, et al.
Pubblicazione: (2024)
di: Zhou, Liuzhi, et al.
Pubblicazione: (2024)
Reducing Communication Overhead in Federated Learning for Network Anomaly Detection with Adaptive Client Selection
di: Marfo, William, et al.
Pubblicazione: (2025)
di: Marfo, William, et al.
Pubblicazione: (2025)
Injecting Adrenaline into LLM Serving: Boosting Resource Utilization and Throughput via Attention Disaggregation
di: Liang, Yunkai, et al.
Pubblicazione: (2025)
di: Liang, Yunkai, et al.
Pubblicazione: (2025)
InkStream: Real-time GNN Inference on Streaming Graphs via Incremental Update
di: Wu, Dan, et al.
Pubblicazione: (2023)
di: Wu, Dan, et al.
Pubblicazione: (2023)
Characterizing WebGPU Dispatch Overhead for LLM Inference Across Four GPU Vendors, Three Backends, and Three Browsers
di: Maczan, Jędrzej
Pubblicazione: (2026)
di: Maczan, Jędrzej
Pubblicazione: (2026)
MiCRO: Near-Zero Cost Gradient Sparsification for Scaling and Accelerating Distributed DNN Training
di: Yoon, Daegun, et al.
Pubblicazione: (2023)
di: Yoon, Daegun, et al.
Pubblicazione: (2023)
Boosting Resource-Constrained Federated Learning Systems with Guessed Updates
di: Boukhari, Mohamed Yassine, et al.
Pubblicazione: (2021)
di: Boukhari, Mohamed Yassine, et al.
Pubblicazione: (2021)
Towards Integrated Fine-tuning and Inference when Generative AI meets Edge Intelligence
di: Chen, Ning, et al.
Pubblicazione: (2024)
di: Chen, Ning, et al.
Pubblicazione: (2024)
MUSE: Multi-Tenant Model Serving With Seamless Model Updates
di: Correia, Cláudio, et al.
Pubblicazione: (2026)
di: Correia, Cláudio, et al.
Pubblicazione: (2026)
Straggler-Resilient Decentralized Learning via Adaptive Asynchronous Updates
di: Xiong, Guojun, et al.
Pubblicazione: (2023)
di: Xiong, Guojun, et al.
Pubblicazione: (2023)
Bandwidth-Aware and Overlap-Weighted Compression for Communication-Efficient Federated Learning
di: Tang, Zichen, et al.
Pubblicazione: (2024)
di: Tang, Zichen, et al.
Pubblicazione: (2024)
A Joint Approach to Local Updating and Gradient Compression for Efficient Asynchronous Federated Learning
di: Song, Jiajun, et al.
Pubblicazione: (2024)
di: Song, Jiajun, et al.
Pubblicazione: (2024)
FedNMUT -- Federated Noisy Model Update Tracking Convergence Analysis
di: Chellapandi, Vishnu Pandi, et al.
Pubblicazione: (2024)
di: Chellapandi, Vishnu Pandi, et al.
Pubblicazione: (2024)
Age-of-Gradient Updates for Federated Learning over Random Access Channels
di: Wu, Yu Heng, et al.
Pubblicazione: (2024)
di: Wu, Yu Heng, et al.
Pubblicazione: (2024)
BGTplanner: Maximizing Training Accuracy for Differentially Private Federated Recommenders via Strategic Privacy Budget Allocation
di: Zhang, Xianzhi, et al.
Pubblicazione: (2024)
di: Zhang, Xianzhi, et al.
Pubblicazione: (2024)
FedAR: Addressing Client Unavailability in Federated Learning with Local Update Approximation and Rectification
di: Jiang, Chutian, et al.
Pubblicazione: (2024)
di: Jiang, Chutian, et al.
Pubblicazione: (2024)
Arctic Inference with Shift Parallelism: Fast and Efficient Open Source Inference System for Enterprise AI
di: Rajbhandari, Samyam, et al.
Pubblicazione: (2025)
di: Rajbhandari, Samyam, et al.
Pubblicazione: (2025)
NestPipe: Large-Scale Recommendation Training on 1,500+ Accelerators via Nested Pipelining
di: Jiang, Zhida, et al.
Pubblicazione: (2026)
di: Jiang, Zhida, et al.
Pubblicazione: (2026)
HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference
di: Zhang, Zeyu, et al.
Pubblicazione: (2025)
di: Zhang, Zeyu, et al.
Pubblicazione: (2025)
TurboGR: An Accelerated Training System for Large-Scale Generative Recommendation
di: Chai, Huichao, et al.
Pubblicazione: (2026)
di: Chai, Huichao, et al.
Pubblicazione: (2026)
Deal: Distributed End-to-End GNN Inference for All Nodes
di: Chen, Shiyang, et al.
Pubblicazione: (2025)
di: Chen, Shiyang, et al.
Pubblicazione: (2025)
RelayGR: Scaling Long-Sequence Generative Recommendation via Cross-Stage Relay-Race Inference
di: Wang, Jiarui, et al.
Pubblicazione: (2026)
di: Wang, Jiarui, et al.
Pubblicazione: (2026)
Designing Large Foundation Models for Efficient Training and Inference: A Survey
di: Liu, Dong, et al.
Pubblicazione: (2024)
di: Liu, Dong, et al.
Pubblicazione: (2024)
Two-dimensional Sparse Parallelism for Large Scale Deep Learning Recommendation Model Training
di: Zhang, Xin, et al.
Pubblicazione: (2025)
di: Zhang, Xin, et al.
Pubblicazione: (2025)
Large Language Model Aided QoS Prediction for Service Recommendation
di: Liu, Huiying, et al.
Pubblicazione: (2024)
di: Liu, Huiying, et al.
Pubblicazione: (2024)
Fair Distributed Cooperative Bandit Learning on Networks for Intelligent Internet of Things Systems (Technical Report)
di: Chen, Ziqun, et al.
Pubblicazione: (2024)
di: Chen, Ziqun, et al.
Pubblicazione: (2024)
Improving the End-to-End Efficiency of Offline Inference for Multi-LLM Applications Based on Sampling and Simulation
di: Fang, Jingzhi, et al.
Pubblicazione: (2025)
di: Fang, Jingzhi, et al.
Pubblicazione: (2025)
FUPareto: Bridging the Forgetting-Utility Gap in Federated Unlearning via Pareto Augmented Optimization
di: Wang, Zeyan, et al.
Pubblicazione: (2026)
di: Wang, Zeyan, et al.
Pubblicazione: (2026)
AMDP: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models Training
di: Chen, Ling, et al.
Pubblicazione: (2026)
di: Chen, Ling, et al.
Pubblicazione: (2026)
Documenti analoghi
-
One Pool, Two Caches: Adaptive HBM Partitioning for Accelerating Generative Recommender Serving
di: Yu, Wenjun, et al.
Pubblicazione: (2026) -
Faster Distributed Inference-Only Recommender Systems via Bounded Lag Synchronous Collectives
di: Dichev, Kiril, et al.
Pubblicazione: (2025) -
TaxBreak: Unmasking the Hidden Costs of LLM Inference Through Overhead Decomposition
di: Vellaisamy, Prabhu, et al.
Pubblicazione: (2026) -
RcLLM: Accelerating Generative Recommendation via Beyond-Prefix KV Caching
di: Zhao, Zhan, et al.
Pubblicazione: (2026) -
Lion Cub: Minimizing Communication Overhead in Distributed Lion
di: Ishikawa, Satoki, et al.
Pubblicazione: (2024)