Faster Distributed Inference-Only Recommender Systems via Bounded Lag Synchronous Collectives
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dichev, Kiril, Pawlowski, Filip, Yzelman, Albert-Jan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HiCR, an Abstract Model for Distributed Heterogeneous Programming
von: Martin, Sergio Miguel, et al.
Veröffentlicht: (2025)
von: Martin, Sergio Miguel, et al.
Veröffentlicht: (2025)
Effective implementation of the High Performance Conjugate Gradient benchmark on GraphBLAS
von: Scolari, Alberto, et al.
Veröffentlicht: (2023)
von: Scolari, Alberto, et al.
Veröffentlicht: (2023)
Near-Zero-Overhead Freshness for Recommendation Systems via Inference-Side Model Updates
von: Yu, Wenjun, et al.
Veröffentlicht: (2025)
von: Yu, Wenjun, et al.
Veröffentlicht: (2025)
Nonlinear spectral clustering with C++ GraphBLAS
von: Pasadakis, Dimosthenis, et al.
Veröffentlicht: (2026)
von: Pasadakis, Dimosthenis, et al.
Veröffentlicht: (2026)
Parallel Track Transformers: Enabling Fast GPU Inference with Reduced Synchronization
von: Wang, Chong, et al.
Veröffentlicht: (2026)
von: Wang, Chong, et al.
Veröffentlicht: (2026)
Empowering Distributed Training with Sparsity-driven Data Synchronization
von: Wang, Zhuang, et al.
Veröffentlicht: (2023)
von: Wang, Zhuang, et al.
Veröffentlicht: (2023)
QSync: Quantization-Minimized Synchronous Distributed Training Across Hybrid Devices
von: Zhao, Juntao, et al.
Veröffentlicht: (2024)
von: Zhao, Juntao, et al.
Veröffentlicht: (2024)
Collaborative and Distributed Bayesian Optimization via Consensus: Showcasing the Power of Collaboration for Optimal Design
von: Yue, Xubo, et al.
Veröffentlicht: (2023)
von: Yue, Xubo, et al.
Veröffentlicht: (2023)
Fast Distributed Inference Serving for Large Language Models
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
Priority-Aware Model-Distributed Inference at Edge Networks
von: Li, Teng, et al.
Veröffentlicht: (2024)
von: Li, Teng, et al.
Veröffentlicht: (2024)
When Less is More: Achieving Faster Convergence in Distributed Edge Machine Learning
von: Basani, Advik Raj, et al.
Veröffentlicht: (2024)
von: Basani, Advik Raj, et al.
Veröffentlicht: (2024)
Intelligent Orchestration of Distributed Large Foundation Model Inference at the Edge
von: Koch, Fernando, et al.
Veröffentlicht: (2025)
von: Koch, Fernando, et al.
Veröffentlicht: (2025)
Deal: Distributed End-to-End GNN Inference for All Nodes
von: Chen, Shiyang, et al.
Veröffentlicht: (2025)
von: Chen, Shiyang, et al.
Veröffentlicht: (2025)
TACOS: Topology-Aware Collective Algorithm Synthesizer for Distributed Machine Learning
von: Won, William, et al.
Veröffentlicht: (2023)
von: Won, William, et al.
Veröffentlicht: (2023)
FedFetch: Faster Federated Learning with Adaptive Downstream Prefetching
von: Yan, Qifan, et al.
Veröffentlicht: (2025)
von: Yan, Qifan, et al.
Veröffentlicht: (2025)
Shabari: Delayed Decision-Making for Faster and Efficient Serverless Functions
von: Sinha, Prasoon, et al.
Veröffentlicht: (2024)
von: Sinha, Prasoon, et al.
Veröffentlicht: (2024)
Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing
von: Liu, Mengfan, et al.
Veröffentlicht: (2025)
von: Liu, Mengfan, et al.
Veröffentlicht: (2025)
TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
von: Gond, Raja, et al.
Veröffentlicht: (2025)
von: Gond, Raja, et al.
Veröffentlicht: (2025)
Distributed Speculative Inference (DSI): Speculation Parallelism for Provably Faster Lossless Language Model Inference
von: Timor, Nadav, et al.
Veröffentlicht: (2024)
von: Timor, Nadav, et al.
Veröffentlicht: (2024)
Arctic Inference with Shift Parallelism: Fast and Efficient Open Source Inference System for Enterprise AI
von: Rajbhandari, Samyam, et al.
Veröffentlicht: (2025)
von: Rajbhandari, Samyam, et al.
Veröffentlicht: (2025)
Grappa: Gradient-Only Communication for Scalable Graph Neural Network Training
von: Xu, Chongyang, et al.
Veröffentlicht: (2026)
von: Xu, Chongyang, et al.
Veröffentlicht: (2026)
Converge Faster, Talk Less: Hessian-Informed Federated Zeroth-Order Optimization
von: Li, Zhe, et al.
Veröffentlicht: (2025)
von: Li, Zhe, et al.
Veröffentlicht: (2025)
TurboGR: An Accelerated Training System for Large-Scale Generative Recommendation
von: Chai, Huichao, et al.
Veröffentlicht: (2026)
von: Chai, Huichao, et al.
Veröffentlicht: (2026)
Hybrid Decentralized Optimization: Leveraging Both First- and Zeroth-Order Optimizers for Faster Convergence
von: Ansaripour, Matin, et al.
Veröffentlicht: (2022)
von: Ansaripour, Matin, et al.
Veröffentlicht: (2022)
Hybrid Dual-Batch and Cyclic Progressive Learning for Efficient Distributed Training
von: Lu, Kuan-Wei, et al.
Veröffentlicht: (2025)
von: Lu, Kuan-Wei, et al.
Veröffentlicht: (2025)
ChunkFlow: Communication-Aware Chunked Prefetching for Layerwise Offloading in Distributed Diffusion Transformer Inference
von: Meng, Han, et al.
Veröffentlicht: (2026)
von: Meng, Han, et al.
Veröffentlicht: (2026)
AutoRank: MCDA Based Rank Personalization for LoRA-Enabled Distributed Learning
von: Chen, Shuaijun, et al.
Veröffentlicht: (2024)
von: Chen, Shuaijun, et al.
Veröffentlicht: (2024)
KVDirect: Distributed Disaggregated LLM Inference
von: Chen, Shiyang, et al.
Veröffentlicht: (2024)
von: Chen, Shiyang, et al.
Veröffentlicht: (2024)
Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning
von: Qin, Ruoyu, et al.
Veröffentlicht: (2025)
von: Qin, Ruoyu, et al.
Veröffentlicht: (2025)
Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference
von: Yin, Ruokai, et al.
Veröffentlicht: (2025)
von: Yin, Ruokai, et al.
Veröffentlicht: (2025)
SyncFed: Time-Aware Federated Learning through Explicit Timestamping and Synchronization
von: Gül, Baran Can, et al.
Veröffentlicht: (2025)
von: Gül, Baran Can, et al.
Veröffentlicht: (2025)
NestPipe: Large-Scale Recommendation Training on 1,500+ Accelerators via Nested Pipelining
von: Jiang, Zhida, et al.
Veröffentlicht: (2026)
von: Jiang, Zhida, et al.
Veröffentlicht: (2026)
HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference
von: Tang, Peng, et al.
Veröffentlicht: (2024)
von: Tang, Peng, et al.
Veröffentlicht: (2024)
RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training
von: Gao, Wei, et al.
Veröffentlicht: (2025)
von: Gao, Wei, et al.
Veröffentlicht: (2025)
Collaborative Speculative Inference for Efficient LLM Inference Serving
von: Gao, Luyao, et al.
Veröffentlicht: (2025)
von: Gao, Luyao, et al.
Veröffentlicht: (2025)
BGTplanner: Maximizing Training Accuracy for Differentially Private Federated Recommenders via Strategic Privacy Budget Allocation
von: Zhang, Xianzhi, et al.
Veröffentlicht: (2024)
von: Zhang, Xianzhi, et al.
Veröffentlicht: (2024)
HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
Quasar: Quantized Self-Speculative Acceleration for Rapid Inference via Memory-Efficient Verification
von: Huang, Guang, et al.
Veröffentlicht: (2026)
von: Huang, Guang, et al.
Veröffentlicht: (2026)
MobiZO: Enabling Efficient LLM Fine-Tuning at the Edge via Inference Engines
von: Gao, Lei, et al.
Veröffentlicht: (2024)
von: Gao, Lei, et al.
Veröffentlicht: (2024)
Cornserve: A Distributed Serving System for Any-to-Any Multimodal Models
von: Chung, Jae-Won, et al.
Veröffentlicht: (2026)
von: Chung, Jae-Won, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
HiCR, an Abstract Model for Distributed Heterogeneous Programming
von: Martin, Sergio Miguel, et al.
Veröffentlicht: (2025) -
Effective implementation of the High Performance Conjugate Gradient benchmark on GraphBLAS
von: Scolari, Alberto, et al.
Veröffentlicht: (2023) -
Near-Zero-Overhead Freshness for Recommendation Systems via Inference-Side Model Updates
von: Yu, Wenjun, et al.
Veröffentlicht: (2025) -
Nonlinear spectral clustering with C++ GraphBLAS
von: Pasadakis, Dimosthenis, et al.
Veröffentlicht: (2026) -
Parallel Track Transformers: Enabling Fast GPU Inference with Reduced Synchronization
von: Wang, Chong, et al.
Veröffentlicht: (2026)