RLHFless: Serverless Computing for Efficient RLHF
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wei, Rui, Yu, Hanfei, Jain, Shubham, Sivakumar, Yogarajan, Tiwari, Devesh, Li, Jian, Park, Seung-Jong, Wang, Hao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MoEless: Efficient MoE LLM Serving via Serverless Computing
von: Yu, Hanfei, et al.
Veröffentlicht: (2026)
von: Yu, Hanfei, et al.
Veröffentlicht: (2026)
HotSwap: Enabling Live Dependency Sharing in Serverless Computing
von: Li, Rui, et al.
Veröffentlicht: (2024)
von: Li, Rui, et al.
Veröffentlicht: (2024)
EcoLife: Carbon-Aware Serverless Function Scheduling for Sustainable Computing
von: Jiang, Yankai, et al.
Veröffentlicht: (2024)
von: Jiang, Yankai, et al.
Veröffentlicht: (2024)
Remoe: Towards Efficient and Low-Cost MoE Inference in Serverless Computing
von: Liu, Wentao, et al.
Veröffentlicht: (2025)
von: Liu, Wentao, et al.
Veröffentlicht: (2025)
LLM Inference Serving: Survey of Recent Advances and Opportunities
von: Li, Baolin, et al.
Veröffentlicht: (2024)
von: Li, Baolin, et al.
Veröffentlicht: (2024)
ServerlessLoRA: Minimizing Latency and Cost in Serverless Inference for LoRA-Based LLMs
von: Sui, Yifan, et al.
Veröffentlicht: (2025)
von: Sui, Yifan, et al.
Veröffentlicht: (2025)
Taming Latency-Memory Trade-Off in MoE-Based LLM Serving via Fine-Grained Expert Offloading
von: Yu, Hanfei, et al.
Veröffentlicht: (2025)
von: Yu, Hanfei, et al.
Veröffentlicht: (2025)
Word Frequency Counting Based on Serverless MapReduce
von: Li, Hanzhe, et al.
Veröffentlicht: (2026)
von: Li, Hanzhe, et al.
Veröffentlicht: (2026)
Green or Fast? Learning to Balance Cold Starts and Idle Carbon in Serverless Computing
von: Sun, Bowen, et al.
Veröffentlicht: (2026)
von: Sun, Bowen, et al.
Veröffentlicht: (2026)
ENOVA: Autoscaling towards Cost-effective and Stable Serverless LLM Serving
von: Huang, Tao, et al.
Veröffentlicht: (2024)
von: Huang, Tao, et al.
Veröffentlicht: (2024)
Shard the Gradient, Scale the Model: Serverless Federated Aggregation via Gradient Partitioning
von: Barrak, Amine
Veröffentlicht: (2026)
von: Barrak, Amine
Veröffentlicht: (2026)
Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism
von: Zhao, Long, et al.
Veröffentlicht: (2026)
von: Zhao, Long, et al.
Veröffentlicht: (2026)
FaaSTube: Optimizing GPU-oriented Data Transfer for Serverless Computing
von: Wu, Hao, et al.
Veröffentlicht: (2024)
von: Wu, Hao, et al.
Veröffentlicht: (2024)
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
von: Yu, Minchen, et al.
Veröffentlicht: (2023)
von: Yu, Minchen, et al.
Veröffentlicht: (2023)
DeF-DReL: Systematic Deployment of Serverless Functions in Fog and Cloud environments using Deep Reinforcement Learning
von: Dehury, Chinmaya Kumar, et al.
Veröffentlicht: (2021)
von: Dehury, Chinmaya Kumar, et al.
Veröffentlicht: (2021)
On-demand Cold Start Frequency Reduction with Off-Policy Reinforcement Learning in Serverless Computing
von: Agarwal, Siddharth, et al.
Veröffentlicht: (2023)
von: Agarwal, Siddharth, et al.
Veröffentlicht: (2023)
Input-Based Ensemble-Learning Method for Dynamic Memory Configuration of Serverless Computing Functions
von: Agarwal, Siddharth, et al.
Veröffentlicht: (2024)
von: Agarwal, Siddharth, et al.
Veröffentlicht: (2024)
Making Serverless Computing Extensible: A Case Study of Serverless Data Analytics
von: Yu, Minchen, et al.
Veröffentlicht: (2025)
von: Yu, Minchen, et al.
Veröffentlicht: (2025)
Towards Energy-Efficient Serverless Computing with Hardware Isolation
von: Carl, Natalie, et al.
Veröffentlicht: (2025)
von: Carl, Natalie, et al.
Veröffentlicht: (2025)
WaterWise: Co-optimizing Carbon- and Water-Footprint Toward Environmentally Sustainable Cloud Computing
von: Jiang, Yankai, et al.
Veröffentlicht: (2025)
von: Jiang, Yankai, et al.
Veröffentlicht: (2025)
Lightweight, Secure and Stateful Serverless Computing with PSL
von: Thomas, Alexander, et al.
Veröffentlicht: (2024)
von: Thomas, Alexander, et al.
Veröffentlicht: (2024)
Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
von: Hao, Zixu, et al.
Veröffentlicht: (2025)
von: Hao, Zixu, et al.
Veröffentlicht: (2025)
Toward Sustainable GenAI using Generation Directives for Carbon-Friendly Large Language Model Inference
von: Li, Baolin, et al.
Veröffentlicht: (2024)
von: Li, Baolin, et al.
Veröffentlicht: (2024)
ReaL: Efficient RLHF Training of Large Language Models with Parameter Reallocation
von: Mei, Zhiyu, et al.
Veröffentlicht: (2024)
von: Mei, Zhiyu, et al.
Veröffentlicht: (2024)
Sustainable Supercomputing for AI: GPU Power Capping at HPC Scale
von: Zhao, Dan, et al.
Veröffentlicht: (2024)
von: Zhao, Dan, et al.
Veröffentlicht: (2024)
MoESys: A Distributed and Efficient Mixture-of-Experts Training and Inference System for Internet Services
von: Yu, Dianhai, et al.
Veröffentlicht: (2022)
von: Yu, Dianhai, et al.
Veröffentlicht: (2022)
Multi-Event Triggers for Serverless Computing
von: Carl, Natalie, et al.
Veröffentlicht: (2025)
von: Carl, Natalie, et al.
Veröffentlicht: (2025)
Litmus: Fair Pricing for Serverless Computing
von: Pei, Qi, et al.
Veröffentlicht: (2024)
von: Pei, Qi, et al.
Veröffentlicht: (2024)
Serverless Computing: Architecture, Concepts, and Applications
von: Ghorbian, Mohsen, et al.
Veröffentlicht: (2025)
von: Ghorbian, Mohsen, et al.
Veröffentlicht: (2025)
FSD-Inference: Fully Serverless Distributed Inference with Scalable Cloud Communication
von: Oakley, Joe, et al.
Veröffentlicht: (2024)
von: Oakley, Joe, et al.
Veröffentlicht: (2024)
Zenix: Efficient Execution of Bulky Serverless Applications
von: Guo, Zhiyuan, et al.
Veröffentlicht: (2022)
von: Guo, Zhiyuan, et al.
Veröffentlicht: (2022)
An Upload-Efficient Scheme for Transferring Knowledge From a Server-Side Pre-trained Generator to Clients in Heterogeneous Federated Learning
von: Zhang, Jianqing, et al.
Veröffentlicht: (2024)
von: Zhang, Jianqing, et al.
Veröffentlicht: (2024)
Towards an Introspective Dynamic Model of Globally Distributed Computing Infrastructures
von: Kilic, Ozgur O., et al.
Veröffentlicht: (2025)
von: Kilic, Ozgur O., et al.
Veröffentlicht: (2025)
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models
von: Cheng, Jialiang, et al.
Veröffentlicht: (2024)
von: Cheng, Jialiang, et al.
Veröffentlicht: (2024)
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
von: Liu, Dong, et al.
Veröffentlicht: (2025)
von: Liu, Dong, et al.
Veröffentlicht: (2025)
Tangram: Accelerating Serverless LLM Loading through GPU Memory Reuse and Affinity
von: Zhu, Wenbin, et al.
Veröffentlicht: (2025)
von: Zhu, Wenbin, et al.
Veröffentlicht: (2025)
Training Heterogeneous Client Models using Knowledge Distillation in Serverless Federated Learning
von: Chadha, Mohak, et al.
Veröffentlicht: (2024)
von: Chadha, Mohak, et al.
Veröffentlicht: (2024)
Energy Efficient Scheduling for Serverless Systems
von: Tsenos, Michail, et al.
Veröffentlicht: (2024)
von: Tsenos, Michail, et al.
Veröffentlicht: (2024)
FaasMeter: Energy-First Serverless Computing
von: Rehman, Abdul, et al.
Veröffentlicht: (2024)
von: Rehman, Abdul, et al.
Veröffentlicht: (2024)
Caching Aided Multi-Tenant Serverless Computing
von: Qiao, Chu, et al.
Veröffentlicht: (2024)
von: Qiao, Chu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MoEless: Efficient MoE LLM Serving via Serverless Computing
von: Yu, Hanfei, et al.
Veröffentlicht: (2026) -
HotSwap: Enabling Live Dependency Sharing in Serverless Computing
von: Li, Rui, et al.
Veröffentlicht: (2024) -
EcoLife: Carbon-Aware Serverless Function Scheduling for Sustainable Computing
von: Jiang, Yankai, et al.
Veröffentlicht: (2024) -
Remoe: Towards Efficient and Low-Cost MoE Inference in Serverless Computing
von: Liu, Wentao, et al.
Veröffentlicht: (2025) -
LLM Inference Serving: Survey of Recent Advances and Opportunities
von: Li, Baolin, et al.
Veröffentlicht: (2024)