RL over Commodity Networks: Overcoming the Bandwidth Barrier with Lossless Sparse Deltas
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ruan, Chaoyi, Luo, Geng, Wan, Xinyi, Zhao, Long, Wang, Qinghe, Zhu, Jiaan, Xu, Duling, Xu, Guanbin, Wei, Dehui, Liu, Xiang, Li, Cheng, Sun, Haifeng, Miao, Congcong, Li, Jialin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cortex: Achieving Low-Latency, Cost-Efficient Remote Data Access For LLM via Semantic-Aware Knowledge Caching
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism
von: Zhao, Long, et al.
Veröffentlicht: (2026)
von: Zhao, Long, et al.
Veröffentlicht: (2026)
Revisiting Parameter Server in LLM Post-Training
von: Wan, Xinyi, et al.
Veröffentlicht: (2026)
von: Wan, Xinyi, et al.
Veröffentlicht: (2026)
Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training
von: Xu, Guanbin, et al.
Veröffentlicht: (2026)
von: Xu, Guanbin, et al.
Veröffentlicht: (2026)
Reaching Agreement Among Reasoning LLM Agents
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)
SparseRL-Sync: Lossless Weight Synchronization with ~100x Less Communication
von: Hu, Lucas, et al.
Veröffentlicht: (2026)
von: Hu, Lucas, et al.
Veröffentlicht: (2026)
Distributed Delta-Coloring under Bandwidth Limitations
von: Maus, Yannic, et al.
Veröffentlicht: (2024)
von: Maus, Yannic, et al.
Veröffentlicht: (2024)
Hiding Communication Cost in Distributed LLM Training via Micro-batch Co-execution
von: Wang, Haiquan, et al.
Veröffentlicht: (2024)
von: Wang, Haiquan, et al.
Veröffentlicht: (2024)
PipeOffload: Improving Scalability of Pipeline Parallelism with Memory Optimization
von: Wan, Xinyi, et al.
Veröffentlicht: (2025)
von: Wan, Xinyi, et al.
Veröffentlicht: (2025)
Bandwidth-Aware Network Topology Optimization for Decentralized Learning
von: Shen, Yipeng, et al.
Veröffentlicht: (2025)
von: Shen, Yipeng, et al.
Veröffentlicht: (2025)
Bandwidth-Aware and Cost-Efficient Pipeline Parallel Scheduling in Geo-Distributed LLM Training
von: Zhang, Han, et al.
Veröffentlicht: (2026)
von: Zhang, Han, et al.
Veröffentlicht: (2026)
Vault: Decentralized Storage Made Durable
von: Sun, Guangda, et al.
Veröffentlicht: (2023)
von: Sun, Guangda, et al.
Veröffentlicht: (2023)
Supermassive Blockchain
von: Sun, Guangda, et al.
Veröffentlicht: (2026)
von: Sun, Guangda, et al.
Veröffentlicht: (2026)
SparseMap: Loop Mapping for Sparse CNNs on Streaming Coarse-grained Reconfigurable Array
von: Ni, Xiaobing, et al.
Veröffentlicht: (2024)
von: Ni, Xiaobing, et al.
Veröffentlicht: (2024)
UCCL-Zip: Lossless Compression Supercharged GPU Communication
von: Ma, Shuang, et al.
Veröffentlicht: (2026)
von: Ma, Shuang, et al.
Veröffentlicht: (2026)
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
von: Sun, Mingyu, et al.
Veröffentlicht: (2025)
von: Sun, Mingyu, et al.
Veröffentlicht: (2025)
Computation-Bandwidth-Memory Trade-offs: A Unified Paradigm for AI Infrastructure
von: Fan, Yuankai, et al.
Veröffentlicht: (2025)
von: Fan, Yuankai, et al.
Veröffentlicht: (2025)
Oases: Efficient Large-Scale Model Training on Commodity Servers via Overlapped and Automated Tensor Model Parallelism
von: Li, Shengwei, et al.
Veröffentlicht: (2023)
von: Li, Shengwei, et al.
Veröffentlicht: (2023)
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers
von: Lu, Yao, et al.
Veröffentlicht: (2026)
von: Lu, Yao, et al.
Veröffentlicht: (2026)
TCDM Burst Access: Breaking the Bandwidth Barrier in Shared-L1 RVV Clusters Beyond 1000 FPUs
von: Shen, Diyou, et al.
Veröffentlicht: (2025)
von: Shen, Diyou, et al.
Veröffentlicht: (2025)
ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs
von: Yang, Jinwu, et al.
Veröffentlicht: (2026)
von: Yang, Jinwu, et al.
Veröffentlicht: (2026)
Building State Machine Replication Using Practical Network Synchrony
von: Wan, Yiliang, et al.
Veröffentlicht: (2025)
von: Wan, Yiliang, et al.
Veröffentlicht: (2025)
Floating-Point Data Transformation for Lossless Compression
von: Jamalidinan, Samirasadat, et al.
Veröffentlicht: (2025)
von: Jamalidinan, Samirasadat, et al.
Veröffentlicht: (2025)
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning
von: Xu, Lang, et al.
Veröffentlicht: (2025)
von: Xu, Lang, et al.
Veröffentlicht: (2025)
PipeMax: Enhancing Offline LLM Inference on Commodity GPU Servers
von: Zhang, Hongbin, et al.
Veröffentlicht: (2026)
von: Zhang, Hongbin, et al.
Veröffentlicht: (2026)
Performance Evaluation of Hashing Algorithms on Commodity Hardware
von: Pandya, Marut
Veröffentlicht: (2024)
von: Pandya, Marut
Veröffentlicht: (2024)
ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios
von: Hu, Xinyi, et al.
Veröffentlicht: (2026)
von: Hu, Xinyi, et al.
Veröffentlicht: (2026)
OrchestrRL: Dynamic Compute and Network Orchestration for Disaggregated RL
von: Tan, Xin, et al.
Veröffentlicht: (2026)
von: Tan, Xin, et al.
Veröffentlicht: (2026)
Accelerating Distributed Deep Learning using Lossless Homomorphic Compression
von: Li, Haoyu, et al.
Veröffentlicht: (2024)
von: Li, Haoyu, et al.
Veröffentlicht: (2024)
The Carnot Bound: Limits and Possibilities for Bandwidth-Efficient Consensus
von: Lewis-Pye, Andrew, et al.
Veröffentlicht: (2026)
von: Lewis-Pye, Andrew, et al.
Veröffentlicht: (2026)
Corrected with the Latest Version: Make Robust Asynchronous Federated Learning Possible
von: Lu, Chaoyi, et al.
Veröffentlicht: (2025)
von: Lu, Chaoyi, et al.
Veröffentlicht: (2025)
Strategies to Measure Energy Consumption Using RAPL During Workflow Execution on Commodity Clusters
von: Thamm, Philipp, et al.
Veröffentlicht: (2025)
von: Thamm, Philipp, et al.
Veröffentlicht: (2025)
Load Balancing Using Sparse Communication
von: Mendelson, Gal, et al.
Veröffentlicht: (2022)
von: Mendelson, Gal, et al.
Veröffentlicht: (2022)
Polar: Agentic RL on Any Harness at Scale
von: Xu, Binfeng, et al.
Veröffentlicht: (2026)
von: Xu, Binfeng, et al.
Veröffentlicht: (2026)
Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI
von: Khalilov, Mikhail, et al.
Veröffentlicht: (2024)
von: Khalilov, Mikhail, et al.
Veröffentlicht: (2024)
Low Latency, High Bandwidth Streaming of Experimental Data with EJFAT
von: Baldin, Ilya, et al.
Veröffentlicht: (2025)
von: Baldin, Ilya, et al.
Veröffentlicht: (2025)
Schedule-Level Shared-Prefix Reuse for LLM RL Training
von: Li, Pengbo, et al.
Veröffentlicht: (2026)
von: Li, Pengbo, et al.
Veröffentlicht: (2026)
On the Bandwidth Consumption of Blockchains
von: Lebedev, Andrei, et al.
Veröffentlicht: (2026)
von: Lebedev, Andrei, et al.
Veröffentlicht: (2026)
Federated Fine-Tuning of Sparsely-Activated Large Language Models on Resource-Constrained Devices
von: Chen, Fahao, et al.
Veröffentlicht: (2025)
von: Chen, Fahao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Cortex: Achieving Low-Latency, Cost-Efficient Remote Data Access For LLM via Semantic-Aware Knowledge Caching
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025) -
Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism
von: Zhao, Long, et al.
Veröffentlicht: (2026) -
Revisiting Parameter Server in LLM Post-Training
von: Wan, Xinyi, et al.
Veröffentlicht: (2026) -
Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training
von: Xu, Guanbin, et al.
Veröffentlicht: (2026) -
Reaching Agreement Among Reasoning LLM Agents
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2025)