SpecInF: Exploiting Idle GPU Resources in Distributed DL Training via Speculative Inference Filling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lv, Cunchi, Shi, Xiao, Liang, Dong, Tan, Wenting, Zhao, Xiaofang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dilu: Enabling GPU Resourcing-on-Demand for Serverless DL Serving via Introspective Elasticity
von: Lv, Cunchi, et al.
Veröffentlicht: (2025)
von: Lv, Cunchi, et al.
Veröffentlicht: (2025)
Cloudless-Training: A Framework to Improve Efficiency of Geo-Distributed ML Training
von: Tan, Wenting, et al.
Veröffentlicht: (2023)
von: Tan, Wenting, et al.
Veröffentlicht: (2023)
The Energy Cost of Execution-Idle in GPU Clusters
von: Lei, Yiran, et al.
Veröffentlicht: (2026)
von: Lei, Yiran, et al.
Veröffentlicht: (2026)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
von: Liu, Xing, et al.
Veröffentlicht: (2025)
von: Liu, Xing, et al.
Veröffentlicht: (2025)
SpecFed: Accelerating Federated LLM Inference with Speculative Decoding and Compressed Transmission
von: Zheng, Ce, et al.
Veröffentlicht: (2026)
von: Zheng, Ce, et al.
Veröffentlicht: (2026)
SpecEE: Accelerating Large Language Model Inference with Speculative Early Exiting
von: Xu, Jiaming, et al.
Veröffentlicht: (2025)
von: Xu, Jiaming, et al.
Veröffentlicht: (2025)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
von: Tran, Phuong, et al.
Veröffentlicht: (2025)
von: Tran, Phuong, et al.
Veröffentlicht: (2025)
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
von: Yu, Minchen, et al.
Veröffentlicht: (2023)
von: Yu, Minchen, et al.
Veröffentlicht: (2023)
FlexSpec: Frozen Drafts Meet Evolving Targets in Edge-Cloud Collaborative LLM Speculative Decoding
von: Li, Yuchen, et al.
Veröffentlicht: (2026)
von: Li, Yuchen, et al.
Veröffentlicht: (2026)
Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
On Harnessing Idle Compute at the Edge for Foundation Model Training
von: Xue, Leyang, et al.
Veröffentlicht: (2025)
von: Xue, Leyang, et al.
Veröffentlicht: (2025)
GPU Memory and Utilization Estimation for Training-Aware Resource Management: Opportunities and Limitations
von: Yousefzadeh-Asl-Miandoab, Ehsan, et al.
Veröffentlicht: (2026)
von: Yousefzadeh-Asl-Miandoab, Ehsan, et al.
Veröffentlicht: (2026)
HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU Clusters
von: Liang, Antian, et al.
Veröffentlicht: (2025)
von: Liang, Antian, et al.
Veröffentlicht: (2025)
Frenzy: A Memory-Aware Serverless LLM Training System for Heterogeneous GPU Clusters
von: Chang, Zihan, et al.
Veröffentlicht: (2024)
von: Chang, Zihan, et al.
Veröffentlicht: (2024)
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference
von: Xie, Jincheng, et al.
Veröffentlicht: (2026)
von: Xie, Jincheng, et al.
Veröffentlicht: (2026)
ReSpec: Towards Optimizing Speculative Decoding in Reinforcement Learning Systems
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
Minions: Accelerating Large Language Model Inference with Aggregated Speculative Execution
von: Wang, Siqi, et al.
Veröffentlicht: (2024)
von: Wang, Siqi, et al.
Veröffentlicht: (2024)
PowerTrip: Exploiting Federated Heterogeneous Datacenter Power for Distributed ML Training
von: Mehboob, Talha, et al.
Veröffentlicht: (2025)
von: Mehboob, Talha, et al.
Veröffentlicht: (2025)
SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
von: Miao, Xupeng, et al.
Veröffentlicht: (2023)
von: Miao, Xupeng, et al.
Veröffentlicht: (2023)
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
von: Han, Yunhe, et al.
Veröffentlicht: (2026)
von: Han, Yunhe, et al.
Veröffentlicht: (2026)
KIS-S: A GPU-Aware Kubernetes Inference Simulator with RL-Based Auto-Scaling
von: Zhang, Guilin, et al.
Veröffentlicht: (2025)
von: Zhang, Guilin, et al.
Veröffentlicht: (2025)
Distributed Speculative Execution for Resilient Cloud Applications
von: Li, Tianyu, et al.
Veröffentlicht: (2024)
von: Li, Tianyu, et al.
Veröffentlicht: (2024)
SparDL: Distributed Deep Learning Training with Efficient Sparse Communication
von: Zhao, Minjun, et al.
Veröffentlicht: (2023)
von: Zhao, Minjun, et al.
Veröffentlicht: (2023)
Accelerating OpenPangu Inference on NPU via Speculative Decoding
von: Dai, Yuntao, et al.
Veröffentlicht: (2026)
von: Dai, Yuntao, et al.
Veröffentlicht: (2026)
FEPLB: Exploiting Copy Engines for Nearly Free MoE Load Balancing in Distributed Training
von: Qi, Shuyao, et al.
Veröffentlicht: (2026)
von: Qi, Shuyao, et al.
Veröffentlicht: (2026)
SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
Accelerating Distributed MoE Training and Inference with Lina
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
SpecMemo: Speculative Decoding is in Your Pocket
von: Yildirim, Selin, et al.
Veröffentlicht: (2025)
von: Yildirim, Selin, et al.
Veröffentlicht: (2025)
Distributed Semi-Speculative Parallel Anisotropic Mesh Adaptation
von: Garner, Kevin, et al.
Veröffentlicht: (2026)
von: Garner, Kevin, et al.
Veröffentlicht: (2026)
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism
von: Shen, Yuhao, et al.
Veröffentlicht: (2025)
von: Shen, Yuhao, et al.
Veröffentlicht: (2025)
SpecRouter: Adaptive Routing for Multi-Level Speculative Decoding in Large Language Models
von: Wu, Hang, et al.
Veröffentlicht: (2025)
von: Wu, Hang, et al.
Veröffentlicht: (2025)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
SPIN: Accelerating Large Language Model Inference with Heterogeneous Speculative Models
von: Chen, Fahao, et al.
Veröffentlicht: (2025)
von: Chen, Fahao, et al.
Veröffentlicht: (2025)
ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud Environments
von: Lee, Munkyu, et al.
Veröffentlicht: (2024)
von: Lee, Munkyu, et al.
Veröffentlicht: (2024)
Optimal Resource Efficiency with Fairness in Heterogeneous GPU Clusters
von: Mo, Zizhao, et al.
Veröffentlicht: (2024)
von: Mo, Zizhao, et al.
Veröffentlicht: (2024)
Agora: Bridging the GPU Cloud Resource-Price Disconnect
von: McDougall, Ian, et al.
Veröffentlicht: (2025)
von: McDougall, Ian, et al.
Veröffentlicht: (2025)
Understanding GPU Resource Interference One Level Deeper
von: Elvinger, Paul, et al.
Veröffentlicht: (2025)
von: Elvinger, Paul, et al.
Veröffentlicht: (2025)
VDCores: Resource Decoupled Programming and Execution for Asynchronous GPU
von: He, Zijian, et al.
Veröffentlicht: (2026)
von: He, Zijian, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Dilu: Enabling GPU Resourcing-on-Demand for Serverless DL Serving via Introspective Elasticity
von: Lv, Cunchi, et al.
Veröffentlicht: (2025) -
Cloudless-Training: A Framework to Improve Efficiency of Geo-Distributed ML Training
von: Tan, Wenting, et al.
Veröffentlicht: (2023) -
The Energy Cost of Execution-Idle in GPU Clusters
von: Lei, Yiran, et al.
Veröffentlicht: (2026) -
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
von: Liu, Xing, et al.
Veröffentlicht: (2025) -
SpecFed: Accelerating Federated LLM Inference with Speculative Decoding and Compressed Transmission
von: Zheng, Ce, et al.
Veröffentlicht: (2026)