Learning to Keep a Promise: Scaling Language Model Decoding Parallelism with Learned Asynchronous Decoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jin, Tian, Cheng, Ellie Y., Ankner, Zack, Saunshi, Nikunj, Elias, Blake M., Yazdanbakhsh, Amir, Ragan-Kelley, Jonathan, Subramanian, Suvinay, Carbin, Michael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient Data-Parallel Continual Learning with Asynchronous Distributed Rehearsal Buffers
von: Bouvier, Thomas, et al.
Veröffentlicht: (2024)
von: Bouvier, Thomas, et al.
Veröffentlicht: (2024)
NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding
von: Chen, Jiefei, et al.
Veröffentlicht: (2026)
von: Chen, Jiefei, et al.
Veröffentlicht: (2026)
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
von: Wei, Jinhui, et al.
Veröffentlicht: (2025)
von: Wei, Jinhui, et al.
Veröffentlicht: (2025)
Benchmarking the Parallel 1D Heat Equation Solver in Chapel, Charm++, C++, HPX, Go, Julia, Python, Rust, Swift, and Java
von: Diehl, Patrick, et al.
Veröffentlicht: (2023)
von: Diehl, Patrick, et al.
Veröffentlicht: (2023)
COSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems
von: Raju, Aditi, et al.
Veröffentlicht: (2025)
von: Raju, Aditi, et al.
Veröffentlicht: (2025)
Boosting Blockchain Throughput: Parallel EVM Execution with Asynchronous Storage for Reddio
von: Qi, Xiaodong, et al.
Veröffentlicht: (2025)
von: Qi, Xiaodong, et al.
Veröffentlicht: (2025)
Asynchronous Secure Federated Learning with Byzantine aggregators
von: Del Pozzo, Antonella, et al.
Veröffentlicht: (2026)
von: Del Pozzo, Antonella, et al.
Veröffentlicht: (2026)
AMUSD: Asynchronous Multi-Device Speculative Decoding for LLM Acceleration
von: McDanel, Bradley
Veröffentlicht: (2024)
von: McDanel, Bradley
Veröffentlicht: (2024)
APEX: Asynchronous Parallel CPU-GPU Execution for Online LLM Inference on Constrained GPUs
von: Fan, Jiakun, et al.
Veröffentlicht: (2025)
von: Fan, Jiakun, et al.
Veröffentlicht: (2025)
Towards Adaptive Asynchronous Federated Learning for Human Activity Recognition
von: Gajanin, Rastko, et al.
Veröffentlicht: (2024)
von: Gajanin, Rastko, et al.
Veröffentlicht: (2024)
RAGO: Systematic Performance Optimization for Retrieval-Augmented Generation Serving
von: Jiang, Wenqi, et al.
Veröffentlicht: (2025)
von: Jiang, Wenqi, et al.
Veröffentlicht: (2025)
An Asynchronous Distributed-Memory Parallel Algorithm for k-mer Counting
von: Hati, Souvadra, et al.
Veröffentlicht: (2025)
von: Hati, Souvadra, et al.
Veröffentlicht: (2025)
Fault-Tolerant Decentralized Distributed Asynchronous Federated Learning with Adaptive Termination Detection
von: Akkinepally, Phani Sahasra, et al.
Veröffentlicht: (2025)
von: Akkinepally, Phani Sahasra, et al.
Veröffentlicht: (2025)
Asynchronous Federated Learning Using Outdated Local Updates Over TDMA Channel
von: Song, Jaeyoung, et al.
Veröffentlicht: (2024)
von: Song, Jaeyoung, et al.
Veröffentlicht: (2024)
Federated Semi-Supervised and Semi-Asynchronous Learning for Anomaly Detection in IoT Networks
von: Zhai, Wenbin, et al.
Veröffentlicht: (2023)
von: Zhai, Wenbin, et al.
Veröffentlicht: (2023)
Accelerating OpenPangu Inference on NPU via Speculative Decoding
von: Dai, Yuntao, et al.
Veröffentlicht: (2026)
von: Dai, Yuntao, et al.
Veröffentlicht: (2026)
EchoPFL: Asynchronous Personalized Federated Learning on Mobile Devices with On-Demand Staleness Control
von: Li, Xiaochen, et al.
Veröffentlicht: (2024)
von: Li, Xiaochen, et al.
Veröffentlicht: (2024)
FLASH Viterbi: Fast and Adaptive Viterbi Decoding for Modern Data Systems
von: Deng, Ziheng, et al.
Veröffentlicht: (2025)
von: Deng, Ziheng, et al.
Veröffentlicht: (2025)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving
von: Wu, Hanjiang, et al.
Veröffentlicht: (2026)
von: Wu, Hanjiang, et al.
Veröffentlicht: (2026)
Nesterov Method for Asynchronous Pipeline Parallel Optimization
von: Ajanthan, Thalaiyasingam, et al.
Veröffentlicht: (2025)
von: Ajanthan, Thalaiyasingam, et al.
Veröffentlicht: (2025)
SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism
von: Shen, Yuhao, et al.
Veröffentlicht: (2025)
von: Shen, Yuhao, et al.
Veröffentlicht: (2025)
Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding
von: Bhatia, Nidhi, et al.
Veröffentlicht: (2025)
von: Bhatia, Nidhi, et al.
Veröffentlicht: (2025)
FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
von: Chen, Wenyan, et al.
Veröffentlicht: (2026)
von: Chen, Wenyan, et al.
Veröffentlicht: (2026)
ClusterFusion++: Expanding Cluster-Level Fusion to Full Transformer-Block Decoding
von: Jin, ChiHeng, et al.
Veröffentlicht: (2026)
von: Jin, ChiHeng, et al.
Veröffentlicht: (2026)
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
von: Wang, Chao, et al.
Veröffentlicht: (2025)
von: Wang, Chao, et al.
Veröffentlicht: (2025)
Air-FedGA: A Grouping Asynchronous Federated Learning Mechanism Exploiting Over-the-air Computation
von: Ma, Qianpiao, et al.
Veröffentlicht: (2025)
von: Ma, Qianpiao, et al.
Veröffentlicht: (2025)
Resource-efficient Parallel Split Learning in Heterogeneous Edge Computing
von: Zhang, Mingjin, et al.
Veröffentlicht: (2024)
von: Zhang, Mingjin, et al.
Veröffentlicht: (2024)
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
von: Zhang, Yida, et al.
Veröffentlicht: (2026)
von: Zhang, Yida, et al.
Veröffentlicht: (2026)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
von: Zhong, Yinmin, et al.
Veröffentlicht: (2024)
von: Zhong, Yinmin, et al.
Veröffentlicht: (2024)
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
von: Tran, Phuong, et al.
Veröffentlicht: (2025)
von: Tran, Phuong, et al.
Veröffentlicht: (2025)
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
von: Han, Yunhe, et al.
Veröffentlicht: (2026)
von: Han, Yunhe, et al.
Veröffentlicht: (2026)
ReSpec: Towards Optimizing Speculative Decoding in Reinforcement Learning Systems
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
von: Chen, Xing, et al.
Veröffentlicht: (2025)
von: Chen, Xing, et al.
Veröffentlicht: (2025)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
von: Chen, Liangkun, et al.
Veröffentlicht: (2025)
von: Chen, Liangkun, et al.
Veröffentlicht: (2025)
Fast Matrix Multiplications for Lookup Table-Quantized LLMs
von: Guo, Han, et al.
Veröffentlicht: (2024)
von: Guo, Han, et al.
Veröffentlicht: (2024)
Understanding Communication Backends in Cross-Silo Federated Learning
von: Ziashahabi, Amir, et al.
Veröffentlicht: (2026)
von: Ziashahabi, Amir, et al.
Veröffentlicht: (2026)
Communication-Computation Pipeline Parallel Split Learning over Wireless Edge Networks
von: Liu, Chenyu, et al.
Veröffentlicht: (2025)
von: Liu, Chenyu, et al.
Veröffentlicht: (2025)
JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials
von: Wang, Hongyu, et al.
Veröffentlicht: (2026)
von: Wang, Hongyu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Efficient Data-Parallel Continual Learning with Asynchronous Distributed Rehearsal Buffers
von: Bouvier, Thomas, et al.
Veröffentlicht: (2024) -
NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding
von: Chen, Jiefei, et al.
Veröffentlicht: (2026) -
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025) -
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
von: Wei, Jinhui, et al.
Veröffentlicht: (2025) -
Benchmarking the Parallel 1D Heat Equation Solver in Chapel, Charm++, C++, HPX, Go, Julia, Python, Rust, Swift, and Java
von: Diehl, Patrick, et al.
Veröffentlicht: (2023)