ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Xinyi, Shen, Yuhao, Zhang, Baolin, Zhang, Hengxin, Dai, Jun, Ge, Shuang, Chen, Lei, Li, Yue, Wan, Mingcheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Accelerating OpenPangu Inference on NPU via Speculative Decoding
von: Dai, Yuntao, et al.
Veröffentlicht: (2026)
von: Dai, Yuntao, et al.
Veröffentlicht: (2026)
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
von: Zhang, Yida, et al.
Veröffentlicht: (2026)
von: Zhang, Yida, et al.
Veröffentlicht: (2026)
SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism
von: Shen, Yuhao, et al.
Veröffentlicht: (2025)
von: Shen, Yuhao, et al.
Veröffentlicht: (2025)
HiCoCS: High Concurrency Cross-Sharding on Permissioned Blockchains
von: Yang, Lingxiao, et al.
Veröffentlicht: (2025)
von: Yang, Lingxiao, et al.
Veröffentlicht: (2025)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
von: Han, Yunhe, et al.
Veröffentlicht: (2026)
von: Han, Yunhe, et al.
Veröffentlicht: (2026)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
von: Chen, Liangkun, et al.
Veröffentlicht: (2025)
von: Chen, Liangkun, et al.
Veröffentlicht: (2025)
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
von: Wei, Jinhui, et al.
Veröffentlicht: (2025)
von: Wei, Jinhui, et al.
Veröffentlicht: (2025)
FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
von: Chen, Wenyan, et al.
Veröffentlicht: (2026)
von: Chen, Wenyan, et al.
Veröffentlicht: (2026)
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
von: Tran, Phuong, et al.
Veröffentlicht: (2025)
von: Tran, Phuong, et al.
Veröffentlicht: (2025)
Nightjar: Dynamic Adaptive Speculative Decoding for Large Language Models Serving
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
FlexSpec: Frozen Drafts Meet Evolving Targets in Edge-Cloud Collaborative LLM Speculative Decoding
von: Li, Yuchen, et al.
Veröffentlicht: (2026)
von: Li, Yuchen, et al.
Veröffentlicht: (2026)
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism
von: Qing, Yuhao, et al.
Veröffentlicht: (2025)
von: Qing, Yuhao, et al.
Veröffentlicht: (2025)
PhoenixOS: Concurrent OS-level GPU Checkpoint and Restore with Validated Speculation
von: Wei, Xingda, et al.
Veröffentlicht: (2024)
von: Wei, Xingda, et al.
Veröffentlicht: (2024)
ReSpec: Towards Optimizing Speculative Decoding in Reinforcement Learning Systems
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
DSD: A Distributed Speculative Decoding Solution for Edge-Cloud Agile Large Model Serving
von: Yu, Fengze, et al.
Veröffentlicht: (2025)
von: Yu, Fengze, et al.
Veröffentlicht: (2025)
SparseMap: Loop Mapping for Sparse CNNs on Streaming Coarse-grained Reconfigurable Array
von: Ni, Xiaobing, et al.
Veröffentlicht: (2024)
von: Ni, Xiaobing, et al.
Veröffentlicht: (2024)
StableShard: Stable and Scalable Blockchain Sharding with High Concurrency via Collaborative Committees
von: Li, Mingzhe, et al.
Veröffentlicht: (2024)
von: Li, Mingzhe, et al.
Veröffentlicht: (2024)
Striking the Right Balance between Compute and Copy: Improving LLM Inferencing Under Speculative Decoding
von: Ramachandran, Arun, et al.
Veröffentlicht: (2025)
von: Ramachandran, Arun, et al.
Veröffentlicht: (2025)
SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding
von: Abramovich, Talor, et al.
Veröffentlicht: (2026)
von: Abramovich, Talor, et al.
Veröffentlicht: (2026)
History-Independent Concurrent Objects
von: Attiya, Hagit, et al.
Veröffentlicht: (2024)
von: Attiya, Hagit, et al.
Veröffentlicht: (2024)
ElasticMoE: An Efficient Auto Scaling Method for Mixture-of-Experts Models
von: Singh, Gursimran, et al.
Veröffentlicht: (2025)
von: Singh, Gursimran, et al.
Veröffentlicht: (2025)
AdaOper: Energy-efficient and Responsive Concurrent DNN Inference on Mobile Devices
von: Lin, Zheng, et al.
Veröffentlicht: (2024)
von: Lin, Zheng, et al.
Veröffentlicht: (2024)
Proving Highly-Concurrent Traversals Correct
von: Feldman, Yotam M. Y., et al.
Veröffentlicht: (2020)
von: Feldman, Yotam M. Y., et al.
Veröffentlicht: (2020)
AsyncSparse: Accelerating Sparse Matrix-Matrix Multiplication on Asynchronous GPU Architectures
von: Liu, Jie, et al.
Veröffentlicht: (2026)
von: Liu, Jie, et al.
Veröffentlicht: (2026)
RL over Commodity Networks: Overcoming the Bandwidth Barrier with Lossless Sparse Deltas
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2026)
von: Ruan, Chaoyi, et al.
Veröffentlicht: (2026)
Epoch-based Optimistic Concurrency Control in Geo-replicated Databases
von: Mao, Yunhao, et al.
Veröffentlicht: (2026)
von: Mao, Yunhao, et al.
Veröffentlicht: (2026)
MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts
von: Wang, Wenfeng, et al.
Veröffentlicht: (2025)
von: Wang, Wenfeng, et al.
Veröffentlicht: (2025)
Distributed Speculative Execution for Resilient Cloud Applications
von: Li, Tianyu, et al.
Veröffentlicht: (2024)
von: Li, Tianyu, et al.
Veröffentlicht: (2024)
PromptTuner: SLO-Aware Elastic System for LLM Prompt Tuning
von: Gao, Wei, et al.
Veröffentlicht: (2026)
von: Gao, Wei, et al.
Veröffentlicht: (2026)
OTAS: An Elastic Transformer Serving System via Token Adaptation
von: Chen, Jinyu, et al.
Veröffentlicht: (2024)
von: Chen, Jinyu, et al.
Veröffentlicht: (2024)
Resolving Conflicts with Grace: Dynamically Concurrent Universality
von: Kuznetsov, Petr, et al.
Veröffentlicht: (2025)
von: Kuznetsov, Petr, et al.
Veröffentlicht: (2025)
A Study of Synchronization Methods for Concurrent Size
von: Kas-Sharir, Hen, et al.
Veröffentlicht: (2025)
von: Kas-Sharir, Hen, et al.
Veröffentlicht: (2025)
EcoLife: Carbon-Aware Serverless Function Scheduling for Sustainable Computing
von: Jiang, Yankai, et al.
Veröffentlicht: (2024)
von: Jiang, Yankai, et al.
Veröffentlicht: (2024)
SpecFed: Accelerating Federated LLM Inference with Speculative Decoding and Compressed Transmission
von: Zheng, Ce, et al.
Veröffentlicht: (2026)
von: Zheng, Ce, et al.
Veröffentlicht: (2026)
Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
von: Song, Jingwei, et al.
Veröffentlicht: (2025)
von: Song, Jingwei, et al.
Veröffentlicht: (2025)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
von: Liu, Xing, et al.
Veröffentlicht: (2025)
von: Liu, Xing, et al.
Veröffentlicht: (2025)
Tasking framework for Adaptive Speculative Parallel Mesh Generation
von: Tsolakis, Christos, et al.
Veröffentlicht: (2024)
von: Tsolakis, Christos, et al.
Veröffentlicht: (2024)
Distributed Semi-Speculative Parallel Anisotropic Mesh Adaptation
von: Garner, Kevin, et al.
Veröffentlicht: (2026)
von: Garner, Kevin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Accelerating OpenPangu Inference on NPU via Speculative Decoding
von: Dai, Yuntao, et al.
Veröffentlicht: (2026) -
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
von: Zhang, Yida, et al.
Veröffentlicht: (2026) -
SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism
von: Shen, Yuhao, et al.
Veröffentlicht: (2025) -
HiCoCS: High Concurrency Cross-Sharding on Permissioned Blockchains
von: Yang, Lingxiao, et al.
Veröffentlicht: (2025) -
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)