SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Abramovich, Talor, Ashkenazi, Maor, Putterman, Izzy, Chislett, Benjamin, Mitra, Tiyasa, Rouhani, Bita Darvish, Zilberstein, Ran, Geifman, Yonatan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding
von: Bhatia, Nidhi, et al.
Veröffentlicht: (2025)
von: Bhatia, Nidhi, et al.
Veröffentlicht: (2025)
Beyond the Buzz: A Pragmatic Take on Inference Disaggregation
von: Mitra, Tiyasa, et al.
Veröffentlicht: (2025)
von: Mitra, Tiyasa, et al.
Veröffentlicht: (2025)
Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding
von: Iso, Hayate, et al.
Veröffentlicht: (2026)
von: Iso, Hayate, et al.
Veröffentlicht: (2026)
Accelerating OpenPangu Inference on NPU via Speculative Decoding
von: Dai, Yuntao, et al.
Veröffentlicht: (2026)
von: Dai, Yuntao, et al.
Veröffentlicht: (2026)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
von: Yu, Yanpeng, et al.
Veröffentlicht: (2025)
von: Yu, Yanpeng, et al.
Veröffentlicht: (2025)
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
von: Wei, Jinhui, et al.
Veröffentlicht: (2025)
von: Wei, Jinhui, et al.
Veröffentlicht: (2025)
FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
von: Chen, Wenyan, et al.
Veröffentlicht: (2026)
von: Chen, Wenyan, et al.
Veröffentlicht: (2026)
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
von: Zhang, Yida, et al.
Veröffentlicht: (2026)
von: Zhang, Yida, et al.
Veröffentlicht: (2026)
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
von: Tran, Phuong, et al.
Veröffentlicht: (2025)
von: Tran, Phuong, et al.
Veröffentlicht: (2025)
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
von: Han, Yunhe, et al.
Veröffentlicht: (2026)
von: Han, Yunhe, et al.
Veröffentlicht: (2026)
Privacy-Preserving and Incentive-Driven Relay-Based Framework for Cross-Domain Blockchain Interoperability
von: Moradi, Saeed, et al.
Veröffentlicht: (2025)
von: Moradi, Saeed, et al.
Veröffentlicht: (2025)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
von: Chen, Liangkun, et al.
Veröffentlicht: (2025)
von: Chen, Liangkun, et al.
Veröffentlicht: (2025)
Proof-of-Collaborative-Learning: A Multi-winner Federated Learning Consensus Algorithm
von: Sokhankhosh, Amirreza, et al.
Veröffentlicht: (2024)
von: Sokhankhosh, Amirreza, et al.
Veröffentlicht: (2024)
FlexSpec: Frozen Drafts Meet Evolving Targets in Edge-Cloud Collaborative LLM Speculative Decoding
von: Li, Yuchen, et al.
Veröffentlicht: (2026)
von: Li, Yuchen, et al.
Veröffentlicht: (2026)
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
ReSpec: Towards Optimizing Speculative Decoding in Reinforcement Learning Systems
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
Nightjar: Dynamic Adaptive Speculative Decoding for Large Language Models Serving
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
ElastiBench: Scalable Continuous Benchmarking on Cloud FaaS Platforms
von: Schirmer, Trever, et al.
Veröffentlicht: (2024)
von: Schirmer, Trever, et al.
Veröffentlicht: (2024)
SpecFed: Accelerating Federated LLM Inference with Speculative Decoding and Compressed Transmission
von: Zheng, Ce, et al.
Veröffentlicht: (2026)
von: Zheng, Ce, et al.
Veröffentlicht: (2026)
Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
von: Song, Jingwei, et al.
Veröffentlicht: (2025)
von: Song, Jingwei, et al.
Veröffentlicht: (2025)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
von: Liu, Xing, et al.
Veröffentlicht: (2025)
von: Liu, Xing, et al.
Veröffentlicht: (2025)
Enhancing Split Learning with Sharded and Blockchain-Enabled SplitFed Approaches
von: Sokhankhosh, Amirreza, et al.
Veröffentlicht: (2025)
von: Sokhankhosh, Amirreza, et al.
Veröffentlicht: (2025)
Utility-Driven Speculative Decoding for Mixture-of-Experts
von: Saxena, Anish, et al.
Veröffentlicht: (2025)
von: Saxena, Anish, et al.
Veröffentlicht: (2025)
SpecMemo: Speculative Decoding is in Your Pocket
von: Yildirim, Selin, et al.
Veröffentlicht: (2025)
von: Yildirim, Selin, et al.
Veröffentlicht: (2025)
Ocularone-Bench: Benchmarking DNN Models on GPUs to Assist the Visually Impaired
von: Raj, Suman, et al.
Veröffentlicht: (2025)
von: Raj, Suman, et al.
Veröffentlicht: (2025)
SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism
von: Shen, Yuhao, et al.
Veröffentlicht: (2025)
von: Shen, Yuhao, et al.
Veröffentlicht: (2025)
SpecRouter: Adaptive Routing for Multi-Level Speculative Decoding in Large Language Models
von: Wu, Hang, et al.
Veröffentlicht: (2025)
von: Wu, Hang, et al.
Veröffentlicht: (2025)
Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving
von: Wang, Chao, et al.
Veröffentlicht: (2025)
von: Wang, Chao, et al.
Veröffentlicht: (2025)
MemGUI-Bench: Benchmarking Memory of Mobile GUI Agents in Dynamic Environments
von: Liu, Guangyi, et al.
Veröffentlicht: (2026)
von: Liu, Guangyi, et al.
Veröffentlicht: (2026)
Fast and Cost-effective Speculative Edge-Cloud Decoding with Early Exits
von: Venkatesha, Yeshwanth, et al.
Veröffentlicht: (2025)
von: Venkatesha, Yeshwanth, et al.
Veröffentlicht: (2025)
Striking the Right Balance between Compute and Copy: Improving LLM Inferencing Under Speculative Decoding
von: Ramachandran, Arun, et al.
Veröffentlicht: (2025)
von: Ramachandran, Arun, et al.
Veröffentlicht: (2025)
DSD: A Distributed Speculative Decoding Solution for Edge-Cloud Agile Large Model Serving
von: Yu, Fengze, et al.
Veröffentlicht: (2025)
von: Yu, Fengze, et al.
Veröffentlicht: (2025)
TailBench++: Flexible Multi-Client, Multi-Server Benchmarking for Latency-Critical Workloads
von: Li, Zhilin, et al.
Veröffentlicht: (2025)
von: Li, Zhilin, et al.
Veröffentlicht: (2025)
Distributed Speculative Execution for Resilient Cloud Applications
von: Li, Tianyu, et al.
Veröffentlicht: (2024)
von: Li, Tianyu, et al.
Veröffentlicht: (2024)
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
von: Cheng, Long, et al.
Veröffentlicht: (2026)
von: Cheng, Long, et al.
Veröffentlicht: (2026)
Enabling Blockchain Interoperability Through Network Discovery Services
von: Hassan, Khalid, et al.
Veröffentlicht: (2025)
von: Hassan, Khalid, et al.
Veröffentlicht: (2025)
MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts
von: Wang, Wenfeng, et al.
Veröffentlicht: (2025)
von: Wang, Wenfeng, et al.
Veröffentlicht: (2025)
Tasking framework for Adaptive Speculative Parallel Mesh Generation
von: Tsolakis, Christos, et al.
Veröffentlicht: (2024)
von: Tsolakis, Christos, et al.
Veröffentlicht: (2024)
Distributed Semi-Speculative Parallel Anisotropic Mesh Adaptation
von: Garner, Kevin, et al.
Veröffentlicht: (2026)
von: Garner, Kevin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding
von: Bhatia, Nidhi, et al.
Veröffentlicht: (2025) -
Beyond the Buzz: A Pragmatic Take on Inference Disaggregation
von: Mitra, Tiyasa, et al.
Veröffentlicht: (2025) -
Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding
von: Iso, Hayate, et al.
Veröffentlicht: (2026) -
Accelerating OpenPangu Inference on NPU via Speculative Decoding
von: Dai, Yuntao, et al.
Veröffentlicht: (2026) -
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)