SMART: When is it Actually Worth Expanding a Speculative Tree?
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Lifu, Zhou, Pan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SpecEE: Accelerating Large Language Model Inference with Speculative Early Exiting
by: Xu, Jiaming, et al.
Published: (2025)
by: Xu, Jiaming, et al.
Published: (2025)
Fast LLM Post-training via Decoupled and Fastest-of-N Speculation
by: Cheng, Rongxin, et al.
Published: (2025)
by: Cheng, Rongxin, et al.
Published: (2025)
Nightjar: Dynamic Adaptive Speculative Decoding for Large Language Models Serving
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
Design a Win-Win Strategy That Is Fair to Both Service Providers and Tasks When Rejection Is Not an Option
by: Trabelsi, Yohai, et al.
Published: (2024)
by: Trabelsi, Yohai, et al.
Published: (2024)
SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism
by: Shen, Yuhao, et al.
Published: (2025)
by: Shen, Yuhao, et al.
Published: (2025)
SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding
by: Abramovich, Talor, et al.
Published: (2026)
by: Abramovich, Talor, et al.
Published: (2026)
Act While Thinking: Accelerating LLM Agents via Pattern-Aware Speculative Tool Execution
by: Sui, Yifan, et al.
Published: (2026)
by: Sui, Yifan, et al.
Published: (2026)
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference
by: Xie, Jincheng, et al.
Published: (2026)
by: Xie, Jincheng, et al.
Published: (2026)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
by: Kumar, Satyam, et al.
Published: (2026)
by: Kumar, Satyam, et al.
Published: (2026)
Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
by: Song, Jingwei, et al.
Published: (2025)
by: Song, Jingwei, et al.
Published: (2025)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
by: Liu, Xing, et al.
Published: (2025)
by: Liu, Xing, et al.
Published: (2025)
B-PASTE: Beam-Aware Pattern-Guided Speculative Execution for Resource-Constrained LLM Agents
by: Song, Yanfei
Published: (2026)
by: Song, Yanfei
Published: (2026)
Striking the Right Balance between Compute and Copy: Improving LLM Inferencing Under Speculative Decoding
by: Ramachandran, Arun, et al.
Published: (2025)
by: Ramachandran, Arun, et al.
Published: (2025)
ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive
by: Luo, Xinhao, et al.
Published: (2025)
by: Luo, Xinhao, et al.
Published: (2025)
Fast and Cost-effective Speculative Edge-Cloud Decoding with Early Exits
by: Venkatesha, Yeshwanth, et al.
Published: (2025)
by: Venkatesha, Yeshwanth, et al.
Published: (2025)
Fail Fast, Win Big: Rethinking the Drafting Strategy in Speculative Decoding via Diffusion LLMs
by: Pan, Rui, et al.
Published: (2025)
by: Pan, Rui, et al.
Published: (2025)
Utility-Driven Speculative Decoding for Mixture-of-Experts
by: Saxena, Anish, et al.
Published: (2025)
by: Saxena, Anish, et al.
Published: (2025)
SpecMemo: Speculative Decoding is in Your Pocket
by: Yildirim, Selin, et al.
Published: (2025)
by: Yildirim, Selin, et al.
Published: (2025)
Why Smaller Is Slower? Dimensional Misalignment in Compressed LLMs
by: Xin, Jihao, et al.
Published: (2026)
by: Xin, Jihao, et al.
Published: (2026)
xDiT: an Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism
by: Fang, Jiarui, et al.
Published: (2024)
by: Fang, Jiarui, et al.
Published: (2024)
When Speculation Spills Secrets: Side Channels via Speculative Decoding In LLMs
by: Wei, Jiankun, et al.
Published: (2024)
by: Wei, Jiankun, et al.
Published: (2024)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
by: Pan, Xinglin, et al.
Published: (2025)
by: Pan, Xinglin, et al.
Published: (2025)
LLM-42: Enabling Determinism in LLM Inference with Verified Speculation
by: Gond, Raja, et al.
Published: (2026)
by: Gond, Raja, et al.
Published: (2026)
Speculative Actions: A Lossless Framework for Faster Agentic Systems
by: Ye, Naimeng, et al.
Published: (2025)
by: Ye, Naimeng, et al.
Published: (2025)
The intelligent prediction and assessment of financial information risk in the cloud computing model
by: Wang, Yufu, et al.
Published: (2024)
by: Wang, Yufu, et al.
Published: (2024)
ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios
by: Hu, Xinyi, et al.
Published: (2026)
by: Hu, Xinyi, et al.
Published: (2026)
ScaleSim: Serving Large-Scale Multi-Agent Simulation with Invocation Distance-Based Memory Management
by: Pan, Zaifeng, et al.
Published: (2026)
by: Pan, Zaifeng, et al.
Published: (2026)
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
by: Wang, Haodong, et al.
Published: (2025)
by: Wang, Haodong, et al.
Published: (2025)
Position: LLM Serving Needs Mathematical Optimization and Algorithmic Foundations, Not Just Heuristics
by: Zhou, Zijie
Published: (2026)
by: Zhou, Zijie
Published: (2026)
Data-Juicer 2.0: Cloud-Scale Adaptive Data Processing for and with Foundation Models
by: Chen, Daoyuan, et al.
Published: (2024)
by: Chen, Daoyuan, et al.
Published: (2024)
Remoe: Towards Efficient and Low-Cost MoE Inference in Serverless Computing
by: Liu, Wentao, et al.
Published: (2025)
by: Liu, Wentao, et al.
Published: (2025)
FreeRide: Harvesting Bubbles in Pipeline Parallelism
by: Zhang, Jiashu, et al.
Published: (2024)
by: Zhang, Jiashu, et al.
Published: (2024)
ClusterRCA: An End-to-End Approach for Network Fault Localization and Classification for HPC System
by: Sun, Yongqian, et al.
Published: (2025)
by: Sun, Yongqian, et al.
Published: (2025)
Research on Model Parallelism and Data Parallelism Optimization Methods in Large Language Model-Based Recommendation Systems
by: Yang, Haowei, et al.
Published: (2025)
by: Yang, Haowei, et al.
Published: (2025)
PlanetServe: A Decentralized, Scalable, and Privacy-Preserving Overlay for Democratizing Large Language Model Serving
by: Fang, Fei, et al.
Published: (2025)
by: Fang, Fei, et al.
Published: (2025)
DIAP: A Decentralized Agent Identity Protocol with Zero-Knowledge Proofs and a Hybrid P2P Stack
by: Liu, Yuanjie, et al.
Published: (2025)
by: Liu, Yuanjie, et al.
Published: (2025)
CoRaiS: Lightweight Real-Time Scheduler for Multi-Edge Cooperative Computing
by: Hu, Yujiao, et al.
Published: (2024)
by: Hu, Yujiao, et al.
Published: (2024)
DGRAG: Distributed Graph-based Retrieval-Augmented Generation in Edge-Cloud Systems
by: Zhou, Wenqing, et al.
Published: (2025)
by: Zhou, Wenqing, et al.
Published: (2025)
Distributed Speculative Inference (DSI): Speculation Parallelism for Provably Faster Lossless Language Model Inference
by: Timor, Nadav, et al.
Published: (2024)
by: Timor, Nadav, et al.
Published: (2024)
UCCL-Zip: Lossless Compression Supercharged GPU Communication
by: Ma, Shuang, et al.
Published: (2026)
by: Ma, Shuang, et al.
Published: (2026)
Similar Items
-
SpecEE: Accelerating Large Language Model Inference with Speculative Early Exiting
by: Xu, Jiaming, et al.
Published: (2025) -
Fast LLM Post-training via Decoupled and Fastest-of-N Speculation
by: Cheng, Rongxin, et al.
Published: (2025) -
Nightjar: Dynamic Adaptive Speculative Decoding for Large Language Models Serving
by: Li, Rui, et al.
Published: (2025) -
Design a Win-Win Strategy That Is Fair to Both Service Providers and Tasks When Rejection Is Not an Option
by: Trabelsi, Yohai, et al.
Published: (2024) -
SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism
by: Shen, Yuhao, et al.
Published: (2025)