FlexSpec: Frozen Drafts Meet Evolving Targets in Edge-Cloud Collaborative LLM Speculative Decoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Yuchen, Kong, Rui, Lyu, Zhonghao, Li, Qiyang, Chen, Xinran, Cai, Hengyi, Yan, Lingyong, Wang, Shuaiqiang, Zhao, Jiashu, Zhu, Guangxu, Kong, Linghe, Chen, Guihai, Xiong, Haoyi, Yin, Dawei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AdaFuse: Accelerating Dynamic Adapter Inference via Token-Level Pre-Gating and Fused Kernel Optimization
von: Li, Qiyang, et al.
Veröffentlicht: (2026)
von: Li, Qiyang, et al.
Veröffentlicht: (2026)
Generative Pre-trained Ranking Model with Over-parameterization at Web-Scale (Extended Abstract)
von: Li, Yuchen, et al.
Veröffentlicht: (2024)
von: Li, Yuchen, et al.
Veröffentlicht: (2024)
Pre-trained Graphformer-based Ranking at Web-scale Search (Extended Abstract)
von: Li, Yuchen, et al.
Veröffentlicht: (2024)
von: Li, Yuchen, et al.
Veröffentlicht: (2024)
SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism
von: Shen, Yuhao, et al.
Veröffentlicht: (2025)
von: Shen, Yuhao, et al.
Veröffentlicht: (2025)
Improving SAM for Camouflaged Object Detection via Dual Stream Adapters
von: Liu, Jiaming, et al.
Veröffentlicht: (2025)
von: Liu, Jiaming, et al.
Veröffentlicht: (2025)
SpecHub: Provable Acceleration to Multi-Draft Speculative Decoding
von: Sun, Ryan, et al.
Veröffentlicht: (2024)
von: Sun, Ryan, et al.
Veröffentlicht: (2024)
When Drafts Evolve: Speculative Decoding Meets Online Learning
von: Qian, Yu-Yang, et al.
Veröffentlicht: (2026)
von: Qian, Yu-Yang, et al.
Veröffentlicht: (2026)
Measuring Maximum Activations in Open Large Language Models
von: Chen, Luxuan, et al.
Veröffentlicht: (2026)
von: Chen, Luxuan, et al.
Veröffentlicht: (2026)
BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning
von: Xu, Yuhang, et al.
Veröffentlicht: (2026)
von: Xu, Yuhang, et al.
Veröffentlicht: (2026)
SpecBlock: Block-Iterative Speculative Decoding with Dynamic Tree Drafting
von: Shi, Weijie, et al.
Veröffentlicht: (2026)
von: Shi, Weijie, et al.
Veröffentlicht: (2026)
EndPrompt: Efficient Long-Context Extension via Terminal Anchoring
von: Tian, Han, et al.
Veröffentlicht: (2026)
von: Tian, Han, et al.
Veröffentlicht: (2026)
ML-SpecQD: Multi-Level Speculative Decoding with Quantized Drafts
von: Georganas, Evangelos, et al.
Veröffentlicht: (2025)
von: Georganas, Evangelos, et al.
Veröffentlicht: (2025)
SpecTr-GBV: Multi-Draft Block Verification Accelerating Speculative Decoding
von: Lin, Yijun, et al.
Veröffentlicht: (2026)
von: Lin, Yijun, et al.
Veröffentlicht: (2026)
SlimSpec: Low-Rank Draft LM-Head for Accelerated Speculative Decoding
von: Plaksin, Anton, et al.
Veröffentlicht: (2026)
von: Plaksin, Anton, et al.
Veröffentlicht: (2026)
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
von: Yang, Penghui, et al.
Veröffentlicht: (2025)
von: Yang, Penghui, et al.
Veröffentlicht: (2025)
FlexDraft: Flexible Speculative Decoding via Attention Tuning and Bonus-Guided Calibration
von: Zhang, Yaojie, et al.
Veröffentlicht: (2026)
von: Zhang, Yaojie, et al.
Veröffentlicht: (2026)
Learning to Draft: Adaptive Speculative Decoding with Reinforcement Learning
von: Zhang, Jiebin, et al.
Veröffentlicht: (2026)
von: Zhang, Jiebin, et al.
Veröffentlicht: (2026)
Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding
von: Shen, Yuhao, et al.
Veröffentlicht: (2026)
von: Shen, Yuhao, et al.
Veröffentlicht: (2026)
SpecMemo: Speculative Decoding is in Your Pocket
von: Yildirim, Selin, et al.
Veröffentlicht: (2025)
von: Yildirim, Selin, et al.
Veröffentlicht: (2025)
EvoSpec: Evolving Speculative Decoding via Real-Time Vocabulary and Parameter Adaptation
von: Zhang, Shuyu, et al.
Veröffentlicht: (2026)
von: Zhang, Shuyu, et al.
Veröffentlicht: (2026)
Not All Preferences Are Created Equal: Stability-Aware and Gradient-Efficient Alignment for Reasoning Models
von: Wu, Hui, et al.
Veröffentlicht: (2026)
von: Wu, Hui, et al.
Veröffentlicht: (2026)
Towards AI Search Paradigm
von: Li, Yuchen, et al.
Veröffentlicht: (2025)
von: Li, Yuchen, et al.
Veröffentlicht: (2025)
DFlare: Scaling Up Draft Capacity for Block Diffusion Speculative Decoding
von: Zhang, Jiebin, et al.
Veröffentlicht: (2026)
von: Zhang, Jiebin, et al.
Veröffentlicht: (2026)
PA-RAG: RAG Alignment via Multi-Perspective Preference Optimization
von: Wu, Jiayi, et al.
Veröffentlicht: (2024)
von: Wu, Jiayi, et al.
Veröffentlicht: (2024)
Beyond Monolithic Architectures: A Multi-Agent Search and Knowledge Optimization Framework for Agentic Search
von: Chen, Yiqun, et al.
Veröffentlicht: (2026)
von: Chen, Yiqun, et al.
Veröffentlicht: (2026)
Make Every Draft Count: Hidden State based Speculative Decoding
von: Chen, Yuetao, et al.
Veröffentlicht: (2026)
von: Chen, Yuetao, et al.
Veröffentlicht: (2026)
ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding
von: Kang, Jialiang, et al.
Veröffentlicht: (2025)
von: Kang, Jialiang, et al.
Veröffentlicht: (2025)
PEARL: Parallel Speculative Decoding with Adaptive Draft Length
von: Liu, Tianyu, et al.
Veröffentlicht: (2024)
von: Liu, Tianyu, et al.
Veröffentlicht: (2024)
ReSpec: Towards Optimizing Speculative Decoding in Reinforcement Learning Systems
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
MoE-Spec: Expert Budgeting for Efficient Speculative Decoding
von: McDanel, Bradley, et al.
Veröffentlicht: (2026)
von: McDanel, Bradley, et al.
Veröffentlicht: (2026)
HiSpec: Hierarchical Speculative Decoding for LLMs
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)
TriSpec: Ternary Speculative Decoding via Lightweight Proxy Verification
von: Jiang, Haoyun, et al.
Veröffentlicht: (2026)
von: Jiang, Haoyun, et al.
Veröffentlicht: (2026)
Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding
von: Zhang, Jun, et al.
Veröffentlicht: (2023)
von: Zhang, Jun, et al.
Veröffentlicht: (2023)
Image Super-Resolution with Text Prompt Diffusion
von: Chen, Zheng, et al.
Veröffentlicht: (2023)
von: Chen, Zheng, et al.
Veröffentlicht: (2023)
DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding
von: Li, Guanghao, et al.
Veröffentlicht: (2025)
von: Li, Guanghao, et al.
Veröffentlicht: (2025)
SpecMamba: Accelerating Mamba Inference on FPGA with Speculative Decoding
von: Zhong, Linfeng, et al.
Veröffentlicht: (2025)
von: Zhong, Linfeng, et al.
Veröffentlicht: (2025)
SpecVLM: Fast Speculative Decoding in Vision-Language Models
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
SJD-PAC: Accelerating Speculative Jacobi Decoding via Proactive Drafting and Adaptive Continuation
von: Kang, Jialiang, et al.
Veröffentlicht: (2026)
von: Kang, Jialiang, et al.
Veröffentlicht: (2026)
SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning
von: Ji, Yicheng, et al.
Veröffentlicht: (2025)
von: Ji, Yicheng, et al.
Veröffentlicht: (2025)
NanoSpec: Accelerating Speculative Decoding using Minimalist In-Context Vocabularies
von: Chen, Zhiyang, et al.
Veröffentlicht: (2026)
von: Chen, Zhiyang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
AdaFuse: Accelerating Dynamic Adapter Inference via Token-Level Pre-Gating and Fused Kernel Optimization
von: Li, Qiyang, et al.
Veröffentlicht: (2026) -
Generative Pre-trained Ranking Model with Over-parameterization at Web-Scale (Extended Abstract)
von: Li, Yuchen, et al.
Veröffentlicht: (2024) -
Pre-trained Graphformer-based Ranking at Web-scale Search (Extended Abstract)
von: Li, Yuchen, et al.
Veröffentlicht: (2024) -
SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism
von: Shen, Yuhao, et al.
Veröffentlicht: (2025) -
Improving SAM for Camouflaged Object Detection via Dual Stream Adapters
von: Liu, Jiaming, et al.
Veröffentlicht: (2025)