Dynamic Delayed Tree Expansion For Improved Multi-Path Speculative Decoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Thomas, Rahul, Kitanovski, Teo, Goldblum, Micah, Pal, Arka |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Knowing What You Know Is Not Enough: Large Language Model Confidences Don't Align With Their Actions
von: Pal, Arka, et al.
Veröffentlicht: (2025)
von: Pal, Arka, et al.
Veröffentlicht: (2025)
Greedy Multi-Path Block Verification for Faster Decoding in Speculative Sampling
von: Thomas, Rahul, et al.
Veröffentlicht: (2026)
von: Thomas, Rahul, et al.
Veröffentlicht: (2026)
Global Resolution: Optimal Multi-Draft Speculative Sampling via Convex Minimization
von: Thomas, Rahul Krishna, et al.
Veröffentlicht: (2025)
von: Thomas, Rahul Krishna, et al.
Veröffentlicht: (2025)
Cascade: Token-Sharded Private LLM Inference
von: Thomas, Rahul, et al.
Veröffentlicht: (2025)
von: Thomas, Rahul, et al.
Veröffentlicht: (2025)
An Attack to Break Permutation-Based Private Third-Party Inference Schemes for LLMs
von: Thomas, Rahul, et al.
Veröffentlicht: (2025)
von: Thomas, Rahul, et al.
Veröffentlicht: (2025)
vTune: Verifiable Fine-Tuning for LLMs Through Backdooring
von: Zhang, Eva, et al.
Veröffentlicht: (2024)
von: Zhang, Eva, et al.
Veröffentlicht: (2024)
Privacy-Preserving Mechanisms Enable Cheap Verifiable Inference of LLMs
von: Pal, Arka, et al.
Veröffentlicht: (2026)
von: Pal, Arka, et al.
Veröffentlicht: (2026)
Speculative Speculative Decoding
von: Kumar, Tanishq, et al.
Veröffentlicht: (2026)
von: Kumar, Tanishq, et al.
Veröffentlicht: (2026)
The No Free Lunch Theorem, Kolmogorov Complexity, and the Role of Inductive Biases in Machine Learning
von: Goldblum, Micah, et al.
Veröffentlicht: (2023)
von: Goldblum, Micah, et al.
Veröffentlicht: (2023)
DySpec: Faster Speculative Decoding with Dynamic Token Tree Structure
von: Xiong, Yunfan, et al.
Veröffentlicht: (2024)
von: Xiong, Yunfan, et al.
Veröffentlicht: (2024)
Multi-Token Prediction via Self-Distillation
von: Kirchenbauer, John, et al.
Veröffentlicht: (2026)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2026)
Decoding Speculative Decoding
von: Yan, Minghao, et al.
Veröffentlicht: (2024)
von: Yan, Minghao, et al.
Veröffentlicht: (2024)
Large Language Models Must Be Taught to Know What They Don't Know
von: Kapoor, Sanyam, et al.
Veröffentlicht: (2024)
von: Kapoor, Sanyam, et al.
Veröffentlicht: (2024)
Draft, Verify, and Improve: Toward Training-Aware Speculative Decoding
von: Bhansali, Shrenik, et al.
Veröffentlicht: (2025)
von: Bhansali, Shrenik, et al.
Veröffentlicht: (2025)
Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful
von: Marek, Martin, et al.
Veröffentlicht: (2025)
von: Marek, Martin, et al.
Veröffentlicht: (2025)
Compute Better Spent: Replacing Dense Layers with Structured Matrices
von: Qiu, Shikai, et al.
Veröffentlicht: (2024)
von: Qiu, Shikai, et al.
Veröffentlicht: (2024)
Adaptive Retention & Correction: Test-Time Training for Continual Learning
von: Chen, Haoran, et al.
Veröffentlicht: (2024)
von: Chen, Haoran, et al.
Veröffentlicht: (2024)
On Training in Imagination
von: Timor, Nadav, et al.
Veröffentlicht: (2026)
von: Timor, Nadav, et al.
Veröffentlicht: (2026)
Traversal Verification for Speculative Tree Decoding
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks
von: Li, Ang, et al.
Veröffentlicht: (2025)
von: Li, Ang, et al.
Veröffentlicht: (2025)
The Lie Derivative for Measuring Learned Equivariance
von: Gruver, Nate, et al.
Veröffentlicht: (2022)
von: Gruver, Nate, et al.
Veröffentlicht: (2022)
Unlocking Tokens as Data Points for Generalization Bounds on Larger Language Models
von: Lotfi, Sanae, et al.
Veröffentlicht: (2024)
von: Lotfi, Sanae, et al.
Veröffentlicht: (2024)
Speculative Decoding Across Languages
von: Paudel, Nirajan, et al.
Veröffentlicht: (2026)
von: Paudel, Nirajan, et al.
Veröffentlicht: (2026)
Speculative Safety-Aware Decoding
von: Wang, Xuekang, et al.
Veröffentlicht: (2025)
von: Wang, Xuekang, et al.
Veröffentlicht: (2025)
DistillSpec: Improving Speculative Decoding via Knowledge Distillation
von: Zhou, Yongchao, et al.
Veröffentlicht: (2023)
von: Zhou, Yongchao, et al.
Veröffentlicht: (2023)
Online Speculative Decoding
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
Yggdrasil: Bridging Dynamic Speculation and Static Runtime for Latency-Optimal Tree-Based LLM Decoding
von: Guan, Yue, et al.
Veröffentlicht: (2025)
von: Guan, Yue, et al.
Veröffentlicht: (2025)
Fast Inference via Hierarchical Speculative Decoding
von: Mohri, Clara, et al.
Veröffentlicht: (2025)
von: Mohri, Clara, et al.
Veröffentlicht: (2025)
SpecPV: Improving Self-Speculative Decoding for Long-Context Generation via Partial Verification
von: Tan, Zhendong, et al.
Veröffentlicht: (2025)
von: Tan, Zhendong, et al.
Veröffentlicht: (2025)
TapOut: A Bandit-Based Approach to Dynamic Speculative Decoding
von: Sridhar, Aditya, et al.
Veröffentlicht: (2025)
von: Sridhar, Aditya, et al.
Veröffentlicht: (2025)
SpecMER: Fast Protein Generation with K-mer Guided Speculative Decoding
von: Walton, Thomas, et al.
Veröffentlicht: (2025)
von: Walton, Thomas, et al.
Veröffentlicht: (2025)
STree: Speculative Tree Decoding for Hybrid State-Space Models
von: Wu, Yangchao, et al.
Veröffentlicht: (2025)
von: Wu, Yangchao, et al.
Veröffentlicht: (2025)
Scaling Speculative Decoding with Lookahead Reasoning
von: Fu, Yichao, et al.
Veröffentlicht: (2025)
von: Fu, Yichao, et al.
Veröffentlicht: (2025)
UniVer: A Unified Perspective for Multi-step and Multi-draft Speculative Decoding
von: Weng, Yepeng, et al.
Veröffentlicht: (2026)
von: Weng, Yepeng, et al.
Veröffentlicht: (2026)
Steering Pretrained Drafters during Speculative Decoding
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2025)
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2025)
Mixture of Attentions For Speculative Decoding
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
Polybasic Speculative Decoding Through a Theoretical Perspective
von: Wang, Ruilin, et al.
Veröffentlicht: (2025)
von: Wang, Ruilin, et al.
Veröffentlicht: (2025)
FastEagle: Cascaded Drafting for Accelerating Speculative Decoding
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
SPIRe: Boosting LLM Inference Throughput with Speculative Decoding
von: Neelam, Sanjit, et al.
Veröffentlicht: (2025)
von: Neelam, Sanjit, et al.
Veröffentlicht: (2025)
Accelerating Time Series Foundation Models with Speculative Decoding
von: Subbaraman, Pranav, et al.
Veröffentlicht: (2025)
von: Subbaraman, Pranav, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Knowing What You Know Is Not Enough: Large Language Model Confidences Don't Align With Their Actions
von: Pal, Arka, et al.
Veröffentlicht: (2025) -
Greedy Multi-Path Block Verification for Faster Decoding in Speculative Sampling
von: Thomas, Rahul, et al.
Veröffentlicht: (2026) -
Global Resolution: Optimal Multi-Draft Speculative Sampling via Convex Minimization
von: Thomas, Rahul Krishna, et al.
Veröffentlicht: (2025) -
Cascade: Token-Sharded Private LLM Inference
von: Thomas, Rahul, et al.
Veröffentlicht: (2025) -
An Attack to Break Permutation-Based Private Third-Party Inference Schemes for LLMs
von: Thomas, Rahul, et al.
Veröffentlicht: (2025)