Break the Sequential Dependency of LLM Inference Using Lookahead Decoding
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Fu, Yichao, Bailis, Peter, Stoica, Ion, Zhang, Hao |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Scaling Speculative Decoding with Lookahead Reasoning
par: Fu, Yichao, et autres
Publié: (2025)
par: Fu, Yichao, et autres
Publié: (2025)
Online Speculative Decoding
par: Liu, Xiaoxuan, et autres
Publié: (2023)
par: Liu, Xiaoxuan, et autres
Publié: (2023)
Are More LLM Calls All You Need? Towards Scaling Laws of Compound Inference Systems
par: Chen, Lingjiao, et autres
Publié: (2024)
par: Chen, Lingjiao, et autres
Publié: (2024)
Efficiently Scaling LLM Reasoning with Certaindex
par: Fu, Yichao, et autres
Publié: (2024)
par: Fu, Yichao, et autres
Publié: (2024)
Optimizing Model Selection for Compound AI Systems
par: Chen, Lingjiao, et autres
Publié: (2025)
par: Chen, Lingjiao, et autres
Publié: (2025)
Progressive Mixed-Precision Decoding for Efficient LLM Inference
par: Chen, Hao Mark, et autres
Publié: (2024)
par: Chen, Hao Mark, et autres
Publié: (2024)
Causal Attention with Lookahead Keys
par: Song, Zhuoqing, et autres
Publié: (2025)
par: Song, Zhuoqing, et autres
Publié: (2025)
Efficient LLM Scheduling by Learning to Rank
par: Fu, Yichao, et autres
Publié: (2024)
par: Fu, Yichao, et autres
Publié: (2024)
Faster LLM Inference via Sequential Monte Carlo
par: Emara, Yahya, et autres
Publié: (2026)
par: Emara, Yahya, et autres
Publié: (2026)
Hardware-Aware Parallel Prompt Decoding for Memory-Efficient Acceleration of LLM Inference
par: Chen, Hao Mark, et autres
Publié: (2024)
par: Chen, Hao Mark, et autres
Publié: (2024)
Thinking into the Future: Latent Lookahead Training for Transformers
par: Noci, Lorenzo, et autres
Publié: (2026)
par: Noci, Lorenzo, et autres
Publié: (2026)
Optimized Multi-Token Joint Decoding with Auxiliary Model for LLM Inference
par: Qin, Zongyue, et autres
Publié: (2024)
par: Qin, Zongyue, et autres
Publié: (2024)
Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
par: Cai, Tianle, et autres
Publié: (2024)
par: Cai, Tianle, et autres
Publié: (2024)
Breaking the Language Barrier: Can Direct Inference Outperform Pre-Translation in Multilingual LLM Applications?
par: Intrator, Yotam, et autres
Publié: (2024)
par: Intrator, Yotam, et autres
Publié: (2024)
Pie: Pooling CPU Memory for LLM Inference
par: Xu, Yi, et autres
Publié: (2024)
par: Xu, Yi, et autres
Publié: (2024)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
par: Tan, Sijun, et autres
Publié: (2024)
par: Tan, Sijun, et autres
Publié: (2024)
RouteLLM: Learning to Route LLMs with Preference Data
par: Ong, Isaac, et autres
Publié: (2024)
par: Ong, Isaac, et autres
Publié: (2024)
Post-Training Sparse Attention with Double Sparsity
par: Yang, Shuo, et autres
Publié: (2024)
par: Yang, Shuo, et autres
Publié: (2024)
Calibrated Speculative Decoding: Frequency-Guided Candidate Selection for Efficient Inference
par: Zhou, Xuwen, et autres
Publié: (2026)
par: Zhou, Xuwen, et autres
Publié: (2026)
When Linear Attention Meets Autoregressive Decoding: Towards More Effective and Efficient Linearized Large Language Models
par: You, Haoran, et autres
Publié: (2024)
par: You, Haoran, et autres
Publié: (2024)
Prompt-to-Leaderboard
par: Frick, Evan, et autres
Publié: (2025)
par: Frick, Evan, et autres
Publié: (2025)
R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents
par: Jain, Naman, et autres
Publié: (2025)
par: Jain, Naman, et autres
Publié: (2025)
MPC-Minimized Secure LLM Inference
par: Rathee, Deevashwer, et autres
Publié: (2024)
par: Rathee, Deevashwer, et autres
Publié: (2024)
LoPA: Scaling dLLM Inference via Lookahead Parallel Decoding
par: Xu, Chenkai, et autres
Publié: (2025)
par: Xu, Chenkai, et autres
Publié: (2025)
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
par: Wen, Zhuofan, et autres
Publié: (2024)
par: Wen, Zhuofan, et autres
Publié: (2024)
Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies
par: Timor, Nadav, et autres
Publié: (2025)
par: Timor, Nadav, et autres
Publié: (2025)
Single Character Perturbations Break LLM Alignment
par: Lin, Leon, et autres
Publié: (2024)
par: Lin, Leon, et autres
Publié: (2024)
Controlled LLM Decoding via Discrete Auto-regressive Biasing
par: Pynadath, Patrick, et autres
Publié: (2025)
par: Pynadath, Patrick, et autres
Publié: (2025)
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency
par: Li, Ruixiao, et autres
Publié: (2025)
par: Li, Ruixiao, et autres
Publié: (2025)
Twilight: Adaptive Attention Sparsity with Hierarchical Top-$p$ Pruning
par: Lin, Chaofan, et autres
Publié: (2025)
par: Lin, Chaofan, et autres
Publié: (2025)
Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length
par: Ma, Xuezhe, et autres
Publié: (2024)
par: Ma, Xuezhe, et autres
Publié: (2024)
Imagine-then-Plan: Agent Learning from Adaptive Lookahead with World Models
par: Liu, Youwei, et autres
Publié: (2026)
par: Liu, Youwei, et autres
Publié: (2026)
Clover: Regressive Lightweight Speculative Decoding with Sequential Knowledge
par: Xiao, Bin, et autres
Publié: (2024)
par: Xiao, Bin, et autres
Publié: (2024)
FlashDecoding++: Faster Large Language Model Inference on GPUs
par: Hong, Ke, et autres
Publié: (2023)
par: Hong, Ke, et autres
Publié: (2023)
Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
par: Roytburg, Dani, et autres
Publié: (2025)
par: Roytburg, Dani, et autres
Publié: (2025)
Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding
par: Yang, Seongjun, et autres
Publié: (2023)
par: Yang, Seongjun, et autres
Publié: (2023)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
par: Jiang, Xuanlin, et autres
Publié: (2024)
par: Jiang, Xuanlin, et autres
Publié: (2024)
LLM In-Context Recall is Prompt Dependent
par: Machlab, Daniel, et autres
Publié: (2024)
par: Machlab, Daniel, et autres
Publié: (2024)
Transfer Q Star: Principled Decoding for LLM Alignment
par: Chakraborty, Souradip, et autres
Publié: (2024)
par: Chakraborty, Souradip, et autres
Publié: (2024)
GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents
par: Shetty, Manish, et autres
Publié: (2025)
par: Shetty, Manish, et autres
Publié: (2025)
Documents similaires
-
Scaling Speculative Decoding with Lookahead Reasoning
par: Fu, Yichao, et autres
Publié: (2025) -
Online Speculative Decoding
par: Liu, Xiaoxuan, et autres
Publié: (2023) -
Are More LLM Calls All You Need? Towards Scaling Laws of Compound Inference Systems
par: Chen, Lingjiao, et autres
Publié: (2024) -
Efficiently Scaling LLM Reasoning with Certaindex
par: Fu, Yichao, et autres
Publié: (2024) -
Optimizing Model Selection for Compound AI Systems
par: Chen, Lingjiao, et autres
Publié: (2025)