The N-Grammys: Accelerating Autoregressive Inference with Learning-Free Batched Speculation
Fuente:
arXiv
Saved in:
| Main Authors: | Stewart, Lawrence, Trager, Matthew, Gonugondla, Sujan Kumar, Soatto, Stefano |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BASS: Batched Attention-optimized Speculative Sampling
by: Qian, Haifeng, et al.
Published: (2024)
by: Qian, Haifeng, et al.
Published: (2024)
e1: Learning Adaptive Control of Reasoning Effort
by: Kleinman, Michael, et al.
Published: (2025)
by: Kleinman, Michael, et al.
Published: (2025)
Compositional Structures in Neural Embedding and Interaction Decompositions
by: Trager, Matthew, et al.
Published: (2024)
by: Trager, Matthew, et al.
Published: (2024)
Experience-Guided Adaptation of Inference-Time Reasoning Strategies
by: Stein, Adam, et al.
Published: (2025)
by: Stein, Adam, et al.
Published: (2025)
Learning When to Attend: Conditional Memory Access for Long-Context LLMs
by: Choudhary, Sakshi, et al.
Published: (2026)
by: Choudhary, Sakshi, et al.
Published: (2026)
STree: Speculative Tree Decoding for Hybrid State-Space Models
by: Wu, Yangchao, et al.
Published: (2025)
by: Wu, Yangchao, et al.
Published: (2025)
PICASO: Permutation-Invariant Context Composition with State Space Models
by: Liu, Tian Yu, et al.
Published: (2025)
by: Liu, Tian Yu, et al.
Published: (2025)
EvoMAS: Evolutionary Generation of Multi-Agent Systems
by: Hu, Yuntong, et al.
Published: (2026)
by: Hu, Yuntong, et al.
Published: (2026)
Linear Spaces of Meanings: Compositional Structures in Vision-Language Models
by: Trager, Matthew, et al.
Published: (2023)
by: Trager, Matthew, et al.
Published: (2023)
Interpretable Measures of Conceptual Similarity by Complexity-Constrained Descriptive Auto-Encoding
by: Achille, Alessandro, et al.
Published: (2024)
by: Achille, Alessandro, et al.
Published: (2024)
Reviving Any-Subset Autoregressive Models with Principled Parallel Sampling and Speculative Decoding
by: Guo, Gabe, et al.
Published: (2025)
by: Guo, Gabe, et al.
Published: (2025)
Speculating Experts Accelerates Inference for Mixture-of-Experts
by: Madan, Vivan, et al.
Published: (2026)
by: Madan, Vivan, et al.
Published: (2026)
Cycles of Thought: Measuring LLM Confidence through Stable Explanations
by: Becker, Evan, et al.
Published: (2024)
by: Becker, Evan, et al.
Published: (2024)
AI Agents as Universal Task Solvers
by: Achille, Alessandro, et al.
Published: (2025)
by: Achille, Alessandro, et al.
Published: (2025)
Learning to Focus: Focal Attention for Selective and Scalable Transformers
by: Ram, Dhananjay, et al.
Published: (2025)
by: Ram, Dhananjay, et al.
Published: (2025)
Robust Planning for Autonomous Driving via Mixed Adversarial Diffusion Predictions
by: Zhao, Albert, et al.
Published: (2025)
by: Zhao, Albert, et al.
Published: (2025)
Multi-Modal Hallucination Control by Visual Information Grounding
by: Favero, Alessandro, et al.
Published: (2024)
by: Favero, Alessandro, et al.
Published: (2024)
Bifurcated Attention: Accelerating Massively Parallel Decoding with Shared Prefixes in LLMs
by: Athiwaratkun, Ben, et al.
Published: (2024)
by: Athiwaratkun, Ben, et al.
Published: (2024)
Critical Learning Periods Emerge Even in Deep Linear Networks
by: Kleinman, Michael, et al.
Published: (2023)
by: Kleinman, Michael, et al.
Published: (2023)
Compiler-Assisted Speculative Sampling for Accelerated LLM Inference on Heterogeneous Edge Devices
by: Mesa, Alejandro Ruiz y, et al.
Published: (2026)
by: Mesa, Alejandro Ruiz y, et al.
Published: (2026)
SpecPipe: Accelerating Pipeline Parallelism-based LLM Inference with Speculative Decoding
by: Yin, Haofei, et al.
Published: (2025)
by: Yin, Haofei, et al.
Published: (2025)
Algebra Unveils Deep Learning -- An Invitation to Neuroalgebraic Geometry
by: Marchetti, Giovanni Luca, et al.
Published: (2025)
by: Marchetti, Giovanni Luca, et al.
Published: (2025)
B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading Memory
by: Zancato, Luca, et al.
Published: (2024)
by: Zancato, Luca, et al.
Published: (2024)
Tangent Transformers for Composition, Privacy and Removal
by: Liu, Tian Yu, et al.
Published: (2023)
by: Liu, Tian Yu, et al.
Published: (2023)
Speculative Speculative Decoding
by: Kumar, Tanishq, et al.
Published: (2026)
by: Kumar, Tanishq, et al.
Published: (2026)
SpecDiff: Accelerating Diffusion Model Inference with Self-Speculation
by: Pan, Jiayi, et al.
Published: (2025)
by: Pan, Jiayi, et al.
Published: (2025)
Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement
by: Jeon, Wonseok, et al.
Published: (2024)
by: Jeon, Wonseok, et al.
Published: (2024)
Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding
by: Zhao, Yilong, et al.
Published: (2025)
by: Zhao, Yilong, et al.
Published: (2025)
CATS: Cascaded Adaptive Tree Speculation for Memory-Limited LLM Inference Acceleration
by: Han, Yuning, et al.
Published: (2026)
by: Han, Yuning, et al.
Published: (2026)
Attention Drift: What Autoregressive Speculative Decoding Models Learn
by: Eldenk, Doğaç, et al.
Published: (2026)
by: Eldenk, Doğaç, et al.
Published: (2026)
Heat Death of Generative Models in Closed-Loop Learning
by: Marchi, Matteo, et al.
Published: (2024)
by: Marchi, Matteo, et al.
Published: (2024)
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
by: Wen, Zhuofan, et al.
Published: (2024)
by: Wen, Zhuofan, et al.
Published: (2024)
Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies
by: Timor, Nadav, et al.
Published: (2025)
by: Timor, Nadav, et al.
Published: (2025)
Lever: Speculative LLM Inference on Smartphones
by: Wang, Tuowei, et al.
Published: (2026)
by: Wang, Tuowei, et al.
Published: (2026)
Diversified Batch Selection for Training Acceleration
by: Hong, Feng, et al.
Published: (2024)
by: Hong, Feng, et al.
Published: (2024)
Accelerating Inference of Discrete Autoregressive Normalizing Flows by Selective Jacobi Decoding
by: Zhang, Jiaru, et al.
Published: (2025)
by: Zhang, Jiaru, et al.
Published: (2025)
Accelerated Diffusion Models via Speculative Sampling
by: De Bortoli, Valentin, et al.
Published: (2025)
by: De Bortoli, Valentin, et al.
Published: (2025)
Fast Inference via Hierarchical Speculative Decoding
by: Mohri, Clara, et al.
Published: (2025)
by: Mohri, Clara, et al.
Published: (2025)
CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMs
by: Ning, Zhiyuan, et al.
Published: (2025)
by: Ning, Zhiyuan, et al.
Published: (2025)
SRT: Accelerating Reinforcement Learning via Speculative Rollout with Tree-Structured Cache
by: Chang, Chi-Chih, et al.
Published: (2026)
by: Chang, Chi-Chih, et al.
Published: (2026)
Similar Items
-
BASS: Batched Attention-optimized Speculative Sampling
by: Qian, Haifeng, et al.
Published: (2024) -
e1: Learning Adaptive Control of Reasoning Effort
by: Kleinman, Michael, et al.
Published: (2025) -
Compositional Structures in Neural Embedding and Interaction Decompositions
by: Trager, Matthew, et al.
Published: (2024) -
Experience-Guided Adaptation of Inference-Time Reasoning Strategies
by: Stein, Adam, et al.
Published: (2025) -
Learning When to Attend: Conditional Memory Access for Long-Context LLMs
by: Choudhary, Sakshi, et al.
Published: (2026)