Judge Decoding: Faster Speculative Sampling Requires Going Beyond Model Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Bachmann, Gregor, Anagnostidis, Sotiris, Pumarola, Albert, Georgopoulos, Markos, Sanakoyeu, Artsiom, Du, Yuming, Schönfeld, Edgar, Thabet, Ali, Kohler, Jonas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FlexiDiT: Your Diffusion Transformer Can Easily Generate High-Quality Samples with Less Compute
by: Anagnostidis, Sotiris, et al.
Published: (2025)
by: Anagnostidis, Sotiris, et al.
Published: (2025)
Autoregressive Distillation of Diffusion Transformers
by: Kim, Yeongmin, et al.
Published: (2025)
by: Kim, Yeongmin, et al.
Published: (2025)
Imagine Flash: Accelerating Emu Diffusion Models with Backward Distillation
by: Kohler, Jonas, et al.
Published: (2024)
by: Kohler, Jonas, et al.
Published: (2024)
Navigating Scaling Laws: Compute Optimality in Adaptive Model Training
by: Anagnostidis, Sotiris, et al.
Published: (2023)
by: Anagnostidis, Sotiris, et al.
Published: (2023)
SneakPeek: Future-Guided Instructional Streaming Video Generation
by: Hong, Cheeun, et al.
Published: (2025)
by: Hong, Cheeun, et al.
Published: (2025)
A Language Model's Guide Through Latent Space
by: von Rütte, Dimitri, et al.
Published: (2024)
by: von Rütte, Dimitri, et al.
Published: (2024)
How Susceptible are LLMs to Influence in Prompts?
by: Anagnostidis, Sotiris, et al.
Published: (2024)
by: Anagnostidis, Sotiris, et al.
Published: (2024)
SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
by: Yoon, Kanghoon, et al.
Published: (2025)
by: Yoon, Kanghoon, et al.
Published: (2025)
GRIFFIN: Effective Token Alignment for Faster Speculative Decoding
by: Hu, Shijing, et al.
Published: (2025)
by: Hu, Shijing, et al.
Published: (2025)
Greedy Multi-Path Block Verification for Faster Decoding in Speculative Sampling
by: Thomas, Rahul, et al.
Published: (2026)
by: Thomas, Rahul, et al.
Published: (2026)
Faster Cascades via Speculative Decoding
by: Narasimhan, Harikrishna, et al.
Published: (2024)
by: Narasimhan, Harikrishna, et al.
Published: (2024)
SpecDiff-2: Scaling Diffusion Drafter Alignment For Faster Speculative Decoding
by: Sandler, Jameson, et al.
Published: (2025)
by: Sandler, Jameson, et al.
Published: (2025)
Bespoke Non-Stationary Solvers for Fast Sampling of Diffusion and Flow Models
by: Shaul, Neta, et al.
Published: (2024)
by: Shaul, Neta, et al.
Published: (2024)
Alignment-Augmented Speculative Decoding with Alignment Sampling and Conditional Verification
by: Wang, Jikai, et al.
Published: (2025)
by: Wang, Jikai, et al.
Published: (2025)
Dynamic Depth Decoding: Faster Speculative Decoding for LLMs
by: Brown, Oscar, et al.
Published: (2024)
by: Brown, Oscar, et al.
Published: (2024)
Using Motion Cues to Supervise Single-Frame Body Pose and Shape Estimation in Low Data Regimes
by: Davydov, Andrey, et al.
Published: (2024)
by: Davydov, Andrey, et al.
Published: (2024)
MegaPortrait: Revisiting Diffusion Control for High-fidelity Portrait Generation
by: Yang, Han, et al.
Published: (2024)
by: Yang, Han, et al.
Published: (2024)
Self-Speculative Biased Decoding for Faster Re-Translation
by: Zeng, Linxiao, et al.
Published: (2025)
by: Zeng, Linxiao, et al.
Published: (2025)
Speculative Decoding via Early-exiting for Faster LLM Inference with Thompson Sampling Control Mechanism
by: Liu, Jiahao, et al.
Published: (2024)
by: Liu, Jiahao, et al.
Published: (2024)
Annealed Relaxation of Speculative Decoding for Faster Autoregressive Image Generation
by: Li, Xingyao, et al.
Published: (2026)
by: Li, Xingyao, et al.
Published: (2026)
Boosting Lossless Speculative Decoding via Feature Sampling and Partial Alignment Distillation
by: Gui, Lujun, et al.
Published: (2024)
by: Gui, Lujun, et al.
Published: (2024)
Cache Me if You Can: Accelerating Diffusion Models through Block Caching
by: Wimbauer, Felix, et al.
Published: (2023)
by: Wimbauer, Felix, et al.
Published: (2023)
Ouroboros: Generating Longer Drafts Phrase by Phrase for Faster Speculative Decoding
by: Zhao, Weilin, et al.
Published: (2024)
by: Zhao, Weilin, et al.
Published: (2024)
Balancing Coverage and Draft Latency in Vocabulary Trimming for Faster Speculative Decoding
by: Shoham, Ofir Ben
Published: (2026)
by: Shoham, Ofir Ben
Published: (2026)
DySpec: Faster Speculative Decoding with Dynamic Token Tree Structure
by: Xiong, Yunfan, et al.
Published: (2024)
by: Xiong, Yunfan, et al.
Published: (2024)
Speculative Decoding for Multi-Sample Inference
by: Li, Yiwei, et al.
Published: (2025)
by: Li, Yiwei, et al.
Published: (2025)
Multi-Drafter Speculative Decoding with Alignment Feedback
by: Kim, Taehyeon, et al.
Published: (2026)
by: Kim, Taehyeon, et al.
Published: (2026)
Think Before You Accept: Semantic Reflective Verification for Faster Speculative Decoding
by: Wang, Yixuan, et al.
Published: (2025)
by: Wang, Yixuan, et al.
Published: (2025)
Exploring Magnitude Preservation and Rotation Modulation in Diffusion Transformers
by: Bill, Eric Tillman, et al.
Published: (2025)
by: Bill, Eric Tillman, et al.
Published: (2025)
EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
by: Li, Yuhui, et al.
Published: (2024)
by: Li, Yuhui, et al.
Published: (2024)
Speculative Decoding and Beyond: An In-Depth Survey of Techniques
by: Hu, Yunhai, et al.
Published: (2025)
by: Hu, Yunhai, et al.
Published: (2025)
Out-of-Vocabulary Sampling Boosts Speculative Decoding
by: Timor, Nadav, et al.
Published: (2025)
by: Timor, Nadav, et al.
Published: (2025)
Speculative Speculative Decoding
by: Kumar, Tanishq, et al.
Published: (2026)
by: Kumar, Tanishq, et al.
Published: (2026)
IC-Portrait: In-Context Matching for View-Consistent Personalized Portrait
by: Yang, Han, et al.
Published: (2025)
by: Yang, Han, et al.
Published: (2025)
Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers
by: Anagnostidis, Sotiris, et al.
Published: (2023)
by: Anagnostidis, Sotiris, et al.
Published: (2023)
Beyond the Target: From Imitation to Collaboration in Speculative Decoding
by: Li, Jinze, et al.
Published: (2026)
by: Li, Jinze, et al.
Published: (2026)
Decoding Speculative Decoding
by: Yan, Minghao, et al.
Published: (2024)
by: Yan, Minghao, et al.
Published: (2024)
Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding
by: Lin, Zijian, et al.
Published: (2025)
by: Lin, Zijian, et al.
Published: (2025)
When Quantum and Classical Models Disagree: Learning Beyond Minimum Norm Least Square
by: Thabet, Slimane, et al.
Published: (2024)
by: Thabet, Slimane, et al.
Published: (2024)
XR-MBT: Multi-modal Full Body Tracking for XR through Self-Supervision with Learned Depth Point Cloud Registration
by: Rozumnyi, Denys, et al.
Published: (2024)
by: Rozumnyi, Denys, et al.
Published: (2024)
Similar Items
-
FlexiDiT: Your Diffusion Transformer Can Easily Generate High-Quality Samples with Less Compute
by: Anagnostidis, Sotiris, et al.
Published: (2025) -
Autoregressive Distillation of Diffusion Transformers
by: Kim, Yeongmin, et al.
Published: (2025) -
Imagine Flash: Accelerating Emu Diffusion Models with Backward Distillation
by: Kohler, Jonas, et al.
Published: (2024) -
Navigating Scaling Laws: Compute Optimality in Adaptive Model Training
by: Anagnostidis, Sotiris, et al.
Published: (2023) -
SneakPeek: Future-Guided Instructional Streaming Video Generation
by: Hong, Cheeun, et al.
Published: (2025)