Steering Pretrained Drafters during Speculative Decoding
Fuente:
arXiv
Saved in:
| Main Authors: | Berdoz, Frédéric, Rheinboldt, Peer, Wattenhofer, Roger |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Alignment-Aware Decoding
by: Berdoz, Frédéric, et al.
Published: (2025)
by: Berdoz, Frédéric, et al.
Published: (2025)
An Interpretable Latency Model for Speculative Decoding in LLM Serving
by: Kong, Linghao, et al.
Published: (2026)
by: Kong, Linghao, et al.
Published: (2026)
Light Differentiable Logic Gate Networks
by: Rüttgers, Lukas, et al.
Published: (2025)
by: Rüttgers, Lukas, et al.
Published: (2025)
Mind the Gap: Removing the Discretization Gap in Differentiable Logic Gate Networks
by: Yousefi, Shakir, et al.
Published: (2025)
by: Yousefi, Shakir, et al.
Published: (2025)
Can AI Agents Agree?
by: Berdoz, Frédéric, et al.
Published: (2026)
by: Berdoz, Frédéric, et al.
Published: (2026)
BitLogic: Training Framework for Gradient-Based FPGA-Native Neural Networks
by: Bührer, Simon, et al.
Published: (2026)
by: Bührer, Simon, et al.
Published: (2026)
Subliminal Signals in Preference Labels
by: Magistrali, Isotta, et al.
Published: (2026)
by: Magistrali, Isotta, et al.
Published: (2026)
Reasoning Boosts Opinion Alignment in LLMs
by: Berdoz, Frédéric, et al.
Published: (2026)
by: Berdoz, Frédéric, et al.
Published: (2026)
N-vium: Mixture-of-Exits Transformer for Accelerated Exact Generation
by: Lorenc, Aleksander, et al.
Published: (2026)
by: Lorenc, Aleksander, et al.
Published: (2026)
ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
by: Xiao, Zilin, et al.
Published: (2024)
by: Xiao, Zilin, et al.
Published: (2024)
High-Fidelity Speech Enhancement via Discrete Audio Tokens
by: Lanzendörfer, Luca A., et al.
Published: (2025)
by: Lanzendörfer, Luca A., et al.
Published: (2025)
Text-to-Scene with Large Reasoning Models
by: Berdoz, Frédéric, et al.
Published: (2025)
by: Berdoz, Frédéric, et al.
Published: (2025)
WorldSpeech: A Multilingual Speech Corpus from Around the World
by: Asonitis, Antonis, et al.
Published: (2026)
by: Asonitis, Antonis, et al.
Published: (2026)
Recurrent Drafter for Fast Speculative Decoding in Large Language Models
by: Cheng, Yunfei, et al.
Published: (2024)
by: Cheng, Yunfei, et al.
Published: (2024)
Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMs
by: Liu, Hongyi, et al.
Published: (2025)
by: Liu, Hongyi, et al.
Published: (2025)
Coupling without Communication and Drafter-Invariant Speculative Decoding
by: Daliri, Majid, et al.
Published: (2024)
by: Daliri, Majid, et al.
Published: (2024)
Single-Thread JPEG Decoder Benchmarks Mis-Evaluate ML Data Loaders
by: Iglovikov, Vladimir, et al.
Published: (2026)
by: Iglovikov, Vladimir, et al.
Published: (2026)
Reducing Latency of LLM Search Agent via Speculation-based Algorithm-System Co-Design
by: Huang, Zixiao, et al.
Published: (2025)
by: Huang, Zixiao, et al.
Published: (2025)
Leveraging Speculative Sampling and KV-Cache Optimizations Together for Generative AI using OpenVINO
by: Barad, Haim, et al.
Published: (2023)
by: Barad, Haim, et al.
Published: (2023)
Can an AI Agent Safely Run a Government? Existence of Probably Approximately Aligned Policies
by: Berdoz, Frédéric, et al.
Published: (2024)
by: Berdoz, Frédéric, et al.
Published: (2024)
OmniDraft: A Cross-vocabulary, Online Adaptive Drafter for On-device Speculative Decoding
by: Ramakrishnan, Ramchalam Kinattinkara, et al.
Published: (2025)
by: Ramakrishnan, Ramchalam Kinattinkara, et al.
Published: (2025)
Prefill vs. Decode Bottlenecks: SRAM-Frequency Tradeoffs and the Memory-Bandwidth Ceiling
by: Atmer, Hannah, et al.
Published: (2025)
by: Atmer, Hannah, et al.
Published: (2025)
HiViS: Hiding Visual Tokens from the Drafter for Speculative Decoding in Vision-Language Models
by: Xie, Zhinan, et al.
Published: (2025)
by: Xie, Zhinan, et al.
Published: (2025)
MIREncoder: Multi-modal IR-based Pretrained Embeddings for Performance Optimizations
by: Dutta, Akash, et al.
Published: (2024)
by: Dutta, Akash, et al.
Published: (2024)
MambaNetBurst: Direct Byte-level Network Traffic Classification without Tokenization or Pretraining
by: Kulatilleke, Gayan K., et al.
Published: (2026)
by: Kulatilleke, Gayan K., et al.
Published: (2026)
Accelerating Diffusion LLMs via Adaptive Parallel Decoding
by: Israel, Daniel, et al.
Published: (2025)
by: Israel, Daniel, et al.
Published: (2025)
CPINN-ABPI: Physics-Informed Neural Networks for Accurate Power Estimation in MPSoCs
by: Elshamy, Mohamed R., et al.
Published: (2025)
by: Elshamy, Mohamed R., et al.
Published: (2025)
DistZO2: High-Throughput and Memory-Efficient Zeroth-Order Fine-tuning LLMs with Distributed Parallel Computing
by: Wang, Liangyu, et al.
Published: (2025)
by: Wang, Liangyu, et al.
Published: (2025)
PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation
by: An, Zihao, et al.
Published: (2025)
by: An, Zihao, et al.
Published: (2025)
Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants
by: You, Bozhi, et al.
Published: (2025)
by: You, Bozhi, et al.
Published: (2025)
Enhancing Tropical Cyclone Path Forecasting with an Improved Transformer Network
by: Van Thanh, Nguyen, et al.
Published: (2025)
by: Van Thanh, Nguyen, et al.
Published: (2025)
lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
by: Wang, Haoxin, et al.
Published: (2025)
by: Wang, Haoxin, et al.
Published: (2025)
WCDT: Systematic WCET Optimization for Decision Tree Implementations
by: Hölscher, Nils, et al.
Published: (2025)
by: Hölscher, Nils, et al.
Published: (2025)
MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025)
by: Chitty-Venkata, Krishna Teja, et al.
Published: (2025)
MLPerf Automotive
by: Shojaei, Radoyeh, et al.
Published: (2025)
by: Shojaei, Radoyeh, et al.
Published: (2025)
Accuracy and Consumption analysis from a compressed model by CompactifAI from Multiverse Computing
by: Fovet, Damien, et al.
Published: (2025)
by: Fovet, Damien, et al.
Published: (2025)
V-Seek: Accelerating LLM Reasoning on Open-hardware Server-class RISC-V Platforms
by: Rodrigo, Javier J. Poveda, et al.
Published: (2025)
by: Rodrigo, Javier J. Poveda, et al.
Published: (2025)
PrETi: Predicting Execution Time in Early Stage with LLVM and Machine Learning
by: Xu, Risheng, et al.
Published: (2025)
by: Xu, Risheng, et al.
Published: (2025)
ZO2: Scalable Zeroth-Order Fine-Tuning for Extremely Large Language Models with Limited GPU Memory
by: Wang, Liangyu, et al.
Published: (2025)
by: Wang, Liangyu, et al.
Published: (2025)
Cloud Computing Energy Consumption Prediction Based on Kernel Extreme Learning Machine Algorithm Improved by Vector Weighted Average Algorithm
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
Similar Items
-
Alignment-Aware Decoding
by: Berdoz, Frédéric, et al.
Published: (2025) -
An Interpretable Latency Model for Speculative Decoding in LLM Serving
by: Kong, Linghao, et al.
Published: (2026) -
Light Differentiable Logic Gate Networks
by: Rüttgers, Lukas, et al.
Published: (2025) -
Mind the Gap: Removing the Discretization Gap in Differentiable Logic Gate Networks
by: Yousefi, Shakir, et al.
Published: (2025) -
Can AI Agents Agree?
by: Berdoz, Frédéric, et al.
Published: (2026)