Mamba Drafters for Speculative Decoding
Fuente:
arXiv
Saved in:
| Main Authors: | Choi, Daewon, Oh, Seunghyuk, Dingliwal, Saket, Tack, Jihoon, Kim, Kyuyoung, Song, Woomin, Kim, Seojin, Han, Insu, Shin, Jinwoo, Galstyan, Aram, Katiyar, Shubham, Bodapati, Sravan Babu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accelerated Test-Time Scaling with Model-Free Speculative Sampling
by: Song, Woomin, et al.
Published: (2025)
by: Song, Woomin, et al.
Published: (2025)
Think Clearly: Improving Reasoning via Redundant Token Pruning
by: Choi, Daewon, et al.
Published: (2025)
by: Choi, Daewon, et al.
Published: (2025)
IdleSpec: Exploiting Idle Time via Speculative Planning for LLM Agents
by: Choi, Daewon, et al.
Published: (2026)
by: Choi, Daewon, et al.
Published: (2026)
ExComm: Exploration-Stage Communication for Error-Resilient Agentic Test-Time Scaling
by: Song, Woomin, et al.
Published: (2026)
by: Song, Woomin, et al.
Published: (2026)
Compress, Gather, and Recompute: REFORMing Long-Context Processing in Transformers
by: Song, Woomin, et al.
Published: (2025)
by: Song, Woomin, et al.
Published: (2025)
Optimized Feature Generation for Tabular Data via LLMs with Decision Tree Reasoning
by: Nam, Jaehyun, et al.
Published: (2024)
by: Nam, Jaehyun, et al.
Published: (2024)
Sparsified State-Space Models are Efficient Highway Networks
by: Song, Woomin, et al.
Published: (2025)
by: Song, Woomin, et al.
Published: (2025)
SeRA: Self-Reviewing and Alignment of Large Language Models using Implicit Reward Margins
by: Ko, Jongwoo, et al.
Published: (2024)
by: Ko, Jongwoo, et al.
Published: (2024)
ReVISE: Learning to Refine at Test-Time via Intrinsic Self-Verification
by: Lee, Hyunseok, et al.
Published: (2025)
by: Lee, Hyunseok, et al.
Published: (2025)
Scalable and Robust LLM Unlearning by Correcting Responses with Retrieved Exclusions
by: Kim, Junbeom, et al.
Published: (2025)
by: Kim, Junbeom, et al.
Published: (2025)
Adaptive Video Understanding Agent: Enhancing efficiency with dynamic frame sampling and feedback-driven reasoning
by: Jeoung, Sullam, et al.
Published: (2024)
by: Jeoung, Sullam, et al.
Published: (2024)
Multi-Drafter Speculative Decoding with Alignment Feedback
by: Kim, Taehyeon, et al.
Published: (2026)
by: Kim, Taehyeon, et al.
Published: (2026)
Personalized Language Models via Privacy-Preserving Evolutionary Model Merging
by: Kim, Kyuyoung, et al.
Published: (2025)
by: Kim, Kyuyoung, et al.
Published: (2025)
Learning to Contextualize Web Pages for Enhanced Decision Making by LLM Agents
by: Lee, Dongjun, et al.
Published: (2025)
by: Lee, Dongjun, et al.
Published: (2025)
Hierarchical Context Merging: Better Long Context Understanding for Pre-trained LLMs
by: Song, Woomin, et al.
Published: (2024)
by: Song, Woomin, et al.
Published: (2024)
Tabular Transfer Learning via Prompting LLMs
by: Nam, Jaehyun, et al.
Published: (2024)
by: Nam, Jaehyun, et al.
Published: (2024)
ReMoDetect: Reward Models Recognize Aligned LLM's Generations
by: Lee, Hyunseok, et al.
Published: (2024)
by: Lee, Hyunseok, et al.
Published: (2024)
RedacBench: Can AI Erase Your Secrets?
by: Jeon, Hyunjun, et al.
Published: (2026)
by: Jeon, Hyunjun, et al.
Published: (2026)
Self-Refining Language Model Anonymizers via Adversarial Distillation
by: Kim, Kyuyoung, et al.
Published: (2025)
by: Kim, Kyuyoung, et al.
Published: (2025)
Towards Fast Multilingual LLM Inference: Speculative Decoding and Specialized Drafters
by: Yi, Euiin, et al.
Published: (2024)
by: Yi, Euiin, et al.
Published: (2024)
Steering Pretrained Drafters during Speculative Decoding
by: Berdoz, Frédéric, et al.
Published: (2025)
by: Berdoz, Frédéric, et al.
Published: (2025)
Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding
by: Park, Jihoon, et al.
Published: (2025)
by: Park, Jihoon, et al.
Published: (2025)
Coupling without Communication and Drafter-Invariant Speculative Decoding
by: Daliri, Majid, et al.
Published: (2024)
by: Daliri, Majid, et al.
Published: (2024)
ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
by: Xiao, Zilin, et al.
Published: (2024)
by: Xiao, Zilin, et al.
Published: (2024)
Beyond Correctness: Learning Robust Reasoning via Transfer
by: Lee, Hyunseok, et al.
Published: (2026)
by: Lee, Hyunseok, et al.
Published: (2026)
Graph Signal Processing Meets Mamba2: Adaptive Filter Bank via Delta Modulation
by: Shin, Yehjin, et al.
Published: (2026)
by: Shin, Yehjin, et al.
Published: (2026)
Recurrent Drafter for Fast Speculative Decoding in Large Language Models
by: Cheng, Yunfei, et al.
Published: (2024)
by: Cheng, Yunfei, et al.
Published: (2024)
Confidence-aware Denoised Fine-tuning of Off-the-shelf Models for Certified Robustness
by: Jang, Suhyeok, et al.
Published: (2024)
by: Jang, Suhyeok, et al.
Published: (2024)
Training Text-to-Molecule Models with Context-Aware Tokenization
by: Kim, Seojin, et al.
Published: (2025)
by: Kim, Seojin, et al.
Published: (2025)
Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMs
by: Liu, Hongyi, et al.
Published: (2025)
by: Liu, Hongyi, et al.
Published: (2025)
Data-Efficient Molecular Generation with Hierarchical Textual Inversion
by: Kim, Seojin, et al.
Published: (2024)
by: Kim, Seojin, et al.
Published: (2024)
SpecDiff-2: Scaling Diffusion Drafter Alignment For Faster Speculative Decoding
by: Sandler, Jameson, et al.
Published: (2025)
by: Sandler, Jameson, et al.
Published: (2025)
Online Adaptation of Language Models with a Memory of Amortized Contexts
by: Tack, Jihoon, et al.
Published: (2024)
by: Tack, Jihoon, et al.
Published: (2024)
OmniDraft: A Cross-vocabulary, Online Adaptive Drafter for On-device Speculative Decoding
by: Ramakrishnan, Ramchalam Kinattinkara, et al.
Published: (2025)
by: Ramakrishnan, Ramchalam Kinattinkara, et al.
Published: (2025)
KnapSpec: Self-Speculative Decoding via Adaptive Layer Selection as a Knapsack Problem
by: Cha, Seongjin, et al.
Published: (2026)
by: Cha, Seongjin, et al.
Published: (2026)
Margin Matching Preference Optimization: Enhanced Model Alignment with Granular Feedback
by: Kim, Kyuyoung, et al.
Published: (2024)
by: Kim, Kyuyoung, et al.
Published: (2024)
FontAdapter: Instant Font Adaptation in Visual Text Generation
by: Koo, Myungkyu, et al.
Published: (2025)
by: Koo, Myungkyu, et al.
Published: (2025)
HiViS: Hiding Visual Tokens from the Drafter for Speculative Decoding in Vision-Language Models
by: Xie, Zhinan, et al.
Published: (2025)
by: Xie, Zhinan, et al.
Published: (2025)
Speculative Verification: Exploiting Information Gain to Refine Speculative Decoding
by: Kim, Sungkyun, et al.
Published: (2025)
by: Kim, Sungkyun, et al.
Published: (2025)
Context Length Alone Hurts LLM Performance Despite Perfect Retrieval
by: Du, Yufeng, et al.
Published: (2025)
by: Du, Yufeng, et al.
Published: (2025)
Similar Items
-
Accelerated Test-Time Scaling with Model-Free Speculative Sampling
by: Song, Woomin, et al.
Published: (2025) -
Think Clearly: Improving Reasoning via Redundant Token Pruning
by: Choi, Daewon, et al.
Published: (2025) -
IdleSpec: Exploiting Idle Time via Speculative Planning for LLM Agents
by: Choi, Daewon, et al.
Published: (2026) -
ExComm: Exploration-Stage Communication for Error-Resilient Agentic Test-Time Scaling
by: Song, Woomin, et al.
Published: (2026) -
Compress, Gather, and Recompute: REFORMing Long-Context Processing in Transformers
by: Song, Woomin, et al.
Published: (2025)