Mamba Drafters for Speculative Decoding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Choi, Daewon, Oh, Seunghyuk, Dingliwal, Saket, Tack, Jihoon, Kim, Kyuyoung, Song, Woomin, Kim, Seojin, Han, Insu, Shin, Jinwoo, Galstyan, Aram, Katiyar, Shubham, Bodapati, Sravan Babu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918042815954944
author Choi, Daewon
Oh, Seunghyuk
Dingliwal, Saket
Tack, Jihoon
Kim, Kyuyoung
Song, Woomin
Kim, Seojin
Han, Insu
Shin, Jinwoo
Galstyan, Aram
Katiyar, Shubham
Bodapati, Sravan Babu
author_facet Choi, Daewon
Oh, Seunghyuk
Dingliwal, Saket
Tack, Jihoon
Kim, Kyuyoung
Song, Woomin
Kim, Seojin
Han, Insu
Shin, Jinwoo
Galstyan, Aram
Katiyar, Shubham
Bodapati, Sravan Babu
contents Speculative decoding has emerged as a promising approach to accelerating large language model (LLM) generation using a fast drafter while maintaining alignment with the target model's distribution. However, existing approaches face a trade-off: external drafters offer flexibility but can suffer from slower drafting, while self-speculation methods use drafters tailored to the target model but require re-training. In this paper, we introduce novel drafters based on Mamba, a state-of-the-art state space model (SSM), as a solution that combines the best aspects of both approaches. By leveraging the linear structure of SSMs, our approach avoids the quadratic complexity inherent in traditional Transformer-based methods, enabling faster drafting and lower memory usage while maintaining the flexibility to work across different target models. We further enhance efficiency with a novel test-time tree search algorithm for generating high-quality draft candidates. Our empirical evaluation demonstrates that Mamba-based drafters not only outperform existing external drafting methods but are also comparable to state-of-the-art self-speculation approaches while using less memory and maintaining their cross-model adaptability.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01206
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mamba Drafters for Speculative Decoding
Choi, Daewon
Oh, Seunghyuk
Dingliwal, Saket
Tack, Jihoon
Kim, Kyuyoung
Song, Woomin
Kim, Seojin
Han, Insu
Shin, Jinwoo
Galstyan, Aram
Katiyar, Shubham
Bodapati, Sravan Babu
Computation and Language
Artificial Intelligence
Speculative decoding has emerged as a promising approach to accelerating large language model (LLM) generation using a fast drafter while maintaining alignment with the target model's distribution. However, existing approaches face a trade-off: external drafters offer flexibility but can suffer from slower drafting, while self-speculation methods use drafters tailored to the target model but require re-training. In this paper, we introduce novel drafters based on Mamba, a state-of-the-art state space model (SSM), as a solution that combines the best aspects of both approaches. By leveraging the linear structure of SSMs, our approach avoids the quadratic complexity inherent in traditional Transformer-based methods, enabling faster drafting and lower memory usage while maintaining the flexibility to work across different target models. We further enhance efficiency with a novel test-time tree search algorithm for generating high-quality draft candidates. Our empirical evaluation demonstrates that Mamba-based drafters not only outperform existing external drafting methods but are also comparable to state-of-the-art self-speculation approaches while using less memory and maintaining their cross-model adaptability.
title Mamba Drafters for Speculative Decoding
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.01206