M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Junxiong, Li, Wen-Ding, Paliotta, Daniele, Ritter, Daniel, Rush, Alexander M., Dao, Tri
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911143596916736
author Wang, Junxiong
Li, Wen-Ding
Paliotta, Daniele
Ritter, Daniel
Rush, Alexander M.
Dao, Tri
author_facet Wang, Junxiong
Li, Wen-Ding
Paliotta, Daniele
Ritter, Daniel
Rush, Alexander M.
Dao, Tri
contents Effective reasoning is crucial to solving complex mathematical problems. Recent large language models (LLMs) have boosted performance by scaling test-time computation through long chain-of-thought reasoning. However, transformer-based models are inherently limited in extending context length due to their quadratic computational complexity and linear memory requirements. In this paper, we introduce a novel hybrid linear RNN reasoning model, M1, built on the Mamba architecture, which allows memory-efficient inference. Our approach leverages a distillation process from existing reasoning models and is further enhanced through RL training. Experimental results on the AIME and MATH benchmarks show that M1 not only outperforms previous linear RNN models but also matches the performance of state-of-the-art Deepseek R1 distilled reasoning models at a similar scale. We also compare our generation speed with a highly performant general purpose inference engine, vLLM, and observe more than a 3x speedup compared to a same size transformer. With throughput speedup, we are able to achieve higher accuracy compared to DeepSeek R1 distilled transformer reasoning models under a fixed generation time budget using self-consistency voting. Overall, we introduce a hybrid Mamba reasoning model and provide a more effective approach to scaling test-time generation using self-consistency or long chain of thought reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2504_10449
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models
Wang, Junxiong
Li, Wen-Ding
Paliotta, Daniele
Ritter, Daniel
Rush, Alexander M.
Dao, Tri
Machine Learning
Effective reasoning is crucial to solving complex mathematical problems. Recent large language models (LLMs) have boosted performance by scaling test-time computation through long chain-of-thought reasoning. However, transformer-based models are inherently limited in extending context length due to their quadratic computational complexity and linear memory requirements. In this paper, we introduce a novel hybrid linear RNN reasoning model, M1, built on the Mamba architecture, which allows memory-efficient inference. Our approach leverages a distillation process from existing reasoning models and is further enhanced through RL training. Experimental results on the AIME and MATH benchmarks show that M1 not only outperforms previous linear RNN models but also matches the performance of state-of-the-art Deepseek R1 distilled reasoning models at a similar scale. We also compare our generation speed with a highly performant general purpose inference engine, vLLM, and observe more than a 3x speedup compared to a same size transformer. With throughput speedup, we are able to achieve higher accuracy compared to DeepSeek R1 distilled transformer reasoning models under a fixed generation time budget using self-consistency voting. Overall, we introduce a hybrid Mamba reasoning model and provide a more effective approach to scaling test-time generation using self-consistency or long chain of thought reasoning.
title M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models
topic Machine Learning
url https://arxiv.org/abs/2504.10449