Rethinking Token Reduction for State Space Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhan, Zheng, Wu, Yushu, Kong, Zhenglun, Yang, Changdi, Gong, Yifan, Shen, Xuan, Lin, Xue, Zhao, Pu, Wang, Yanzhi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909354409590784
author Zhan, Zheng
Wu, Yushu
Kong, Zhenglun
Yang, Changdi
Gong, Yifan
Shen, Xuan
Lin, Xue
Zhao, Pu
Wang, Yanzhi
author_facet Zhan, Zheng
Wu, Yushu
Kong, Zhenglun
Yang, Changdi
Gong, Yifan
Shen, Xuan
Lin, Xue
Zhao, Pu
Wang, Yanzhi
contents Recent advancements in State Space Models (SSMs) have attracted significant interest, particularly in models optimized for parallel training and handling long-range dependencies. Architectures like Mamba have scaled to billions of parameters with selective SSM. To facilitate broader applications using Mamba, exploring its efficiency is crucial. While token reduction techniques offer a straightforward post-training strategy, we find that applying existing methods directly to SSMs leads to substantial performance drops. Through insightful analysis, we identify the reasons for this failure and the limitations of current techniques. In response, we propose a tailored, unified post-training token reduction method for SSMs. Our approach integrates token importance and similarity, thus taking advantage of both pruning and merging, to devise a fine-grained intra-layer token reduction strategy. Extensive experiments show that our method improves the average accuracy by 5.7% to 13.1% on six benchmarks with Mamba-2 compared to existing methods, while significantly reducing computational demands and memory requirements.
format Preprint
id arxiv_https___arxiv_org_abs_2410_14725
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Rethinking Token Reduction for State Space Models
Zhan, Zheng
Wu, Yushu
Kong, Zhenglun
Yang, Changdi
Gong, Yifan
Shen, Xuan
Lin, Xue
Zhao, Pu
Wang, Yanzhi
Machine Learning
Computation and Language
Recent advancements in State Space Models (SSMs) have attracted significant interest, particularly in models optimized for parallel training and handling long-range dependencies. Architectures like Mamba have scaled to billions of parameters with selective SSM. To facilitate broader applications using Mamba, exploring its efficiency is crucial. While token reduction techniques offer a straightforward post-training strategy, we find that applying existing methods directly to SSMs leads to substantial performance drops. Through insightful analysis, we identify the reasons for this failure and the limitations of current techniques. In response, we propose a tailored, unified post-training token reduction method for SSMs. Our approach integrates token importance and similarity, thus taking advantage of both pruning and merging, to devise a fine-grained intra-layer token reduction strategy. Extensive experiments show that our method improves the average accuracy by 5.7% to 13.1% on six benchmarks with Mamba-2 compared to existing methods, while significantly reducing computational demands and memory requirements.
title Rethinking Token Reduction for State Space Models
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2410.14725