LightMamba: Efficient Mamba Acceleration on FPGA with Quantization and Hardware Co-design
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Renjie, Xu, Songqiang, Zhong, Linfeng, Yang, Zebin, Guo, Qingyu, Wang, Yuan, Wang, Runsheng, Li, Meng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SpecMamba: Accelerating Mamba Inference on FPGA with Speculative Decoding
by: Zhong, Linfeng, et al.
Published: (2025)
by: Zhong, Linfeng, et al.
Published: (2025)
FPQVAR: Floating Point Quantization for Visual Autoregressive Model with FPGA Hardware Co-design
by: Wei, Renjie, et al.
Published: (2025)
by: Wei, Renjie, et al.
Published: (2025)
FastMamba: A High-Speed and Efficient Mamba Accelerator on FPGA with Accurate Quantization
by: Wang, Aotao, et al.
Published: (2025)
by: Wang, Aotao, et al.
Published: (2025)
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
by: Xu, Zukang, et al.
Published: (2025)
by: Xu, Zukang, et al.
Published: (2025)
ProPD: Dynamic Token Tree Pruning and Generation for LLM Parallel Decoding
by: Zhong, Shuzhang, et al.
Published: (2024)
by: Zhong, Shuzhang, et al.
Published: (2024)
HG-PIPE: Vision Transformer Acceleration with Hybrid-Grained Pipeline
by: Guo, Qingyu, et al.
Published: (2024)
by: Guo, Qingyu, et al.
Published: (2024)
ReMamba: Equip Mamba with Effective Long-Sequence Modeling
by: Yuan, Danlong, et al.
Published: (2024)
by: Yuan, Danlong, et al.
Published: (2024)
SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding
by: Wei, Linye, et al.
Published: (2025)
by: Wei, Linye, et al.
Published: (2025)
DenseMamba: State Space Models with Dense Hidden Connection for Efficient Large Language Models
by: He, Wei, et al.
Published: (2024)
by: He, Wei, et al.
Published: (2024)
ML-Mamba: Efficient Multi-Modal Large Language Model Utilizing Mamba-2
by: Huang, Wenjun, et al.
Published: (2024)
by: Huang, Wenjun, et al.
Published: (2024)
RankMamba: Benchmarking Mamba's Document Ranking Performance in the Era of Transformers
by: Xu, Zhichao
Published: (2024)
by: Xu, Zhichao
Published: (2024)
MambaPEFT: Exploring Parameter-Efficient Fine-Tuning for Mamba
by: Yoshimura, Masakazu, et al.
Published: (2024)
by: Yoshimura, Masakazu, et al.
Published: (2024)
BWTA: Accurate and Efficient Binarized Transformer by Algorithm-Hardware Co-design
by: Ding, Yifu, et al.
Published: (2026)
by: Ding, Yifu, et al.
Published: (2026)
PrivQuant: Communication-Efficient Private Inference with Quantized Network/Protocol Co-Optimization
by: Xu, Tianshi, et al.
Published: (2024)
by: Xu, Tianshi, et al.
Published: (2024)
ViM-Q: Scalable Algorithm-Hardware Co-Design for Vision Mamba Model Inference on FPGA
by: Lyu, Shengzhe, et al.
Published: (2026)
by: Lyu, Shengzhe, et al.
Published: (2026)
Exploring the Limitations of Mamba in COPY and CoT Reasoning
by: Ren, Ruifeng, et al.
Published: (2024)
by: Ren, Ruifeng, et al.
Published: (2024)
DuplexMamba: Enhancing Real-time Speech Conversations with Duplex and Streaming Capabilities
by: Lu, Xiangyu, et al.
Published: (2025)
by: Lu, Xiangyu, et al.
Published: (2025)
Orchestrating Dual-Boundaries: An Arithmetic Intensity Inspired Acceleration Framework for Diffusion Language Models
by: Wei, Linye, et al.
Published: (2025)
by: Wei, Linye, et al.
Published: (2025)
Design and Implementation of an FPGA-Based Hardware Accelerator for Transformer
by: Li, Richie, et al.
Published: (2025)
by: Li, Richie, et al.
Published: (2025)
Differential Mamba
by: Schneider, Nadav, et al.
Published: (2025)
by: Schneider, Nadav, et al.
Published: (2025)
An Optimizing Framework on MLIR for Efficient FPGA-based Accelerator Generation
by: Zhang, Weichuang, et al.
Published: (2024)
by: Zhang, Weichuang, et al.
Published: (2024)
EQO: Exploring Ultra-Efficient Private Inference with Winograd-Based Protocol and Quantization Co-Optimization
by: Zeng, Wenxuan, et al.
Published: (2024)
by: Zeng, Wenxuan, et al.
Published: (2024)
ASCEND: Accurate yet Efficient End-to-End Stochastic Computing Acceleration of Vision Transformer
by: Xie, Tong, et al.
Published: (2024)
by: Xie, Tong, et al.
Published: (2024)
LongMamba: Enhancing Mamba's Long Context Capabilities via Training-Free Receptive Field Enlargement
by: Ye, Zhifan, et al.
Published: (2025)
by: Ye, Zhifan, et al.
Published: (2025)
GeneMamba: An Efficient and Effective Foundation Model on Single Cell Data
by: Qi, Cong, et al.
Published: (2025)
by: Qi, Cong, et al.
Published: (2025)
Tiny Recursive Reasoning with Mamba-2 Attention Hybrid
by: Wang, Wenlong, et al.
Published: (2026)
by: Wang, Wenlong, et al.
Published: (2026)
HEQuant: Marrying Homomorphic Encryption and Quantization for Communication-Efficient Private Inference
by: Xu, Tianshi, et al.
Published: (2024)
by: Xu, Tianshi, et al.
Published: (2024)
Cement2: Temporal Hardware Transactions for High-Level and Efficient FPGA Programming
by: Xiao, Youwei, et al.
Published: (2025)
by: Xiao, Youwei, et al.
Published: (2025)
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
Q-MambaIR: Accurate Quantized Mamba for Efficient Image Restoration
by: Chen, Yujie, et al.
Published: (2025)
by: Chen, Yujie, et al.
Published: (2025)
An Exploration of Mamba for Speech Self-Supervised Models
by: Lin, Tzu-Quan, et al.
Published: (2025)
by: Lin, Tzu-Quan, et al.
Published: (2025)
BioMamba: Domain-Adaptive Biomedical Language Models
by: Yue, Ling, et al.
Published: (2024)
by: Yue, Ling, et al.
Published: (2024)
Mamba Drafters for Speculative Decoding
by: Choi, Daewon, et al.
Published: (2025)
by: Choi, Daewon, et al.
Published: (2025)
Locating and Editing Factual Associations in Mamba
by: Sharma, Arnab Sen, et al.
Published: (2024)
by: Sharma, Arnab Sen, et al.
Published: (2024)
HADES: Hardware Accelerated Decoding for Efficient Speculation in Large Language Models
by: Yang, Ze, et al.
Published: (2024)
by: Yang, Ze, et al.
Published: (2024)
MambaScope: Coarse-to-Fine Scoping for Efficient Vision Mamba
by: Liu, Shanhui, et al.
Published: (2025)
by: Liu, Shanhui, et al.
Published: (2025)
DocMamba: Efficient Document Pre-training with State Space Model
by: Hu, Pengfei, et al.
Published: (2024)
by: Hu, Pengfei, et al.
Published: (2024)
Mamba-Shedder: Post-Transformer Compression for Efficient Selective Structured State Space Models
by: Muñoz, J. Pablo, et al.
Published: (2025)
by: Muñoz, J. Pablo, et al.
Published: (2025)
MambaByte: Token-free Selective State Space Model
by: Wang, Junxiong, et al.
Published: (2024)
by: Wang, Junxiong, et al.
Published: (2024)
Efficient yet Accurate End-to-End SC Accelerator Design
by: Li, Meng, et al.
Published: (2024)
by: Li, Meng, et al.
Published: (2024)
Similar Items
-
SpecMamba: Accelerating Mamba Inference on FPGA with Speculative Decoding
by: Zhong, Linfeng, et al.
Published: (2025) -
FPQVAR: Floating Point Quantization for Visual Autoregressive Model with FPGA Hardware Co-design
by: Wei, Renjie, et al.
Published: (2025) -
FastMamba: A High-Speed and Efficient Mamba Accelerator on FPGA with Accurate Quantization
by: Wang, Aotao, et al.
Published: (2025) -
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
by: Xu, Zukang, et al.
Published: (2025) -
ProPD: Dynamic Token Tree Pruning and Generation for LLM Parallel Decoding
by: Zhong, Shuzhang, et al.
Published: (2024)