LightMamba: Efficient Mamba Acceleration on FPGA with Quantization and Hardware Co-design
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wei, Renjie, Xu, Songqiang, Zhong, Linfeng, Yang, Zebin, Guo, Qingyu, Wang, Yuan, Wang, Runsheng, Li, Meng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SpecMamba: Accelerating Mamba Inference on FPGA with Speculative Decoding
von: Zhong, Linfeng, et al.
Veröffentlicht: (2025)
von: Zhong, Linfeng, et al.
Veröffentlicht: (2025)
FPQVAR: Floating Point Quantization for Visual Autoregressive Model with FPGA Hardware Co-design
von: Wei, Renjie, et al.
Veröffentlicht: (2025)
von: Wei, Renjie, et al.
Veröffentlicht: (2025)
FastMamba: A High-Speed and Efficient Mamba Accelerator on FPGA with Accurate Quantization
von: Wang, Aotao, et al.
Veröffentlicht: (2025)
von: Wang, Aotao, et al.
Veröffentlicht: (2025)
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
von: Xu, Zukang, et al.
Veröffentlicht: (2025)
von: Xu, Zukang, et al.
Veröffentlicht: (2025)
ProPD: Dynamic Token Tree Pruning and Generation for LLM Parallel Decoding
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2024)
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2024)
HG-PIPE: Vision Transformer Acceleration with Hybrid-Grained Pipeline
von: Guo, Qingyu, et al.
Veröffentlicht: (2024)
von: Guo, Qingyu, et al.
Veröffentlicht: (2024)
ReMamba: Equip Mamba with Effective Long-Sequence Modeling
von: Yuan, Danlong, et al.
Veröffentlicht: (2024)
von: Yuan, Danlong, et al.
Veröffentlicht: (2024)
SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding
von: Wei, Linye, et al.
Veröffentlicht: (2025)
von: Wei, Linye, et al.
Veröffentlicht: (2025)
DenseMamba: State Space Models with Dense Hidden Connection for Efficient Large Language Models
von: He, Wei, et al.
Veröffentlicht: (2024)
von: He, Wei, et al.
Veröffentlicht: (2024)
ML-Mamba: Efficient Multi-Modal Large Language Model Utilizing Mamba-2
von: Huang, Wenjun, et al.
Veröffentlicht: (2024)
von: Huang, Wenjun, et al.
Veröffentlicht: (2024)
RankMamba: Benchmarking Mamba's Document Ranking Performance in the Era of Transformers
von: Xu, Zhichao
Veröffentlicht: (2024)
von: Xu, Zhichao
Veröffentlicht: (2024)
MambaPEFT: Exploring Parameter-Efficient Fine-Tuning for Mamba
von: Yoshimura, Masakazu, et al.
Veröffentlicht: (2024)
von: Yoshimura, Masakazu, et al.
Veröffentlicht: (2024)
BWTA: Accurate and Efficient Binarized Transformer by Algorithm-Hardware Co-design
von: Ding, Yifu, et al.
Veröffentlicht: (2026)
von: Ding, Yifu, et al.
Veröffentlicht: (2026)
PrivQuant: Communication-Efficient Private Inference with Quantized Network/Protocol Co-Optimization
von: Xu, Tianshi, et al.
Veröffentlicht: (2024)
von: Xu, Tianshi, et al.
Veröffentlicht: (2024)
ViM-Q: Scalable Algorithm-Hardware Co-Design for Vision Mamba Model Inference on FPGA
von: Lyu, Shengzhe, et al.
Veröffentlicht: (2026)
von: Lyu, Shengzhe, et al.
Veröffentlicht: (2026)
Exploring the Limitations of Mamba in COPY and CoT Reasoning
von: Ren, Ruifeng, et al.
Veröffentlicht: (2024)
von: Ren, Ruifeng, et al.
Veröffentlicht: (2024)
DuplexMamba: Enhancing Real-time Speech Conversations with Duplex and Streaming Capabilities
von: Lu, Xiangyu, et al.
Veröffentlicht: (2025)
von: Lu, Xiangyu, et al.
Veröffentlicht: (2025)
Orchestrating Dual-Boundaries: An Arithmetic Intensity Inspired Acceleration Framework for Diffusion Language Models
von: Wei, Linye, et al.
Veröffentlicht: (2025)
von: Wei, Linye, et al.
Veröffentlicht: (2025)
Design and Implementation of an FPGA-Based Hardware Accelerator for Transformer
von: Li, Richie, et al.
Veröffentlicht: (2025)
von: Li, Richie, et al.
Veröffentlicht: (2025)
Differential Mamba
von: Schneider, Nadav, et al.
Veröffentlicht: (2025)
von: Schneider, Nadav, et al.
Veröffentlicht: (2025)
An Optimizing Framework on MLIR for Efficient FPGA-based Accelerator Generation
von: Zhang, Weichuang, et al.
Veröffentlicht: (2024)
von: Zhang, Weichuang, et al.
Veröffentlicht: (2024)
EQO: Exploring Ultra-Efficient Private Inference with Winograd-Based Protocol and Quantization Co-Optimization
von: Zeng, Wenxuan, et al.
Veröffentlicht: (2024)
von: Zeng, Wenxuan, et al.
Veröffentlicht: (2024)
ASCEND: Accurate yet Efficient End-to-End Stochastic Computing Acceleration of Vision Transformer
von: Xie, Tong, et al.
Veröffentlicht: (2024)
von: Xie, Tong, et al.
Veröffentlicht: (2024)
LongMamba: Enhancing Mamba's Long Context Capabilities via Training-Free Receptive Field Enlargement
von: Ye, Zhifan, et al.
Veröffentlicht: (2025)
von: Ye, Zhifan, et al.
Veröffentlicht: (2025)
GeneMamba: An Efficient and Effective Foundation Model on Single Cell Data
von: Qi, Cong, et al.
Veröffentlicht: (2025)
von: Qi, Cong, et al.
Veröffentlicht: (2025)
Tiny Recursive Reasoning with Mamba-2 Attention Hybrid
von: Wang, Wenlong, et al.
Veröffentlicht: (2026)
von: Wang, Wenlong, et al.
Veröffentlicht: (2026)
HEQuant: Marrying Homomorphic Encryption and Quantization for Communication-Efficient Private Inference
von: Xu, Tianshi, et al.
Veröffentlicht: (2024)
von: Xu, Tianshi, et al.
Veröffentlicht: (2024)
Cement2: Temporal Hardware Transactions for High-Level and Efficient FPGA Programming
von: Xiao, Youwei, et al.
Veröffentlicht: (2025)
von: Xiao, Youwei, et al.
Veröffentlicht: (2025)
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
Q-MambaIR: Accurate Quantized Mamba for Efficient Image Restoration
von: Chen, Yujie, et al.
Veröffentlicht: (2025)
von: Chen, Yujie, et al.
Veröffentlicht: (2025)
An Exploration of Mamba for Speech Self-Supervised Models
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2025)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2025)
BioMamba: Domain-Adaptive Biomedical Language Models
von: Yue, Ling, et al.
Veröffentlicht: (2024)
von: Yue, Ling, et al.
Veröffentlicht: (2024)
Mamba Drafters for Speculative Decoding
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
Locating and Editing Factual Associations in Mamba
von: Sharma, Arnab Sen, et al.
Veröffentlicht: (2024)
von: Sharma, Arnab Sen, et al.
Veröffentlicht: (2024)
HADES: Hardware Accelerated Decoding for Efficient Speculation in Large Language Models
von: Yang, Ze, et al.
Veröffentlicht: (2024)
von: Yang, Ze, et al.
Veröffentlicht: (2024)
MambaScope: Coarse-to-Fine Scoping for Efficient Vision Mamba
von: Liu, Shanhui, et al.
Veröffentlicht: (2025)
von: Liu, Shanhui, et al.
Veröffentlicht: (2025)
DocMamba: Efficient Document Pre-training with State Space Model
von: Hu, Pengfei, et al.
Veröffentlicht: (2024)
von: Hu, Pengfei, et al.
Veröffentlicht: (2024)
Mamba-Shedder: Post-Transformer Compression for Efficient Selective Structured State Space Models
von: Muñoz, J. Pablo, et al.
Veröffentlicht: (2025)
von: Muñoz, J. Pablo, et al.
Veröffentlicht: (2025)
MambaByte: Token-free Selective State Space Model
von: Wang, Junxiong, et al.
Veröffentlicht: (2024)
von: Wang, Junxiong, et al.
Veröffentlicht: (2024)
Efficient yet Accurate End-to-End SC Accelerator Design
von: Li, Meng, et al.
Veröffentlicht: (2024)
von: Li, Meng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SpecMamba: Accelerating Mamba Inference on FPGA with Speculative Decoding
von: Zhong, Linfeng, et al.
Veröffentlicht: (2025) -
FPQVAR: Floating Point Quantization for Visual Autoregressive Model with FPGA Hardware Co-design
von: Wei, Renjie, et al.
Veröffentlicht: (2025) -
FastMamba: A High-Speed and Efficient Mamba Accelerator on FPGA with Accurate Quantization
von: Wang, Aotao, et al.
Veröffentlicht: (2025) -
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
von: Xu, Zukang, et al.
Veröffentlicht: (2025) -
ProPD: Dynamic Token Tree Pruning and Generation for LLM Parallel Decoding
von: Zhong, Shuzhang, et al.
Veröffentlicht: (2024)