FastMamba: A High-Speed and Efficient Mamba Accelerator on FPGA with Accurate Quantization
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Aotao, Shao, Haikuo, Ma, Shaobo, Wang, Zhongfeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Arbitrary Precision Acceleration for Large Language Models on GPU Tensor Cores
by: Ma, Shaobo, et al.
Published: (2024)
by: Ma, Shaobo, et al.
Published: (2024)
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration
by: Ma, Shaobo, et al.
Published: (2025)
by: Ma, Shaobo, et al.
Published: (2025)
An FPGA-Based Reconfigurable Accelerator for Convolution-Transformer Hybrid EfficientViT
by: Shao, Haikuo, et al.
Published: (2024)
by: Shao, Haikuo, et al.
Published: (2024)
AccLLM: Accelerating Long-Context LLM Inference Via Algorithm-Hardware Co-Design
by: Liang, Yanbiao, et al.
Published: (2025)
by: Liang, Yanbiao, et al.
Published: (2025)
M$^2$-ViT: Accelerating Hybrid Vision Transformers with Two-Level Mixed Quantization
by: Liang, Yanbiao, et al.
Published: (2024)
by: Liang, Yanbiao, et al.
Published: (2024)
BETA: Binarized Energy-Efficient Transformer Accelerator at the Edge
by: Ji, Yuhao, et al.
Published: (2024)
by: Ji, Yuhao, et al.
Published: (2024)
SpecMamba: Accelerating Mamba Inference on FPGA with Speculative Decoding
by: Zhong, Linfeng, et al.
Published: (2025)
by: Zhong, Linfeng, et al.
Published: (2025)
MARCA: Mamba Accelerator with ReConfigurable Architecture
by: Li, Jinhao, et al.
Published: (2024)
by: Li, Jinhao, et al.
Published: (2024)
CDM-QTA: Quantized Training Acceleration for Efficient LoRA Fine-Tuning of Diffusion Model
by: Lu, Jinming, et al.
Published: (2025)
by: Lu, Jinming, et al.
Published: (2025)
Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators
by: Kim, Jiyoon, et al.
Published: (2025)
by: Kim, Jiyoon, et al.
Published: (2025)
FPGA-Based Neural Network Accelerators for Space Applications: A Survey
by: Antunes, Pedro, et al.
Published: (2025)
by: Antunes, Pedro, et al.
Published: (2025)
An FPGA-Based Accelerator Enabling Efficient Support for CNNs with Arbitrary Kernel Sizes
by: Wang, Miaoxin, et al.
Published: (2024)
by: Wang, Miaoxin, et al.
Published: (2024)
Architectural Design and Performance Analysis of FPGA based AI Accelerators: A Comprehensive Review
by: Chatterjee, Soumita, et al.
Published: (2026)
by: Chatterjee, Soumita, et al.
Published: (2026)
Embedded FPGA Acceleration of Brain-Like Neural Networks: Online Learning to Scalable Inference
by: Hafiz, Muhammad Ihsan Al, et al.
Published: (2025)
by: Hafiz, Muhammad Ihsan Al, et al.
Published: (2025)
Enable Lightweight and Precision-Scalable Posit/IEEE-754 Arithmetic in RISC-V Cores for Transprecision Computing
by: Li, Qiong, et al.
Published: (2025)
by: Li, Qiong, et al.
Published: (2025)
LLM-Driven Design Space Exploration of FPGA-based Accelerators
by: Sharma, Vinamra, et al.
Published: (2026)
by: Sharma, Vinamra, et al.
Published: (2026)
Optimizing Neural Networks with Learnable Non-Linear Activation Functions via Lookup-Based FPGA Acceleration
by: Yin, Mengyuan, et al.
Published: (2025)
by: Yin, Mengyuan, et al.
Published: (2025)
Heterogeneous SoC Integrating an Open-Source Recurrent SNN Accelerator for Neuromorphic Edge Computing on FPGA
by: Barocci, Michelangelo, et al.
Published: (2026)
by: Barocci, Michelangelo, et al.
Published: (2026)
Fast, Scalable, Energy-Efficient Non-element-wise Matrix Multiplication on FPGA
by: Zhu, Xuqi, et al.
Published: (2024)
by: Zhu, Xuqi, et al.
Published: (2024)
PAI: Fast, Accurate, and Full Benchmark Performance Projection with AI
by: Johnson, Avery, et al.
Published: (2026)
by: Johnson, Avery, et al.
Published: (2026)
Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding
by: Fan, Wang, et al.
Published: (2026)
by: Fan, Wang, et al.
Published: (2026)
Idle is the New Sleep: Configuration-Aware Alternative to Powering Off FPGA-Based DL Accelerators During Inactivity
by: Qian, Chao, et al.
Published: (2024)
by: Qian, Chao, et al.
Published: (2024)
Mamba-X: An End-to-End Vision Mamba Accelerator for Edge Computing Devices
by: Yoon, Dongho, et al.
Published: (2025)
by: Yoon, Dongho, et al.
Published: (2025)
Hardware-Efficient FPGA Implementation of Sigmoid Function Using Mixed-Radix Hyperbolic Rotation CORDIC
by: Panchal, Chintan, et al.
Published: (2026)
by: Panchal, Chintan, et al.
Published: (2026)
Accurate Block Quantization in LLMs with Outliers
by: Trukhanov, Nikita, et al.
Published: (2024)
by: Trukhanov, Nikita, et al.
Published: (2024)
SpeedLLM: An FPGA Co-design of Large Language Model Inference Accelerator
by: Wang, Peipei, et al.
Published: (2025)
by: Wang, Peipei, et al.
Published: (2025)
Bitwise Systolic Array Architecture for Runtime-Reconfigurable Multi-precision Quantized Multiplication on Hardware Accelerators
by: Liu, Yuhao, et al.
Published: (2026)
by: Liu, Yuhao, et al.
Published: (2026)
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format
by: Fang, Chao, et al.
Published: (2024)
by: Fang, Chao, et al.
Published: (2024)
SnipSnap: A Joint Compression Format and Dataflow Co-Optimization Framework for Efficient Sparse LLM Accelerator Design
by: Wu, Junyi, et al.
Published: (2025)
by: Wu, Junyi, et al.
Published: (2025)
Panacea: Novel DNN Accelerator using Accuracy-Preserving Asymmetric Quantization and Energy-Saving Bit-Slice Sparsity
by: Kam, Dongyun, et al.
Published: (2024)
by: Kam, Dongyun, et al.
Published: (2024)
KANtize: Exploring Low-bit Quantization of Kolmogorov-Arnold Networks for Efficient Inference
by: Errabii, Sohaib, et al.
Published: (2026)
by: Errabii, Sohaib, et al.
Published: (2026)
Automatic Generation of Fast and Accurate Performance Models for Deep Neural Network Accelerators
by: Lübeck, Konstantin, et al.
Published: (2024)
by: Lübeck, Konstantin, et al.
Published: (2024)
FPGA Resource-aware Structured Pruning for Real-Time Neural Networks
by: Ramhorst, Benjamin, et al.
Published: (2023)
by: Ramhorst, Benjamin, et al.
Published: (2023)
A Configurable and Efficient Memory Hierarchy for Neural Network Hardware Accelerator
by: Bause, Oliver, et al.
Published: (2024)
by: Bause, Oliver, et al.
Published: (2024)
FASQ: Flexible Accelerated Subspace Quantization for Calibration-Free LLM Compression
by: Qiao, Ye, et al.
Published: (2026)
by: Qiao, Ye, et al.
Published: (2026)
An Efficient Sparse Hardware Accelerator for Spike-Driven Transformer
by: Li, Zhengke, et al.
Published: (2025)
by: Li, Zhengke, et al.
Published: (2025)
KAN-SAs: Efficient Acceleration of Kolmogorov-Arnold Networks on Systolic Arrays
by: Errabii, Sohaib, et al.
Published: (2025)
by: Errabii, Sohaib, et al.
Published: (2025)
Carbon-Efficient 3D DNN Acceleration: Optimizing Performance and Sustainability
by: Panteleaki, Aikaterini Maria, et al.
Published: (2025)
by: Panteleaki, Aikaterini Maria, et al.
Published: (2025)
SWAT: Scalable and Efficient Window Attention-based Transformers Acceleration on FPGAs
by: Bai, Zhenyu, et al.
Published: (2024)
by: Bai, Zhenyu, et al.
Published: (2024)
TSB: Tiny Shared Block for Efficient DNN Deployment on NVCIM Accelerators
by: Qin, Yifan, et al.
Published: (2024)
by: Qin, Yifan, et al.
Published: (2024)
Similar Items
-
Efficient Arbitrary Precision Acceleration for Large Language Models on GPU Tensor Cores
by: Ma, Shaobo, et al.
Published: (2024) -
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration
by: Ma, Shaobo, et al.
Published: (2025) -
An FPGA-Based Reconfigurable Accelerator for Convolution-Transformer Hybrid EfficientViT
by: Shao, Haikuo, et al.
Published: (2024) -
AccLLM: Accelerating Long-Context LLM Inference Via Algorithm-Hardware Co-Design
by: Liang, Yanbiao, et al.
Published: (2025) -
M$^2$-ViT: Accelerating Hybrid Vision Transformers with Two-Level Mixed Quantization
by: Liang, Yanbiao, et al.
Published: (2024)