MARCA: Mamba Accelerator with ReConfigurable Architecture
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Jinhao, Huang, Shan, Xu, Jiaming, Liu, Jun, Ding, Li, Xu, Ningyi, Dai, Guohao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SoftmAP: Software-Hardware Co-design for Integer-Only Softmax on Associative Processors
di: Rakka, Mariam, et al.
Pubblicazione: (2024)
di: Rakka, Mariam, et al.
Pubblicazione: (2024)
ROMA: a Read-Only-Memory-based Accelerator for QLoRA-based On-Device LLM
di: Wang, Wenqiang, et al.
Pubblicazione: (2025)
di: Wang, Wenqiang, et al.
Pubblicazione: (2025)
Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective
di: Li, Jinhao, et al.
Pubblicazione: (2024)
di: Li, Jinhao, et al.
Pubblicazione: (2024)
FastMamba: A High-Speed and Efficient Mamba Accelerator on FPGA with Accurate Quantization
di: Wang, Aotao, et al.
Pubblicazione: (2025)
di: Wang, Aotao, et al.
Pubblicazione: (2025)
A Fully Hardware Implemented Accelerator Design in ReRAM Analog Computing without ADCs
di: Dang, Peng, et al.
Pubblicazione: (2024)
di: Dang, Peng, et al.
Pubblicazione: (2024)
A Configurable and Efficient Memory Hierarchy for Neural Network Hardware Accelerator
di: Bause, Oliver, et al.
Pubblicazione: (2024)
di: Bause, Oliver, et al.
Pubblicazione: (2024)
DeepGate4: Efficient and Effective Representation Learning for Circuit Design at Scale
di: Zheng, Ziyang, et al.
Pubblicazione: (2025)
di: Zheng, Ziyang, et al.
Pubblicazione: (2025)
FlightLLM: Efficient Large Language Model Inference with a Complete Mapping Flow on FPGAs
di: Zeng, Shulin, et al.
Pubblicazione: (2024)
di: Zeng, Shulin, et al.
Pubblicazione: (2024)
CiMNet: Towards Joint Optimization for DNN Architecture and Configuration for Compute-In-Memory Hardware
di: Kundu, Souvik, et al.
Pubblicazione: (2024)
di: Kundu, Souvik, et al.
Pubblicazione: (2024)
Expert Streaming: Accelerating Low-Batch MoE Inference via Multi-chiplet Architecture and Dynamic Expert Trajectory Scheduling
di: Ma, Songchen, et al.
Pubblicazione: (2026)
di: Ma, Songchen, et al.
Pubblicazione: (2026)
Bitwise Systolic Array Architecture for Runtime-Reconfigurable Multi-precision Quantized Multiplication on Hardware Accelerators
di: Liu, Yuhao, et al.
Pubblicazione: (2026)
di: Liu, Yuhao, et al.
Pubblicazione: (2026)
TL-nvSRAM-CIM: Ultra-High-Density Three-Level ReRAM-Assisted Computing-in-nvSRAM with DC-Power Free Restore and Ternary MAC Operations
di: Wang, Dengfeng, et al.
Pubblicazione: (2023)
di: Wang, Dengfeng, et al.
Pubblicazione: (2023)
Idle is the New Sleep: Configuration-Aware Alternative to Powering Off FPGA-Based DL Accelerators During Inactivity
di: Qian, Chao, et al.
Pubblicazione: (2024)
di: Qian, Chao, et al.
Pubblicazione: (2024)
A Survey on Design Methodologies for Accelerating Deep Learning on Heterogeneous Architectures
di: Curzel, Serena, et al.
Pubblicazione: (2023)
di: Curzel, Serena, et al.
Pubblicazione: (2023)
Using the Abstract Computer Architecture Description Language to Model AI Hardware Accelerators
di: Müller, Mika Markus, et al.
Pubblicazione: (2024)
di: Müller, Mika Markus, et al.
Pubblicazione: (2024)
ELSA: An ELastic SNN Inference Architecture for Efficient Neuromorphic Computing
di: You, Kang, et al.
Pubblicazione: (2026)
di: You, Kang, et al.
Pubblicazione: (2026)
SpikeX: Exploring Accelerator Architecture and Network-Hardware Co-Optimization for Sparse Spiking Neural Networks
di: Xu, Boxun, et al.
Pubblicazione: (2025)
di: Xu, Boxun, et al.
Pubblicazione: (2025)
BF-IMNA: A Bit Fluid In-Memory Neural Architecture for Neural Network Acceleration
di: Rakka, Mariam, et al.
Pubblicazione: (2024)
di: Rakka, Mariam, et al.
Pubblicazione: (2024)
Architectural Design and Performance Analysis of FPGA based AI Accelerators: A Comprehensive Review
di: Chatterjee, Soumita, et al.
Pubblicazione: (2026)
di: Chatterjee, Soumita, et al.
Pubblicazione: (2026)
LUMINA: LLM-Guided GPU Architecture Exploration via Bottleneck Analysis
di: Zhang, Tao, et al.
Pubblicazione: (2026)
di: Zhang, Tao, et al.
Pubblicazione: (2026)
ArchPower: Dataset for Architecture-Level Power Modeling of Modern CPU Design
di: Zhang, Qijun, et al.
Pubblicazione: (2025)
di: Zhang, Qijun, et al.
Pubblicazione: (2025)
TriGen: NPU Architecture for End-to-End Acceleration of Large Language Models based on SW-HW Co-Design
di: Lee, Jonghun, et al.
Pubblicazione: (2026)
di: Lee, Jonghun, et al.
Pubblicazione: (2026)
LLM-DSE: Searching Accelerator Parameters with LLM Agents
di: Wang, Hanyu, et al.
Pubblicazione: (2025)
di: Wang, Hanyu, et al.
Pubblicazione: (2025)
SWAT: Scalable and Efficient Window Attention-based Transformers Acceleration on FPGAs
di: Bai, Zhenyu, et al.
Pubblicazione: (2024)
di: Bai, Zhenyu, et al.
Pubblicazione: (2024)
AHASD: Asynchronous Heterogeneous Architecture for LLM Adaptive Drafting Speculative Decoding on Mobile Devices
di: Zirui, Ma, et al.
Pubblicazione: (2026)
di: Zirui, Ma, et al.
Pubblicazione: (2026)
FractalCloud: A Fractal-Inspired Architecture for Efficient Large-Scale Point Cloud Processing
di: Fu, Yuzhe, et al.
Pubblicazione: (2025)
di: Fu, Yuzhe, et al.
Pubblicazione: (2025)
SpecMamba: Accelerating Mamba Inference on FPGA with Speculative Decoding
di: Zhong, Linfeng, et al.
Pubblicazione: (2025)
di: Zhong, Linfeng, et al.
Pubblicazione: (2025)
Comprehensive Design Space Exploration for Tensorized Neural Network Hardware Accelerators
di: Zhang, Jinsong, et al.
Pubblicazione: (2025)
di: Zhang, Jinsong, et al.
Pubblicazione: (2025)
Efficient Orchestrated AI Workflows Execution on Scale-out Spatial Architecture
di: Deng, Jinyi, et al.
Pubblicazione: (2024)
di: Deng, Jinyi, et al.
Pubblicazione: (2024)
Monitor Placement for Fault Localization in Deep Neural Network Accelerators
di: Liu, Wei-Kai
Pubblicazione: (2023)
di: Liu, Wei-Kai
Pubblicazione: (2023)
DPUV4E: High-Throughput DPU Architecture Design for CNN on Versal ACAP
di: Li, Guoyu, et al.
Pubblicazione: (2025)
di: Li, Guoyu, et al.
Pubblicazione: (2025)
Towards Efficient Neuro-Symbolic AI: From Workload Characterization to Hardware Architecture
di: Wan, Zishen, et al.
Pubblicazione: (2024)
di: Wan, Zishen, et al.
Pubblicazione: (2024)
MICSim: A Modular Simulator for Mixed-signal Compute-in-Memory based AI Accelerator
di: Wang, Cong, et al.
Pubblicazione: (2024)
di: Wang, Cong, et al.
Pubblicazione: (2024)
Zero-Space Cost Fault Tolerance for Transformer-based Language Models on ReRAM
di: Li, Bingbing, et al.
Pubblicazione: (2024)
di: Li, Bingbing, et al.
Pubblicazione: (2024)
Enable Lightweight and Precision-Scalable Posit/IEEE-754 Arithmetic in RISC-V Cores for Transprecision Computing
di: Li, Qiong, et al.
Pubblicazione: (2025)
di: Li, Qiong, et al.
Pubblicazione: (2025)
ReChisel: Effective Automatic Chisel Code Generation by LLM with Reflection
di: Niu, Juxin, et al.
Pubblicazione: (2025)
di: Niu, Juxin, et al.
Pubblicazione: (2025)
HG-PIPE: Vision Transformer Acceleration with Hybrid-Grained Pipeline
di: Guo, Qingyu, et al.
Pubblicazione: (2024)
di: Guo, Qingyu, et al.
Pubblicazione: (2024)
GRAU: Generic Reconfigurable Activation Unit Design for Neural Network Hardware Accelerators
di: Liu, Yuhao, et al.
Pubblicazione: (2026)
di: Liu, Yuhao, et al.
Pubblicazione: (2026)
SPADE: Sparse Pillar-based 3D Object Detection Accelerator for Autonomous Driving
di: Lee, Minjae, et al.
Pubblicazione: (2023)
di: Lee, Minjae, et al.
Pubblicazione: (2023)
REASON: Accelerating Probabilistic Logical Reasoning for Scalable Neuro-Symbolic Intelligence
di: Wan, Zishen, et al.
Pubblicazione: (2026)
di: Wan, Zishen, et al.
Pubblicazione: (2026)
Documenti analoghi
-
SoftmAP: Software-Hardware Co-design for Integer-Only Softmax on Associative Processors
di: Rakka, Mariam, et al.
Pubblicazione: (2024) -
ROMA: a Read-Only-Memory-based Accelerator for QLoRA-based On-Device LLM
di: Wang, Wenqiang, et al.
Pubblicazione: (2025) -
Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective
di: Li, Jinhao, et al.
Pubblicazione: (2024) -
FastMamba: A High-Speed and Efficient Mamba Accelerator on FPGA with Accurate Quantization
di: Wang, Aotao, et al.
Pubblicazione: (2025) -
A Fully Hardware Implemented Accelerator Design in ReRAM Analog Computing without ADCs
di: Dang, Peng, et al.
Pubblicazione: (2024)