Mamba-X: An End-to-End Vision Mamba Accelerator for Edge Computing Devices
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yoon, Dongho, Lee, Gungyu, Chang, Jaewon, Lee, Yunjae, Lee, Dongjae, Rhu, Minsoo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PIM-MMU: A Memory Management Unit for Accelerating Data Transfers in Commercial PIM Systems
von: Lee, Dongjae, et al.
Veröffentlicht: (2024)
von: Lee, Dongjae, et al.
Veröffentlicht: (2024)
PreSto: An In-Storage Data Preprocessing System for Training Recommendation Models
von: Lee, Yunjae, et al.
Veröffentlicht: (2024)
von: Lee, Yunjae, et al.
Veröffentlicht: (2024)
PIM-malloc: A Fast and Scalable Dynamic Memory Allocator for Processing-In-Memory (PIM) Architectures
von: Lee, Dongjae, et al.
Veröffentlicht: (2025)
von: Lee, Dongjae, et al.
Veröffentlicht: (2025)
Pathfinding Future PIM Architectures by Demystifying a Commercial PIM Technology
von: Hyun, Bongjoon, et al.
Veröffentlicht: (2023)
von: Hyun, Bongjoon, et al.
Veröffentlicht: (2023)
CIMR-V: An End-to-End SRAM-based CIM Accelerator with RISC-V for AI Edge Device
von: and, Yan-Cheng Guo, et al.
Veröffentlicht: (2025)
von: and, Yan-Cheng Guo, et al.
Veröffentlicht: (2025)
SpecMamba: Accelerating Mamba Inference on FPGA with Speculative Decoding
von: Zhong, Linfeng, et al.
Veröffentlicht: (2025)
von: Zhong, Linfeng, et al.
Veröffentlicht: (2025)
TriGen: NPU Architecture for End-to-End Acceleration of Large Language Models based on SW-HW Co-Design
von: Lee, Jonghun, et al.
Veröffentlicht: (2026)
von: Lee, Jonghun, et al.
Veröffentlicht: (2026)
SLDB: An End-To-End Heterogeneous System-on-Chip Benchmark Suite for LLM-Aided Design
von: Alvanaki, Elisavet Lydia, et al.
Veröffentlicht: (2025)
von: Alvanaki, Elisavet Lydia, et al.
Veröffentlicht: (2025)
The Cost of Dynamic Reasoning: Demystifying AI Agents and Test-Time Scaling from an AI Infrastructure Perspective
von: Kim, Jiin, et al.
Veröffentlicht: (2025)
von: Kim, Jiin, et al.
Veröffentlicht: (2025)
PASCAL: A Phase-Aware Scheduling Algorithm for Serving Reasoning-based Large Language Models
von: Cho, Eunyeong, et al.
Veröffentlicht: (2026)
von: Cho, Eunyeong, et al.
Veröffentlicht: (2026)
AutoGNN: End-to-End Hardware-Driven Graph Preprocessing for Enhanced GNN Performance
von: Kang, Seungkwan, et al.
Veröffentlicht: (2026)
von: Kang, Seungkwan, et al.
Veröffentlicht: (2026)
Efficient yet Accurate End-to-End SC Accelerator Design
von: Li, Meng, et al.
Veröffentlicht: (2024)
von: Li, Meng, et al.
Veröffentlicht: (2024)
PIMCOMP: An End-to-End DNN Compiler for Processing-In-Memory Accelerators
von: Sun, Xiaotian, et al.
Veröffentlicht: (2024)
von: Sun, Xiaotian, et al.
Veröffentlicht: (2024)
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
von: Kim, Hyeseong, et al.
Veröffentlicht: (2026)
von: Kim, Hyeseong, et al.
Veröffentlicht: (2026)
Voyager: An End-to-End Framework for Design-Space Exploration and Generation of DNN Accelerators
von: Prabhu, Kartik, et al.
Veröffentlicht: (2025)
von: Prabhu, Kartik, et al.
Veröffentlicht: (2025)
FastMamba: A High-Speed and Efficient Mamba Accelerator on FPGA with Accurate Quantization
von: Wang, Aotao, et al.
Veröffentlicht: (2025)
von: Wang, Aotao, et al.
Veröffentlicht: (2025)
Computing-In-Memory Aware Model Adaption For Edge Devices
von: Lin, Ming-Han, et al.
Veröffentlicht: (2025)
von: Lin, Ming-Han, et al.
Veröffentlicht: (2025)
ESACT: An End-to-End Sparse Accelerator for Compute-Intensive Transformers via Local Similarity
von: Liu, Hongxiang, et al.
Veröffentlicht: (2025)
von: Liu, Hongxiang, et al.
Veröffentlicht: (2025)
vTrain: A Simulation Framework for Evaluating Cost-effective and Compute-optimal Large Language Model Training
von: Bang, Jehyeon, et al.
Veröffentlicht: (2023)
von: Bang, Jehyeon, et al.
Veröffentlicht: (2023)
MARCA: Mamba Accelerator with ReConfigurable Architecture
von: Li, Jinhao, et al.
Veröffentlicht: (2024)
von: Li, Jinhao, et al.
Veröffentlicht: (2024)
MemIntelli: A Generic End-to-End Simulation Framework for Memristive Intelligent Computing
von: Zhou, Houji, et al.
Veröffentlicht: (2025)
von: Zhou, Houji, et al.
Veröffentlicht: (2025)
PowerFlow-DNN: Compiler-Directed Fine-Grained Power Orchestration for End-to-End Edge AI Inference
von: Chen, Paul, et al.
Veröffentlicht: (2026)
von: Chen, Paul, et al.
Veröffentlicht: (2026)
Debunking the CUDA Myth Towards GPU-based AI Systems
von: Lee, Yunjae, et al.
Veröffentlicht: (2024)
von: Lee, Yunjae, et al.
Veröffentlicht: (2024)
HH-PIM: Dynamic Optimization of Power and Performance with Heterogeneous-Hybrid PIM for Edge AI Devices
von: Jeon, Sangmin, et al.
Veröffentlicht: (2025)
von: Jeon, Sangmin, et al.
Veröffentlicht: (2025)
VitaLLM: A Versatile and Tiny Accelerator for Mixed-Precision LLM Inference on Edge Devices
von: Lin, Zi-Wei, et al.
Veröffentlicht: (2026)
von: Lin, Zi-Wei, et al.
Veröffentlicht: (2026)
PD-Swap: Prefill-Decode Logic Swapping for End-to-End LLM Inference on Edge FPGAs via Dynamic Partial Reconfiguration
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
HyperCroc: End-to-End Open-Source RISC-V MCU with a Plug-In Interface for Domain-Specific Accelerators
von: Sauter, Philippe, et al.
Veröffentlicht: (2026)
von: Sauter, Philippe, et al.
Veröffentlicht: (2026)
End-to-End Transformer Acceleration Through Processing-in-Memory Architectures
von: Yang, Xiaoxuan, et al.
Veröffentlicht: (2025)
von: Yang, Xiaoxuan, et al.
Veröffentlicht: (2025)
TeLLMe v2: An Efficient End-to-End Ternary LLM Prefill and Decode Accelerator with Table-Lookup Matmul on Edge FPGAs
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
DiSC: Resolution-Scalable Acceleration of Diffusion Models by Exploiting Sparsity and Cached Token Reuse with Hash-based Distribution
von: Yoon, Jieon, et al.
Veröffentlicht: (2026)
von: Yoon, Jieon, et al.
Veröffentlicht: (2026)
FIXME: Towards End-to-End Benchmarking of LLM-Aided Design Verification
von: Wan, Gwok-Waa, et al.
Veröffentlicht: (2025)
von: Wan, Gwok-Waa, et al.
Veröffentlicht: (2025)
GenPairX: A Hardware-Algorithm Co-Designed Accelerator for Paired-End Read Mapping
von: Eudine, Julien, et al.
Veröffentlicht: (2026)
von: Eudine, Julien, et al.
Veröffentlicht: (2026)
LoRA-Edge: Tensor-Train-Assisted LoRA for Practical CNN Fine-Tuning on Edge Devices
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025)
von: Kwak, Hyunseok, et al.
Veröffentlicht: (2025)
MiCo: End-to-End Mixed Precision Neural Network Co-Exploration Framework for Edge AI
von: Jiang, Zijun, et al.
Veröffentlicht: (2025)
von: Jiang, Zijun, et al.
Veröffentlicht: (2025)
CMAX-CAMEL: A Coarse-to-Fine Adaptive, Memory-Efficient, and Low-Power Edge Processor for Contrast Maximization
von: Min, Kyeongpil, et al.
Veröffentlicht: (2026)
von: Min, Kyeongpil, et al.
Veröffentlicht: (2026)
MCMComm: Hardware-Software Co-Optimization for End-to-End Communication in Multi-Chip-Modules
von: Raj, Ritik, et al.
Veröffentlicht: (2025)
von: Raj, Ritik, et al.
Veröffentlicht: (2025)
FASE: FPGA-Assisted Syscall Emulation for Rapid End-to-End Processor Performance Validation
von: Meng, Chengzhen, et al.
Veröffentlicht: (2025)
von: Meng, Chengzhen, et al.
Veröffentlicht: (2025)
Towards an End-To-End System for Real-Time Gesture Recognition from Surface Vibrations
von: Hettstedt, Florian, et al.
Veröffentlicht: (2026)
von: Hettstedt, Florian, et al.
Veröffentlicht: (2026)
Modeling Analog-Digital-Converter Energy and Area for Compute-In-Memory Accelerator Design
von: Andrulis, Tanner, et al.
Veröffentlicht: (2024)
von: Andrulis, Tanner, et al.
Veröffentlicht: (2024)
CD-PIM: A High-Bandwidth and Compute-Efficient LPDDR5-Based PIM for Low-Batch LLM Acceleration on Edge-Device
von: Lin, Ye, et al.
Veröffentlicht: (2026)
von: Lin, Ye, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
PIM-MMU: A Memory Management Unit for Accelerating Data Transfers in Commercial PIM Systems
von: Lee, Dongjae, et al.
Veröffentlicht: (2024) -
PreSto: An In-Storage Data Preprocessing System for Training Recommendation Models
von: Lee, Yunjae, et al.
Veröffentlicht: (2024) -
PIM-malloc: A Fast and Scalable Dynamic Memory Allocator for Processing-In-Memory (PIM) Architectures
von: Lee, Dongjae, et al.
Veröffentlicht: (2025) -
Pathfinding Future PIM Architectures by Demystifying a Commercial PIM Technology
von: Hyun, Bongjoon, et al.
Veröffentlicht: (2023) -
CIMR-V: An End-to-End SRAM-based CIM Accelerator with RISC-V for AI Edge Device
von: and, Yan-Cheng Guo, et al.
Veröffentlicht: (2025)