MatrixFlow: System-Accelerator co-design for high-performance transformer applications
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Qunyou, Zapater, Marina, Atienza, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mitigating the Bandwidth Wall via Data-Streaming System-Accelerator Co-Design
von: Liu, Qunyou, et al.
Veröffentlicht: (2026)
von: Liu, Qunyou, et al.
Veröffentlicht: (2026)
Gem5-AcceSys: Enabling System-Level Exploration of Standard Interconnects for Novel Accelerators
von: Liu, Qunyou, et al.
Veröffentlicht: (2025)
von: Liu, Qunyou, et al.
Veröffentlicht: (2025)
SigmaQuant: Hardware-Aware Heterogeneous Quantization Method for Edge DNN Inference
von: Liu, Qunyou, et al.
Veröffentlicht: (2026)
von: Liu, Qunyou, et al.
Veröffentlicht: (2026)
CXLRAMSim v1.0: System-Level Exploration of CXL Memory Expander Cards
von: Pathak, Karan, et al.
Veröffentlicht: (2026)
von: Pathak, Karan, et al.
Veröffentlicht: (2026)
STRELA: STReaming ELAstic CGRA Accelerator for Embedded Systems
von: Vazquez, Daniel, et al.
Veröffentlicht: (2024)
von: Vazquez, Daniel, et al.
Veröffentlicht: (2024)
HAVEN: High-Bandwidth Flash Augmented Vector Engine for Large-Scale Approximate Nearest-Neighbor Search Acceleration
von: Hsu, Po-Kai, et al.
Veröffentlicht: (2026)
von: Hsu, Po-Kai, et al.
Veröffentlicht: (2026)
Exploiting pre-optimized kernels with polyhedral transformations for CGRA compilation
von: Wang, Yuxuan, et al.
Veröffentlicht: (2026)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2026)
GeneTEK: Low-power, high-performance and scalable FPGA architecture for exact unit-cost edit distance
von: Espinosa, Elena, et al.
Veröffentlicht: (2025)
von: Espinosa, Elena, et al.
Veröffentlicht: (2025)
X-HEEP: An Open-Source, Configurable and Extendible RISC-V Microcontroller for the Exploration of Ultra-Low-Power Edge Accelerators
von: Machetti, Simone, et al.
Veröffentlicht: (2024)
von: Machetti, Simone, et al.
Veröffentlicht: (2024)
Quadrilatero: A RISC-V programmable matrix coprocessor for low-power edge applications
von: Cammarata, Danilo, et al.
Veröffentlicht: (2025)
von: Cammarata, Danilo, et al.
Veröffentlicht: (2025)
DRACO: Co-design for DSP-Efficient Rigid Body Dynamics Accelerator
von: Liu, Xingyu, et al.
Veröffentlicht: (2025)
von: Liu, Xingyu, et al.
Veröffentlicht: (2025)
bitSMM: A bit-Serial Matrix Multiplication Accelerator
von: Antunes, Pedro, et al.
Veröffentlicht: (2026)
von: Antunes, Pedro, et al.
Veröffentlicht: (2026)
ADiP: Adaptive-Precision Systolic Array for Matrix Multiplication Acceleration
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2025)
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2025)
SuperUROP: An FPGA-Based Spatial Accelerator for Sparse Matrix Operations
von: Parthasarathy, Rishab
Veröffentlicht: (2025)
von: Parthasarathy, Rishab
Veröffentlicht: (2025)
Just TestIt! An SBST Approach To Automate System-Integration Testing
von: Terzano, Tommaso, et al.
Veröffentlicht: (2025)
von: Terzano, Tommaso, et al.
Veröffentlicht: (2025)
FTTN: Feature-Targeted Testing for Numerical Properties of NVIDIA & AMD Matrix Accelerators
von: Li, Xinyi, et al.
Veröffentlicht: (2024)
von: Li, Xinyi, et al.
Veröffentlicht: (2024)
Co-Design of CNN Accelerators for TinyML using Approximate Matrix Decomposition
von: Morales, José Juan Hernández, et al.
Veröffentlicht: (2026)
von: Morales, José Juan Hernández, et al.
Veröffentlicht: (2026)
GUST: Graph Edge-Coloring Utilization for Accelerating Sparse Matrix Vector Multiplication
von: Gerami, Armin, et al.
Veröffentlicht: (2024)
von: Gerami, Armin, et al.
Veröffentlicht: (2024)
Algorithm-hardware co-design for Energy-Efficient A/D conversion in ReRAM-based accelerators
von: Zhang, Chenguang, et al.
Veröffentlicht: (2024)
von: Zhang, Chenguang, et al.
Veröffentlicht: (2024)
Optimized Spatial Architecture Mapping Flow for Transformer Accelerators
von: Xu, Haocheng, et al.
Veröffentlicht: (2024)
von: Xu, Haocheng, et al.
Veröffentlicht: (2024)
Invited Paper: FEMU: An Open-Source and Configurable Emulation Framework for Prototyping TinyAI Heterogeneous Systems
von: Machetti, Simone, et al.
Veröffentlicht: (2025)
von: Machetti, Simone, et al.
Veröffentlicht: (2025)
SpeedLLM: An FPGA Co-design of Large Language Model Inference Accelerator
von: Wang, Peipei, et al.
Veröffentlicht: (2025)
von: Wang, Peipei, et al.
Veröffentlicht: (2025)
X-HEEP: An Open-Source, Configurable and Extendible RISC-V Platform for TinyAI Applications
von: Machetti, Simone, et al.
Veröffentlicht: (2025)
von: Machetti, Simone, et al.
Veröffentlicht: (2025)
Increasing the Energy-Efficiency of Wearables Using Low-Precision Posit Arithmetic with PHEE
von: Mallasén, David, et al.
Veröffentlicht: (2025)
von: Mallasén, David, et al.
Veröffentlicht: (2025)
Systolic Array Data Flows for Efficient Matrix Multiplication in Deep Neural Networks
von: Raja, Tejas
Veröffentlicht: (2024)
von: Raja, Tejas
Veröffentlicht: (2024)
Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication
von: Shan, Haoxuan, et al.
Veröffentlicht: (2025)
von: Shan, Haoxuan, et al.
Veröffentlicht: (2025)
Systolic Array Acceleration of Diagonal-Optimized Sparse-Sparse Matrix Multiplication for Efficient Quantum Simulation
von: Su, Yuchao, et al.
Veröffentlicht: (2025)
von: Su, Yuchao, et al.
Veröffentlicht: (2025)
D-Legion: A Scalable Many-Core Architecture for Accelerating Matrix Multiplication in Quantized LLMs
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2026)
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2026)
Hardware-Software Co-design for 3D-DRAM-based LLM Serving Accelerator
von: Li, Cong, et al.
Veröffentlicht: (2026)
von: Li, Cong, et al.
Veröffentlicht: (2026)
Inclusive-PIM: Hardware-Software Co-design for Broad Acceleration on Commercial PIM Architectures
von: Alsop, Johnathan, et al.
Veröffentlicht: (2023)
von: Alsop, Johnathan, et al.
Veröffentlicht: (2023)
Towards Zero-Stall Matrix Multiplication on Energy-Efficient RISC-V Clusters for Machine Learning Acceleration
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
Sparse-on-Dense: Area and Energy-Efficient Computing of Sparse Neural Networks on Dense Matrix Multiplication Accelerators
von: Yoon, Hyunsung, et al.
Veröffentlicht: (2026)
von: Yoon, Hyunsung, et al.
Veröffentlicht: (2026)
Theoretical Analysis of the Efficient-Memory Matrix Storage Method for Quantum Emulation Accelerators with Gate Fusion on FPGAs
von: Le, Tran Xuan Hieu, et al.
Veröffentlicht: (2024)
von: Le, Tran Xuan Hieu, et al.
Veröffentlicht: (2024)
Accelerator-assisted Floating-point ASIP for Communication and Positioning in Massive MIMO Systems
von: Attari, Mohammad, et al.
Veröffentlicht: (2025)
von: Attari, Mohammad, et al.
Veröffentlicht: (2025)
LLMulator: Generalizable Cost Modeling for Dataflow Accelerators with Input-Adaptive Control Flow
von: Chang, Kaiyan, et al.
Veröffentlicht: (2025)
von: Chang, Kaiyan, et al.
Veröffentlicht: (2025)
An integrated design of energy and indoor environmental quality monitoring system for effective building performance management
von: Zakka, Vincent Gbouna, et al.
Veröffentlicht: (2025)
von: Zakka, Vincent Gbouna, et al.
Veröffentlicht: (2025)
e-GPU: An Open-Source and Configurable RISC-V Graphic Processing Unit for TinyAI Applications
von: Machetti, Simone, et al.
Veröffentlicht: (2025)
von: Machetti, Simone, et al.
Veröffentlicht: (2025)
Scalable and RISC-V Programmable Near-Memory Computing Architectures for Edge Nodes
von: Caon, Michele, et al.
Veröffentlicht: (2024)
von: Caon, Michele, et al.
Veröffentlicht: (2024)
Performance evaluation of acceleration of convolutional layers on OpenEdgeCGRA
von: Carpentieri, Nicolò, et al.
Veröffentlicht: (2024)
von: Carpentieri, Nicolò, et al.
Veröffentlicht: (2024)
Enhancing Regression Models for Complex Systems Using Evolutionary Techniques for Feature Engineering
von: Arroba, Patricia, et al.
Veröffentlicht: (2024)
von: Arroba, Patricia, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Mitigating the Bandwidth Wall via Data-Streaming System-Accelerator Co-Design
von: Liu, Qunyou, et al.
Veröffentlicht: (2026) -
Gem5-AcceSys: Enabling System-Level Exploration of Standard Interconnects for Novel Accelerators
von: Liu, Qunyou, et al.
Veröffentlicht: (2025) -
SigmaQuant: Hardware-Aware Heterogeneous Quantization Method for Edge DNN Inference
von: Liu, Qunyou, et al.
Veröffentlicht: (2026) -
CXLRAMSim v1.0: System-Level Exploration of CXL Memory Expander Cards
von: Pathak, Karan, et al.
Veröffentlicht: (2026) -
STRELA: STReaming ELAstic CGRA Accelerator for Embedded Systems
von: Vazquez, Daniel, et al.
Veröffentlicht: (2024)