Performance Analysis of Matrix Multiplication for Deep Learning on the Edge
Fuente:
arXiv
Guardado en:
| Autores principales: | Ramírez, Cristian, Castelló, Adrián, Martínez, Héctor, Quintana-Ortí, Enrique S. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Enabling RISC-V Vector Code Generation in MLIR through Custom xDSL Lowerings
por: Lei, Jie, et al.
Publicado: (2026)
por: Lei, Jie, et al.
Publicado: (2026)
GUST: Graph Edge-Coloring Utilization for Accelerating Sparse Matrix Vector Multiplication
por: Gerami, Armin, et al.
Publicado: (2024)
por: Gerami, Armin, et al.
Publicado: (2024)
ADiP: Adaptive-Precision Systolic Array for Matrix Multiplication Acceleration
por: Abdelmaksoud, Ahmed J., et al.
Publicado: (2025)
por: Abdelmaksoud, Ahmed J., et al.
Publicado: (2025)
D-Legion: A Scalable Many-Core Architecture for Accelerating Matrix Multiplication in Quantized LLMs
por: Abdelmaksoud, Ahmed J., et al.
Publicado: (2026)
por: Abdelmaksoud, Ahmed J., et al.
Publicado: (2026)
Systolic Array Data Flows for Efficient Matrix Multiplication in Deep Neural Networks
por: Raja, Tejas
Publicado: (2024)
por: Raja, Tejas
Publicado: (2024)
Empowering Vector Architectures for ML: The CAMP Architecture for Matrix Multiplication
por: Nojehdeh, Mohammadreza Esmali, et al.
Publicado: (2025)
por: Nojehdeh, Mohammadreza Esmali, et al.
Publicado: (2025)
Fair and Square: Replacing One Real Multiplication with a Single Square and One Complex Multiplication with Three Squares When Performing Matrix Multiplication and Convolutions
por: Liguori, Vincenzo
Publicado: (2026)
por: Liguori, Vincenzo
Publicado: (2026)
Edge GPU Aware Multiple AI Model Pipeline for Accelerated MRI Reconstruction and Analysis
por: Majeed, Ashiyana Abdul, et al.
Publicado: (2025)
por: Majeed, Ashiyana Abdul, et al.
Publicado: (2025)
bitSMM: A bit-Serial Matrix Multiplication Accelerator
por: Antunes, Pedro, et al.
Publicado: (2026)
por: Antunes, Pedro, et al.
Publicado: (2026)
Towards Zero-Stall Matrix Multiplication on Energy-Efficient RISC-V Clusters for Machine Learning Acceleration
por: Colagrande, Luca, et al.
Publicado: (2025)
por: Colagrande, Luca, et al.
Publicado: (2025)
Optimizing Structured-Sparse Matrix Multiplication in RISC-V Vector Processors
por: Titopoulos, Vasileios, et al.
Publicado: (2025)
por: Titopoulos, Vasileios, et al.
Publicado: (2025)
A Matrix Decomposition Method for Odd-Type Gaussian Normal Basis Multiplication
por: Phalakarn, Kittiphon, et al.
Publicado: (2025)
por: Phalakarn, Kittiphon, et al.
Publicado: (2025)
DeMM: A Decoupled Matrix Multiplication Engine Supporting Relaxed Structured Sparsity
por: Peltekis, Christodoulos, et al.
Publicado: (2024)
por: Peltekis, Christodoulos, et al.
Publicado: (2024)
Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication
por: Shan, Haoxuan, et al.
Publicado: (2025)
por: Shan, Haoxuan, et al.
Publicado: (2025)
Systolic Array Acceleration of Diagonal-Optimized Sparse-Sparse Matrix Multiplication for Efficient Quantum Simulation
por: Su, Yuchao, et al.
Publicado: (2025)
por: Su, Yuchao, et al.
Publicado: (2025)
Performance evaluation of acceleration of convolutional layers on OpenEdgeCGRA
por: Carpentieri, Nicolò, et al.
Publicado: (2024)
por: Carpentieri, Nicolò, et al.
Publicado: (2024)
MX: Enhancing RISC-V's Vector ISA for Ultra-Low Overhead, Energy-Efficient Matrix Multiplication
por: Perotti, Matteo, et al.
Publicado: (2024)
por: Perotti, Matteo, et al.
Publicado: (2024)
Sparse-on-Dense: Area and Energy-Efficient Computing of Sparse Neural Networks on Dense Matrix Multiplication Accelerators
por: Yoon, Hyunsung, et al.
Publicado: (2026)
por: Yoon, Hyunsung, et al.
Publicado: (2026)
Salient Store: Enabling Smart Storage for Continuous Learning Edge Servers
por: Mishra, Cyan Subhra, et al.
Publicado: (2024)
por: Mishra, Cyan Subhra, et al.
Publicado: (2024)
Design Environment of Quantization-Aware Edge AI Hardware for Few-Shot Learning
por: Kanda, R., et al.
Publicado: (2026)
por: Kanda, R., et al.
Publicado: (2026)
Reconfigurable Digital RRAM Logic Enables In-Situ Pruning and Learning for Edge AI
por: Wang, Songqi, et al.
Publicado: (2025)
por: Wang, Songqi, et al.
Publicado: (2025)
Theoretical Analysis of the Efficient-Memory Matrix Storage Method for Quantum Emulation Accelerators with Gate Fusion on FPGAs
por: Le, Tran Xuan Hieu, et al.
Publicado: (2024)
por: Le, Tran Xuan Hieu, et al.
Publicado: (2024)
Bit-Width-Aware Design Environment for Few-Shot Learning on Edge AI Hardware
por: Kanda, R., et al.
Publicado: (2026)
por: Kanda, R., et al.
Publicado: (2026)
FAME: FPGA Acceleration of Secure Matrix Multiplication with Homomorphic Encryption
por: Xu, Zhihan, et al.
Publicado: (2025)
por: Xu, Zhihan, et al.
Publicado: (2025)
EdgeLLM: A Highly Efficient CPU-FPGA Heterogeneous Edge Accelerator for Large Language Models
por: Huang, Mingqiang, et al.
Publicado: (2024)
por: Huang, Mingqiang, et al.
Publicado: (2024)
Leveraging FPGAs for Homomorphic Matrix-Vector Multiplication in Oblivious Message Retrieval
por: Bosworth, Grant, et al.
Publicado: (2025)
por: Bosworth, Grant, et al.
Publicado: (2025)
EdgeMM: Multi-Core CPU with Heterogeneous AI-Extension and Activation-aware Weight Pruning for Multimodal LLMs at Edge
por: Bai, Kangbo, et al.
Publicado: (2025)
por: Bai, Kangbo, et al.
Publicado: (2025)
Hardware-Centric Analysis of DeepSeek's Multi-Head Latent Attention
por: Geens, Robin, et al.
Publicado: (2025)
por: Geens, Robin, et al.
Publicado: (2025)
Mapping code on Coarse Grained Reconfigurable Arrays using a SAT solver
por: Tirelli, Cristian, et al.
Publicado: (2025)
por: Tirelli, Cristian, et al.
Publicado: (2025)
Benchmarking Deep Learning Convolutions on Energy-constrained CPUs
por: Galvez, Enrique, et al.
Publicado: (2025)
por: Galvez, Enrique, et al.
Publicado: (2025)
Switchable Single/Dual Edge Registers for Pipeline Architecture
por: Singh, Suyash Vardhan, et al.
Publicado: (2024)
por: Singh, Suyash Vardhan, et al.
Publicado: (2024)
MING: An Automated CNN-to-Edge MLIR HLS framework
por: Bi, Jiahong, et al.
Publicado: (2026)
por: Bi, Jiahong, et al.
Publicado: (2026)
Hardware-Aware DNN Compression for Homogeneous Edge Devices
por: Zhang, Kunlong, et al.
Publicado: (2025)
por: Zhang, Kunlong, et al.
Publicado: (2025)
Computing-In-Memory Aware Model Adaption For Edge Devices
por: Lin, Ming-Han, et al.
Publicado: (2025)
por: Lin, Ming-Han, et al.
Publicado: (2025)
Flexible Bit-Truncation Memory for Approximate Applications on the Edge
por: Oswald, William, et al.
Publicado: (2025)
por: Oswald, William, et al.
Publicado: (2025)
Polaris: Multi-Fidelity Design Space Exploration of Deep Learning Accelerators
por: Sakhuja, Chirag, et al.
Publicado: (2024)
por: Sakhuja, Chirag, et al.
Publicado: (2024)
SigDLA: A Deep Learning Accelerator Extension for Signal Processing
por: Fu, Fangfa, et al.
Publicado: (2024)
por: Fu, Fangfa, et al.
Publicado: (2024)
Fast and Practical Strassen's Matrix Multiplication using FPGAs
por: Ahmad, Afzal, et al.
Publicado: (2024)
por: Ahmad, Afzal, et al.
Publicado: (2024)
SuperUROP: An FPGA-Based Spatial Accelerator for Sparse Matrix Operations
por: Parthasarathy, Rishab
Publicado: (2025)
por: Parthasarathy, Rishab
Publicado: (2025)
Analysis of Single Event Induced Bit Faults in a Deep Neural Network Accelerator Pipeline
por: Jonckers, Naïn, et al.
Publicado: (2025)
por: Jonckers, Naïn, et al.
Publicado: (2025)
Ejemplares similares
-
Enabling RISC-V Vector Code Generation in MLIR through Custom xDSL Lowerings
por: Lei, Jie, et al.
Publicado: (2026) -
GUST: Graph Edge-Coloring Utilization for Accelerating Sparse Matrix Vector Multiplication
por: Gerami, Armin, et al.
Publicado: (2024) -
ADiP: Adaptive-Precision Systolic Array for Matrix Multiplication Acceleration
por: Abdelmaksoud, Ahmed J., et al.
Publicado: (2025) -
D-Legion: A Scalable Many-Core Architecture for Accelerating Matrix Multiplication in Quantized LLMs
por: Abdelmaksoud, Ahmed J., et al.
Publicado: (2026) -
Systolic Array Data Flows for Efficient Matrix Multiplication in Deep Neural Networks
por: Raja, Tejas
Publicado: (2024)