Lightweight Software Kernels and Hardware Extensions for Efficient Sparse Deep Neural Networks on Microcontrollers
Fuente:
arXiv
Guardado en:
| Autores principales: | Daghero, Francesco, Pagliari, Daniele Jahier, Conti, Francesco, Benini, Luca, Poncino, Massimo, Burrello, Alessio |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Accelerating Depthwise Separable Convolutions on Ultra-Low-Power Devices
por: Daghero, Francesco, et al.
Publicado: (2024)
por: Daghero, Francesco, et al.
Publicado: (2024)
Optimizing the Deployment of Tiny Transformers on Low-Power MCUs
por: Jung, Victor J. B., et al.
Publicado: (2024)
por: Jung, Victor J. B., et al.
Publicado: (2024)
MATCHA: Efficient Deployment of Deep Neural Networks on Multi-Accelerator Heterogeneous Edge SoCs
por: Russo, Enrico, et al.
Publicado: (2026)
por: Russo, Enrico, et al.
Publicado: (2026)
HTVM: Efficient Neural Network Deployment On Heterogeneous TinyML Platforms
por: Van Delm, Josse, et al.
Publicado: (2024)
por: Van Delm, Josse, et al.
Publicado: (2024)
MATCH: Model-Aware TVM-based Compilation for Heterogeneous Edge Devices
por: Hamdi, Mohamed Amine, et al.
Publicado: (2024)
por: Hamdi, Mohamed Amine, et al.
Publicado: (2024)
Ariel-ML: Computing Parallelization with Embedded Rust for Neural Networks on Heterogeneous Multi-core Microcontrollers
por: Huang, Zhaolan, et al.
Publicado: (2025)
por: Huang, Zhaolan, et al.
Publicado: (2025)
Optimizing Foundation Model Inference on a Many-tiny-core Open-source RISC-V Platform
por: Potocnik, Viviane, et al.
Publicado: (2024)
por: Potocnik, Viviane, et al.
Publicado: (2024)
Is Sparse Matrix Reordering Effective for Sparse Matrix-Vector Multiplication?
por: Asudeh, Omid, et al.
Publicado: (2025)
por: Asudeh, Omid, et al.
Publicado: (2025)
Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations
por: Islam, Tanzima Z., et al.
Publicado: (2024)
por: Islam, Tanzima Z., et al.
Publicado: (2024)
Can Tensor Cores Benefit Memory-Bound Kernels? (No!)
por: Zhang, Lingqi, et al.
Publicado: (2025)
por: Zhang, Lingqi, et al.
Publicado: (2025)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
por: Davis, Joshua H., et al.
Publicado: (2026)
por: Davis, Joshua H., et al.
Publicado: (2026)
Cache Blocking of Distributed-Memory Parallel Matrix Power Kernels
por: Lacey, Dane C., et al.
Publicado: (2024)
por: Lacey, Dane C., et al.
Publicado: (2024)
Hardware-Agnostic and Insightful Efficiency Metrics for Accelerated Systems: Definition and Implementation within TALP
por: Rahimi, Ghazal, et al.
Publicado: (2026)
por: Rahimi, Ghazal, et al.
Publicado: (2026)
Comparing the Performance of Heterogeneous Conjugate Gradient and Cholesky Solvers on Various Hardware Using SYCL
por: Thüring, Tim, et al.
Publicado: (2026)
por: Thüring, Tim, et al.
Publicado: (2026)
FACT: Compositional Kernel Synthesis with a Three-Stage Agentic Workflow
por: Heidari, Sina, et al.
Publicado: (2026)
por: Heidari, Sina, et al.
Publicado: (2026)
Staging Blocked Evaluation over Structured Sparse Matrices
por: Das, Pratyush, et al.
Publicado: (2024)
por: Das, Pratyush, et al.
Publicado: (2024)
AcceleratedKernels.jl: Cross-Architecture Parallel Algorithms from a Unified, Transpiled Codebase
por: Nicusan, Andrei-Leonard, et al.
Publicado: (2025)
por: Nicusan, Andrei-Leonard, et al.
Publicado: (2025)
Portable High-Performance Kernel Generation for a Computational Fluid Dynamics Code with DaCe
por: Andersson, Måns I., et al.
Publicado: (2025)
por: Andersson, Måns I., et al.
Publicado: (2025)
SHIRO: Near-Optimal Communication Strategies for Distributed Sparse Matrix Multiplication
por: Zhuang, Chen, et al.
Publicado: (2025)
por: Zhuang, Chen, et al.
Publicado: (2025)
SmartWatts: Self-Calibrating Software-Defined Power Meter for Containers
por: Fieni, Guillaume, et al.
Publicado: (2020)
por: Fieni, Guillaume, et al.
Publicado: (2020)
Scalable Systems and Software Architectures for High-Performance Computing on cloud platforms
por: Ramesh, Risshab Srinivas
Publicado: (2024)
por: Ramesh, Risshab Srinivas
Publicado: (2024)
Bringing Auto-tuning to HIP: Analysis of Tuning Impact and Difficulty on AMD and Nvidia GPUs
por: Lurati, Milo, et al.
Publicado: (2024)
por: Lurati, Milo, et al.
Publicado: (2024)
Architecture Specific Generation of Large Scale Lattice Boltzmann Methods for Sparse Complex Geometries
por: Suffa, Philipp, et al.
Publicado: (2024)
por: Suffa, Philipp, et al.
Publicado: (2024)
An Auto-tuning Method for Run-time Data Transformation for Sparse Matrix-Vector Multiplication
por: Katagiri, Takahiro, et al.
Publicado: (2024)
por: Katagiri, Takahiro, et al.
Publicado: (2024)
PICO: Performance Insights for Collective Operations
por: Pasqualoni, Saverio, et al.
Publicado: (2025)
por: Pasqualoni, Saverio, et al.
Publicado: (2025)
FalconFS: Distributed File System for Large-Scale Deep Learning Pipeline
por: Xu, Jingwei, et al.
Publicado: (2025)
por: Xu, Jingwei, et al.
Publicado: (2025)
Parallel I/O Characterization and Optimization on Large-Scale HPC Systems: A 360-Degree Survey
por: Ather, Hammad, et al.
Publicado: (2024)
por: Ather, Hammad, et al.
Publicado: (2024)
iSpLib: A Library for Accelerating Graph Neural Networks using Auto-tuned Sparse Operations
por: Anik, Md Saidul Hoque, et al.
Publicado: (2024)
por: Anik, Md Saidul Hoque, et al.
Publicado: (2024)
Constructive community race: full-density spiking neural network model drives neuromorphic computing
por: Senk, Johanna, et al.
Publicado: (2025)
por: Senk, Johanna, et al.
Publicado: (2025)
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
por: Nichols, Daniel, et al.
Publicado: (2025)
por: Nichols, Daniel, et al.
Publicado: (2025)
UPMEM Unleashed: Software Secrets for Speed
por: Chmielewski, Krystian, et al.
Publicado: (2025)
por: Chmielewski, Krystian, et al.
Publicado: (2025)
GhostServe: A Lightweight Checkpointing System in the Shadow for Fault-Tolerant LLM Serving
por: Jayakody, Shakya, et al.
Publicado: (2026)
por: Jayakody, Shakya, et al.
Publicado: (2026)
CoNST: Code Generator for Sparse Tensor Networks
por: Raje, Saurabh, et al.
Publicado: (2024)
por: Raje, Saurabh, et al.
Publicado: (2024)
EdgeProfiler: A Fast Profiling Framework for Lightweight LLMs on Edge Using Analytical Model
por: Pinnock, Alyssa, et al.
Publicado: (2025)
por: Pinnock, Alyssa, et al.
Publicado: (2025)
oneDAL Optimization for ARM Scalable Vector Extension: Maximizing Efficiency for High-Performance Data Science
por: Sharma, Chandan, et al.
Publicado: (2025)
por: Sharma, Chandan, et al.
Publicado: (2025)
Xabclib:A Fully Auto-tuned Sparse Iterative Solver
por: Katagiri, Takahiro, et al.
Publicado: (2024)
por: Katagiri, Takahiro, et al.
Publicado: (2024)
Multi-DNN Inference of Sparse Models on Edge SoCs
por: Luo, Jiawei, et al.
Publicado: (2026)
por: Luo, Jiawei, et al.
Publicado: (2026)
Extracting Practical, Actionable Energy Insights from Supercomputer Telemetry and Logs
por: Cornelius, Melanie, et al.
Publicado: (2025)
por: Cornelius, Melanie, et al.
Publicado: (2025)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
por: Zhuang, Chen, et al.
Publicado: (2024)
por: Zhuang, Chen, et al.
Publicado: (2024)
Profiling and optimization of multi-card GPU machine learning jobs
por: Lawenda, Marcin, et al.
Publicado: (2025)
por: Lawenda, Marcin, et al.
Publicado: (2025)
Ejemplares similares
-
Accelerating Depthwise Separable Convolutions on Ultra-Low-Power Devices
por: Daghero, Francesco, et al.
Publicado: (2024) -
Optimizing the Deployment of Tiny Transformers on Low-Power MCUs
por: Jung, Victor J. B., et al.
Publicado: (2024) -
MATCHA: Efficient Deployment of Deep Neural Networks on Multi-Accelerator Heterogeneous Edge SoCs
por: Russo, Enrico, et al.
Publicado: (2026) -
HTVM: Efficient Neural Network Deployment On Heterogeneous TinyML Platforms
por: Van Delm, Josse, et al.
Publicado: (2024) -
MATCH: Model-Aware TVM-based Compilation for Heterogeneous Edge Devices
por: Hamdi, Mohamed Amine, et al.
Publicado: (2024)