Accelerating Depthwise Separable Convolutions on Ultra-Low-Power Devices
Fuente:
arXiv
Saved in:
| Main Authors: | Daghero, Francesco, Burrello, Alessio, Poncino, Massimo, Macii, Enrico, Pagliari, Daniele Jahier |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lightweight Software Kernels and Hardware Extensions for Efficient Sparse Deep Neural Networks on Microcontrollers
by: Daghero, Francesco, et al.
Published: (2025)
by: Daghero, Francesco, et al.
Published: (2025)
MATCHA: Efficient Deployment of Deep Neural Networks on Multi-Accelerator Heterogeneous Edge SoCs
by: Russo, Enrico, et al.
Published: (2026)
by: Russo, Enrico, et al.
Published: (2026)
Optimizing the Deployment of Tiny Transformers on Low-Power MCUs
by: Jung, Victor J. B., et al.
Published: (2024)
by: Jung, Victor J. B., et al.
Published: (2024)
MATCH: Model-Aware TVM-based Compilation for Heterogeneous Edge Devices
by: Hamdi, Mohamed Amine, et al.
Published: (2024)
by: Hamdi, Mohamed Amine, et al.
Published: (2024)
EDEA: Efficient Dual-Engine Accelerator for Depthwise Separable Convolution with Direct Data Transfer
by: Chen, Yi, et al.
Published: (2025)
by: Chen, Yi, et al.
Published: (2025)
Optimized Deployment of Deep Neural Networks for Visual Pose Estimation on Nano-drones
by: Risso, Matteo, et al.
Published: (2024)
by: Risso, Matteo, et al.
Published: (2024)
Accelerating Transposed Convolutions on FPGA-based Edge Devices
by: Haris, Jude, et al.
Published: (2025)
by: Haris, Jude, et al.
Published: (2025)
HTVM: Efficient Neural Network Deployment On Heterogeneous TinyML Platforms
by: Van Delm, Josse, et al.
Published: (2024)
by: Van Delm, Josse, et al.
Published: (2024)
FedQUIT: On-Device Federated Unlearning via a Quasi-Competent Virtual Teacher
by: Mora, Alessio, et al.
Published: (2024)
by: Mora, Alessio, et al.
Published: (2024)
Unleashing the Power of Continual Learning on Non-Centralized Devices: A Survey
by: Li, Yichen, et al.
Published: (2024)
by: Li, Yichen, et al.
Published: (2024)
Accelerating Local LLMs on Resource-Constrained Edge Devices via Distributed Prompt Caching
by: Matsutani, Hiroki, et al.
Published: (2026)
by: Matsutani, Hiroki, et al.
Published: (2026)
The CAPSARII Approach to Cyber-Secure Wearable, Ultra-Low-Power Networked Sensors for Soldier Health Monitoring
by: Bozzi, Luciano, et al.
Published: (2026)
by: Bozzi, Luciano, et al.
Published: (2026)
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding
by: Zhang, Ziyi, et al.
Published: (2025)
by: Zhang, Ziyi, et al.
Published: (2025)
FedCompass: Efficient Cross-Silo Federated Learning on Heterogeneous Client Devices using a Computing Power Aware Scheduler
by: Li, Zilinghan, et al.
Published: (2023)
by: Li, Zilinghan, et al.
Published: (2023)
Fusing Depthwise and Pointwise Convolutions for Efficient Inference on GPUs
by: Qararyah, Fareed, et al.
Published: (2024)
by: Qararyah, Fareed, et al.
Published: (2024)
High-Dimensional Sparse Data Low-rank Representation via Accelerated Asynchronous Parallel Stochastic Gradient Descent
by: Hu, Qicong, et al.
Published: (2024)
by: Hu, Qicong, et al.
Published: (2024)
Detection of Global Anomalies on Distributed IoT Edges with Device-to-Device Communication
by: Ochiai, Hideya, et al.
Published: (2024)
by: Ochiai, Hideya, et al.
Published: (2024)
AMUSD: Asynchronous Multi-Device Speculative Decoding for LLM Acceleration
by: McDanel, Bradley
Published: (2024)
by: McDanel, Bradley
Published: (2024)
Heterogeneous Federated Learning with Convolutional and Spiking Neural Networks
by: Yu, Yingchao, et al.
Published: (2024)
by: Yu, Yingchao, et al.
Published: (2024)
cuConv: A CUDA Implementation of Convolution for CNN Inference
by: Jordà, Marc, et al.
Published: (2021)
by: Jordà, Marc, et al.
Published: (2021)
Distributed Convolutional Neural Network Training on Mobile and Edge Clusters
by: Rama, Pranav, et al.
Published: (2024)
by: Rama, Pranav, et al.
Published: (2024)
Optimizing Foundation Model Inference on a Many-tiny-core Open-source RISC-V Platform
by: Potocnik, Viviane, et al.
Published: (2024)
by: Potocnik, Viviane, et al.
Published: (2024)
Deep Reinforcement Learning based Online Scheduling Policy for Deep Neural Network Multi-Tenant Multi-Accelerator Systems
by: Blanco, Francesco G., et al.
Published: (2024)
by: Blanco, Francesco G., et al.
Published: (2024)
Towards Fair and Firm Real-Time Scheduling in DNN Multi-Tenant Multi-Accelerator Systems via Reinforcement Learning
by: Russo, Enrico, et al.
Published: (2024)
by: Russo, Enrico, et al.
Published: (2024)
Accelerate Intermittent Deep Inference
by: Zhang, Ziliang
Published: (2024)
by: Zhang, Ziliang
Published: (2024)
HW-SW Optimization of DNNs for Privacy-preserving People Counting on Low-resolution Infrared Arrays
by: Risso, Matteo, et al.
Published: (2024)
by: Risso, Matteo, et al.
Published: (2024)
CUDA Kernel Optimization and Counter-Free Performance Analysis for Depthwise Convolution in Cloud Environments
by: Babak, Huriyeh, et al.
Published: (2026)
by: Babak, Huriyeh, et al.
Published: (2026)
Efficient Distributed Learning over Decentralized Networks with Convoluted Support Vector Machine
by: Chen, Canyi, et al.
Published: (2025)
by: Chen, Canyi, et al.
Published: (2025)
Adaptive Stream Processing on Edge Devices through Active Inference
by: Sedlak, Boris, et al.
Published: (2024)
by: Sedlak, Boris, et al.
Published: (2024)
Device Scheduling and Assignment in Hierarchical Federated Learning for Internet of Things
by: Zhang, Tinghao, et al.
Published: (2024)
by: Zhang, Tinghao, et al.
Published: (2024)
A Robust Federated Learning Framework for Undependable Devices at Scale
by: Wang, Shilong, et al.
Published: (2024)
by: Wang, Shilong, et al.
Published: (2024)
Kraken: Inherently Parallel Transformers For Efficient Multi-Device Inference
by: Prabhakar, Rohan Baskar, et al.
Published: (2024)
by: Prabhakar, Rohan Baskar, et al.
Published: (2024)
Metadata-Guided Adaptable Frequency Scaling across Heterogeneous Applications and Devices
by: Yan, Jinqi, et al.
Published: (2025)
by: Yan, Jinqi, et al.
Published: (2025)
QSync: Quantization-Minimized Synchronous Distributed Training Across Hybrid Devices
by: Zhao, Juntao, et al.
Published: (2024)
by: Zhao, Juntao, et al.
Published: (2024)
Communication-Efficient Device Scheduling for Federated Learning Using Lyapunov Optimization
by: Perazzone, Jake B., et al.
Published: (2025)
by: Perazzone, Jake B., et al.
Published: (2025)
NEST: Network- and Memory-Aware Device Placement For Distributed Deep Learning
by: Wang, Irene, et al.
Published: (2026)
by: Wang, Irene, et al.
Published: (2026)
Lumos: Heterogeneity-aware Federated Graph Learning over Decentralized Devices
by: Pan, Qiying, et al.
Published: (2023)
by: Pan, Qiying, et al.
Published: (2023)
DOPPLER: Dual-Policy Learning for Device Assignment in Asynchronous Dataflow Graphs
by: Yao, Xinyu, et al.
Published: (2025)
by: Yao, Xinyu, et al.
Published: (2025)
Tackling Intertwined Data and Device Heterogeneities in Federated Learning with Unlimited Staleness
by: Wang, Haoming, et al.
Published: (2023)
by: Wang, Haoming, et al.
Published: (2023)
Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference
by: Zhang, Haolin, et al.
Published: (2025)
by: Zhang, Haolin, et al.
Published: (2025)
Similar Items
-
Lightweight Software Kernels and Hardware Extensions for Efficient Sparse Deep Neural Networks on Microcontrollers
by: Daghero, Francesco, et al.
Published: (2025) -
MATCHA: Efficient Deployment of Deep Neural Networks on Multi-Accelerator Heterogeneous Edge SoCs
by: Russo, Enrico, et al.
Published: (2026) -
Optimizing the Deployment of Tiny Transformers on Low-Power MCUs
by: Jung, Victor J. B., et al.
Published: (2024) -
MATCH: Model-Aware TVM-based Compilation for Heterogeneous Edge Devices
by: Hamdi, Mohamed Amine, et al.
Published: (2024) -
EDEA: Efficient Dual-Engine Accelerator for Depthwise Separable Convolution with Direct Data Transfer
by: Chen, Yi, et al.
Published: (2025)