Deeploy: Enabling Energy-Efficient Deployment of Small Language Models On Heterogeneous Microcontrollers
Fuente:
arXiv
Saved in:
| Main Authors: | Scherer, Moritz, Macan, Luka, Jung, Victor, Wiese, Philip, Bompani, Luca, Burrello, Alessio, Conti, Francesco, Benini, Luca |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Toward Attention-based TinyML: A Heterogeneous Accelerated Architecture and Automated Deployment Flow
by: Wiese, Philip, et al.
Published: (2024)
by: Wiese, Philip, et al.
Published: (2024)
TrainDeeploy: Hardware-Accelerated Parameter-Efficient Fine-Tuning of Small Transformer Models at the Extreme Edge
by: Wang, Run, et al.
Published: (2026)
by: Wang, Run, et al.
Published: (2026)
Fused-Tiled Layers: Minimizing Data Movement on RISC-V SoCs with Software-Managed Caches
by: Jung, Victor J. B., et al.
Published: (2025)
by: Jung, Victor J. B., et al.
Published: (2025)
BioTrain: Sub-MB, Sub-50mW On-Device Fine-Tuning for Edge-AI on Biosignals
by: Wang, Run, et al.
Published: (2026)
by: Wang, Run, et al.
Published: (2026)
Trikarenos: A Fault-Tolerant RISC-V-based Microcontroller for CubeSats in 28nm
by: Rogenmoser, Michael, et al.
Published: (2023)
by: Rogenmoser, Michael, et al.
Published: (2023)
MATCHA: Efficient Deployment of Deep Neural Networks on Multi-Accelerator Heterogeneous Edge SoCs
by: Russo, Enrico, et al.
Published: (2026)
by: Russo, Enrico, et al.
Published: (2026)
Circuits and Systems for Embodied AI: Exploring uJ Multi-Modal Perception for Nano-UAVs on the Kraken Shield
by: Potocnik, Viviane, et al.
Published: (2024)
by: Potocnik, Viviane, et al.
Published: (2024)
xTern: Energy-Efficient Ternary Neural Network Inference on RISC-V-Based Edge Systems
by: Rutishauser, Georg, et al.
Published: (2024)
by: Rutishauser, Georg, et al.
Published: (2024)
CHIMERA: A Flexible and Scalable 3.1 TOPS/W AI-MCU with Transformer Accelerator and 563 Gb/s Shared-L2 Memory Subsystem with QoS Guarantees
by: Leone, Lorenzo, et al.
Published: (2026)
by: Leone, Lorenzo, et al.
Published: (2026)
Dual-Issue Execution of Mixed Integer and Floating-Point Workloads on Energy-Efficient In-Order RISC-V Cores
by: Colagrande, Luca, et al.
Published: (2025)
by: Colagrande, Luca, et al.
Published: (2025)
ITA: An Energy-Efficient Attention and Softmax Accelerator for Quantized Transformers
by: İslamoğlu, Gamze, et al.
Published: (2023)
by: İslamoğlu, Gamze, et al.
Published: (2023)
Distributed Inference with Minimal Off-Chip Traffic for Transformers on Low-Power MCUs
by: Bochem, Severin, et al.
Published: (2024)
by: Bochem, Severin, et al.
Published: (2024)
MXDOTP: A RISC-V ISA Extension for Enabling Microscaling (MX) Floating-Point Dot Products
by: İslamoğlu, Gamze, et al.
Published: (2025)
by: İslamoğlu, Gamze, et al.
Published: (2025)
TOP: Towards Open & Predictable Heterogeneous SoCs
by: Valente, Luca, et al.
Published: (2024)
by: Valente, Luca, et al.
Published: (2024)
Late Breaking Results: Boosting Efficient Dual-Issue Execution on Lightweight RISC-V Cores
by: Colagrande, Luca, et al.
Published: (2026)
by: Colagrande, Luca, et al.
Published: (2026)
Enabling Efficient Hybrid Systolic Computation in Shared L1-Memory Manycore Clusters
by: Mazzola, Sergio, et al.
Published: (2024)
by: Mazzola, Sergio, et al.
Published: (2024)
A Dynamic Allocation Scheme for Adaptive Shared-Memory Mapping on Kilo-core RV Clusters for Attention-Based Model Deployment
by: Wang, Bowen, et al.
Published: (2025)
by: Wang, Bowen, et al.
Published: (2025)
Towards Zero-Stall Matrix Multiplication on Energy-Efficient RISC-V Clusters for Machine Learning Acceleration
by: Colagrande, Luca, et al.
Published: (2025)
by: Colagrande, Luca, et al.
Published: (2025)
A Multicast-Capable AXI Crossbar for Many-core Machine Learning Accelerators
by: Colagrande, Luca, et al.
Published: (2025)
by: Colagrande, Luca, et al.
Published: (2025)
SARIS: Accelerating Stencil Computations on Energy-Efficient RISC-V Compute Clusters with Indirect Stream Registers
by: Scheffler, Paul, et al.
Published: (2024)
by: Scheffler, Paul, et al.
Published: (2024)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
Sub-Millisecond Event-Based Eye Tracking on a Resource-Constrained Microcontroller
by: Giordano, Marco, et al.
Published: (2025)
by: Giordano, Marco, et al.
Published: (2025)
Evaluating IOMMU-Based Shared Virtual Addressing for RISC-V Embedded Heterogeneous SoCs
by: Koenig, Cyril, et al.
Published: (2025)
by: Koenig, Cyril, et al.
Published: (2025)
Siracusa: A 16 nm Heterogenous RISC-V SoC for Extended Reality with At-MRAM Neural Engine
by: Prasad, Arpan Suravi, et al.
Published: (2023)
by: Prasad, Arpan Suravi, et al.
Published: (2023)
Optimizing Offload Performance in Heterogeneous MPSoCs
by: Colagrande, Luca, et al.
Published: (2024)
by: Colagrande, Luca, et al.
Published: (2024)
A Flexible Template for Edge Generative AI with High-Accuracy Accelerated Softmax & GELU
by: Belano, Andrea, et al.
Published: (2024)
by: Belano, Andrea, et al.
Published: (2024)
Open-Source Heterogeneous SoCs for AI: The PULP Platform Experience
by: Conti, Francesco, et al.
Published: (2024)
by: Conti, Francesco, et al.
Published: (2024)
A Survey on Design Methodologies for Accelerating Deep Learning on Heterogeneous Architectures
by: Curzel, Serena, et al.
Published: (2023)
by: Curzel, Serena, et al.
Published: (2023)
Late Breaking Results: A RISC-V ISA Extension for Chaining in Scalar Processors
by: Colagrande, Luca, et al.
Published: (2025)
by: Colagrande, Luca, et al.
Published: (2025)
LRSCwait: Enabling Scalable and Efficient Synchronization in Manycore Systems through Polling-Free and Retry-Free Operation
by: Riedel, Samuel, et al.
Published: (2024)
by: Riedel, Samuel, et al.
Published: (2024)
Hybrid Modular Redundancy: Exploring Modular Redundancy Approaches in RISC-V Multi-Core Computing Clusters for Reliable Processing in Space
by: Rogenmoser, Michael, et al.
Published: (2023)
by: Rogenmoser, Michael, et al.
Published: (2023)
MX: Enhancing RISC-V's Vector ISA for Ultra-Low Overhead, Energy-Efficient Matrix Multiplication
by: Perotti, Matteo, et al.
Published: (2024)
by: Perotti, Matteo, et al.
Published: (2024)
HULK-V: a Heterogeneous Ultra-low-power Linux capable RISC-V SoC
by: Valente, Luca, et al.
Published: (2022)
by: Valente, Luca, et al.
Published: (2022)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Large Attention-Based Model Inference on Tile-Based Accelerators
by: Zhang, Chi, et al.
Published: (2026)
by: Zhang, Chi, et al.
Published: (2026)
relOBI: A Reliable Low-latency Interconnect for Tightly-Coupled On-chip Communication
by: Rogenmoser, Michael, et al.
Published: (2025)
by: Rogenmoser, Michael, et al.
Published: (2025)
How to keep pushing ML accelerator performance? Know your rooflines!
by: Verhelst, Marian, et al.
Published: (2025)
by: Verhelst, Marian, et al.
Published: (2025)
O-POPE: High-Frequency Pipelined Outer Product based GEMM acceleration with minimal buffering overhead
by: Cammarata, Danilo, et al.
Published: (2026)
by: Cammarata, Danilo, et al.
Published: (2026)
AraXL: A Physically Scalable, Ultra-Wide RISC-V Vector Processor Design for Fast and Efficient Computation on Long Vectors
by: Purayil, Navaneeth Kunhi, et al.
Published: (2025)
by: Purayil, Navaneeth Kunhi, et al.
Published: (2025)
MiniFloat-NN and ExSdotp: An ISA Extension and a Modular Open Hardware Unit for Low-Precision Training on RISC-V cores
by: Bertaccini, Luca, et al.
Published: (2022)
by: Bertaccini, Luca, et al.
Published: (2022)
AXI-REALM: Safe, Modular and Lightweight Traffic Monitoring and Regulation for Heterogeneous Mixed-Criticality Systems
by: Benz, Thomas, et al.
Published: (2025)
by: Benz, Thomas, et al.
Published: (2025)
Similar Items
-
Toward Attention-based TinyML: A Heterogeneous Accelerated Architecture and Automated Deployment Flow
by: Wiese, Philip, et al.
Published: (2024) -
TrainDeeploy: Hardware-Accelerated Parameter-Efficient Fine-Tuning of Small Transformer Models at the Extreme Edge
by: Wang, Run, et al.
Published: (2026) -
Fused-Tiled Layers: Minimizing Data Movement on RISC-V SoCs with Software-Managed Caches
by: Jung, Victor J. B., et al.
Published: (2025) -
BioTrain: Sub-MB, Sub-50mW On-Device Fine-Tuning for Edge-AI on Biosignals
by: Wang, Run, et al.
Published: (2026) -
Trikarenos: A Fault-Tolerant RISC-V-based Microcontroller for CubeSats in 28nm
by: Rogenmoser, Michael, et al.
Published: (2023)