Distributed Inference with Minimal Off-Chip Traffic for Transformers on Low-Power MCUs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bochem, Severin, Jung, Victor J. B., Prasad, Arpan, Conti, Francesco, Benini, Luca |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Fused-Tiled Layers: Minimizing Data Movement on RISC-V SoCs with Software-Managed Caches
von: Jung, Victor J. B., et al.
Veröffentlicht: (2025)
von: Jung, Victor J. B., et al.
Veröffentlicht: (2025)
MXDOTP: A RISC-V ISA Extension for Enabling Microscaling (MX) Floating-Point Dot Products
von: İslamoğlu, Gamze, et al.
Veröffentlicht: (2025)
von: İslamoğlu, Gamze, et al.
Veröffentlicht: (2025)
TrainDeeploy: Hardware-Accelerated Parameter-Efficient Fine-Tuning of Small Transformer Models at the Extreme Edge
von: Wang, Run, et al.
Veröffentlicht: (2026)
von: Wang, Run, et al.
Veröffentlicht: (2026)
VEXP: A Low-Cost RISC-V ISA Extension for Accelerated Softmax Computation in Transformers
von: Wang, Run, et al.
Veröffentlicht: (2025)
von: Wang, Run, et al.
Veröffentlicht: (2025)
CHIMERA: A Flexible and Scalable 3.1 TOPS/W AI-MCU with Transformer Accelerator and 563 Gb/s Shared-L2 Memory Subsystem with QoS Guarantees
von: Leone, Lorenzo, et al.
Veröffentlicht: (2026)
von: Leone, Lorenzo, et al.
Veröffentlicht: (2026)
Low-Energy On-Device Personalization for MCUs
von: Huang, Yushan, et al.
Veröffentlicht: (2024)
von: Huang, Yushan, et al.
Veröffentlicht: (2024)
Circuits and Systems for Embodied AI: Exploring uJ Multi-Modal Perception for Nano-UAVs on the Kraken Shield
von: Potocnik, Viviane, et al.
Veröffentlicht: (2024)
von: Potocnik, Viviane, et al.
Veröffentlicht: (2024)
A Flexible Template for Edge Generative AI with High-Accuracy Accelerated Softmax & GELU
von: Belano, Andrea, et al.
Veröffentlicht: (2024)
von: Belano, Andrea, et al.
Veröffentlicht: (2024)
Deeploy: Enabling Energy-Efficient Deployment of Small Language Models On Heterogeneous Microcontrollers
von: Scherer, Moritz, et al.
Veröffentlicht: (2024)
von: Scherer, Moritz, et al.
Veröffentlicht: (2024)
A Dynamic Allocation Scheme for Adaptive Shared-Memory Mapping on Kilo-core RV Clusters for Attention-Based Model Deployment
von: Wang, Bowen, et al.
Veröffentlicht: (2025)
von: Wang, Bowen, et al.
Veröffentlicht: (2025)
Work-In-Progress: Accelerating Numpy With OpenBLAS For Open-Source RISC-V Chips
von: Koenig, Cyril, et al.
Veröffentlicht: (2025)
von: Koenig, Cyril, et al.
Veröffentlicht: (2025)
Toward Attention-based TinyML: A Heterogeneous Accelerated Architecture and Automated Deployment Flow
von: Wiese, Philip, et al.
Veröffentlicht: (2024)
von: Wiese, Philip, et al.
Veröffentlicht: (2024)
PELS: A Lightweight and Flexible Peripheral Event Linking System for Ultra-Low Power IoT Processors
von: Ottaviano, Alessandro, et al.
Veröffentlicht: (2023)
von: Ottaviano, Alessandro, et al.
Veröffentlicht: (2023)
ControlPULP: A RISC-V On-Chip Parallel Power Controller for Many-Core HPC Processors with FPGA-Based Hardware-In-The-Loop Power and Thermal Emulation
von: Ottaviano, Alessandro, et al.
Veröffentlicht: (2023)
von: Ottaviano, Alessandro, et al.
Veröffentlicht: (2023)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Large Attention-Based Model Inference on Tile-Based Accelerators
von: Zhang, Chi, et al.
Veröffentlicht: (2026)
von: Zhang, Chi, et al.
Veröffentlicht: (2026)
relOBI: A Reliable Low-latency Interconnect for Tightly-Coupled On-chip Communication
von: Rogenmoser, Michael, et al.
Veröffentlicht: (2025)
von: Rogenmoser, Michael, et al.
Veröffentlicht: (2025)
Siracusa: A 16 nm Heterogenous RISC-V SoC for Extended Reality with At-MRAM Neural Engine
von: Prasad, Arpan Suravi, et al.
Veröffentlicht: (2023)
von: Prasad, Arpan Suravi, et al.
Veröffentlicht: (2023)
Hybrid Modular Redundancy: Exploring Modular Redundancy Approaches in RISC-V Multi-Core Computing Clusters for Reliable Processing in Space
von: Rogenmoser, Michael, et al.
Veröffentlicht: (2023)
von: Rogenmoser, Michael, et al.
Veröffentlicht: (2023)
A Multicast-Capable AXI Crossbar for Many-core Machine Learning Accelerators
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
Late Breaking Results: Boosting Efficient Dual-Issue Execution on Lightweight RISC-V Cores
von: Colagrande, Luca, et al.
Veröffentlicht: (2026)
von: Colagrande, Luca, et al.
Veröffentlicht: (2026)
Dual-Issue Execution of Mixed Integer and Floating-Point Workloads on Energy-Efficient In-Order RISC-V Cores
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
Improving Chip Design Enablement for Universities in Europe -- A Position Paper
von: Krupp, Lukas, et al.
Veröffentlicht: (2025)
von: Krupp, Lukas, et al.
Veröffentlicht: (2025)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
vMCU: Coordinated Memory Management and Kernel Optimization for DNN Inference on MCUs
von: Zheng, Size, et al.
Veröffentlicht: (2024)
von: Zheng, Size, et al.
Veröffentlicht: (2024)
AXI-REALM: Safe, Modular and Lightweight Traffic Monitoring and Regulation for Heterogeneous Mixed-Criticality Systems
von: Benz, Thomas, et al.
Veröffentlicht: (2025)
von: Benz, Thomas, et al.
Veröffentlicht: (2025)
TOP: Towards Open & Predictable Heterogeneous SoCs
von: Valente, Luca, et al.
Veröffentlicht: (2024)
von: Valente, Luca, et al.
Veröffentlicht: (2024)
ITA: An Energy-Efficient Attention and Softmax Accelerator for Quantized Transformers
von: İslamoğlu, Gamze, et al.
Veröffentlicht: (2023)
von: İslamoğlu, Gamze, et al.
Veröffentlicht: (2023)
Trikarenos: A Fault-Tolerant RISC-V-based Microcontroller for CubeSats in 28nm
von: Rogenmoser, Michael, et al.
Veröffentlicht: (2023)
von: Rogenmoser, Michael, et al.
Veröffentlicht: (2023)
SpikeStream: Accelerating Spiking Neural Network Inference on RISC-V Clusters with Sparse Computation Extensions
von: Manoni, Simone, et al.
Veröffentlicht: (2025)
von: Manoni, Simone, et al.
Veröffentlicht: (2025)
TROOP: At-the-Roofline Performance for Vector Processors on Low Operational Intensity Workloads
von: Purayil, Navaneeth Kunhi, et al.
Veröffentlicht: (2025)
von: Purayil, Navaneeth Kunhi, et al.
Veröffentlicht: (2025)
MiniFloat-NN and ExSdotp: An ISA Extension and a Modular Open Hardware Unit for Low-Precision Training on RISC-V cores
von: Bertaccini, Luca, et al.
Veröffentlicht: (2022)
von: Bertaccini, Luca, et al.
Veröffentlicht: (2022)
Late Breaking Results: A RISC-V ISA Extension for Chaining in Scalar Processors
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
Ultra Low-Power SDM-based Circuit-Switching for Networks-on-Chip
von: Zaeemi, Meysam, et al.
Veröffentlicht: (2026)
von: Zaeemi, Meysam, et al.
Veröffentlicht: (2026)
SALSA: Simulated Annealing based Loop-Ordering Scheduler for DNN Accelerators
von: Jung, Victor J. B., et al.
Veröffentlicht: (2023)
von: Jung, Victor J. B., et al.
Veröffentlicht: (2023)
BioTrain: Sub-MB, Sub-50mW On-Device Fine-Tuning for Edge-AI on Biosignals
von: Wang, Run, et al.
Veröffentlicht: (2026)
von: Wang, Run, et al.
Veröffentlicht: (2026)
MX: Enhancing RISC-V's Vector ISA for Ultra-Low Overhead, Energy-Efficient Matrix Multiplication
von: Perotti, Matteo, et al.
Veröffentlicht: (2024)
von: Perotti, Matteo, et al.
Veröffentlicht: (2024)
Enabling Efficient Hybrid Systolic Computation in Shared L1-Memory Manycore Clusters
von: Mazzola, Sergio, et al.
Veröffentlicht: (2024)
von: Mazzola, Sergio, et al.
Veröffentlicht: (2024)
How to keep pushing ML accelerator performance? Know your rooflines!
von: Verhelst, Marian, et al.
Veröffentlicht: (2025)
von: Verhelst, Marian, et al.
Veröffentlicht: (2025)
O-POPE: High-Frequency Pipelined Outer Product based GEMM acceleration with minimal buffering overhead
von: Cammarata, Danilo, et al.
Veröffentlicht: (2026)
von: Cammarata, Danilo, et al.
Veröffentlicht: (2026)
Evaluating IOMMU-Based Shared Virtual Addressing for RISC-V Embedded Heterogeneous SoCs
von: Koenig, Cyril, et al.
Veröffentlicht: (2025)
von: Koenig, Cyril, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Fused-Tiled Layers: Minimizing Data Movement on RISC-V SoCs with Software-Managed Caches
von: Jung, Victor J. B., et al.
Veröffentlicht: (2025) -
MXDOTP: A RISC-V ISA Extension for Enabling Microscaling (MX) Floating-Point Dot Products
von: İslamoğlu, Gamze, et al.
Veröffentlicht: (2025) -
TrainDeeploy: Hardware-Accelerated Parameter-Efficient Fine-Tuning of Small Transformer Models at the Extreme Edge
von: Wang, Run, et al.
Veröffentlicht: (2026) -
VEXP: A Low-Cost RISC-V ISA Extension for Accelerated Softmax Computation in Transformers
von: Wang, Run, et al.
Veröffentlicht: (2025) -
CHIMERA: A Flexible and Scalable 3.1 TOPS/W AI-MCU with Transformer Accelerator and 563 Gb/s Shared-L2 Memory Subsystem with QoS Guarantees
von: Leone, Lorenzo, et al.
Veröffentlicht: (2026)