Toward Attention-based TinyML: A Heterogeneous Accelerated Architecture and Automated Deployment Flow
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wiese, Philip, İslamoğlu, Gamze, Scherer, Moritz, Macan, Luka, Jung, Victor J. B., Burrello, Alessio, Conti, Francesco, Benini, Luca |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Deeploy: Enabling Energy-Efficient Deployment of Small Language Models On Heterogeneous Microcontrollers
von: Scherer, Moritz, et al.
Veröffentlicht: (2024)
von: Scherer, Moritz, et al.
Veröffentlicht: (2024)
CHIMERA: A Flexible and Scalable 3.1 TOPS/W AI-MCU with Transformer Accelerator and 563 Gb/s Shared-L2 Memory Subsystem with QoS Guarantees
von: Leone, Lorenzo, et al.
Veröffentlicht: (2026)
von: Leone, Lorenzo, et al.
Veröffentlicht: (2026)
ITA: An Energy-Efficient Attention and Softmax Accelerator for Quantized Transformers
von: İslamoğlu, Gamze, et al.
Veröffentlicht: (2023)
von: İslamoğlu, Gamze, et al.
Veröffentlicht: (2023)
TrainDeeploy: Hardware-Accelerated Parameter-Efficient Fine-Tuning of Small Transformer Models at the Extreme Edge
von: Wang, Run, et al.
Veröffentlicht: (2026)
von: Wang, Run, et al.
Veröffentlicht: (2026)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
Fused-Tiled Layers: Minimizing Data Movement on RISC-V SoCs with Software-Managed Caches
von: Jung, Victor J. B., et al.
Veröffentlicht: (2025)
von: Jung, Victor J. B., et al.
Veröffentlicht: (2025)
MXDOTP: A RISC-V ISA Extension for Enabling Microscaling (MX) Floating-Point Dot Products
von: İslamoğlu, Gamze, et al.
Veröffentlicht: (2025)
von: İslamoğlu, Gamze, et al.
Veröffentlicht: (2025)
VEXP: A Low-Cost RISC-V ISA Extension for Accelerated Softmax Computation in Transformers
von: Wang, Run, et al.
Veröffentlicht: (2025)
von: Wang, Run, et al.
Veröffentlicht: (2025)
HTVM: Efficient Neural Network Deployment On Heterogeneous TinyML Platforms
von: Van Delm, Josse, et al.
Veröffentlicht: (2024)
von: Van Delm, Josse, et al.
Veröffentlicht: (2024)
VMXDOTP: A RISC-V Vector ISA Extension for Efficient Microscaling (MX) Format Acceleration
von: Wipfli, Max, et al.
Veröffentlicht: (2026)
von: Wipfli, Max, et al.
Veröffentlicht: (2026)
Optimizing the Deployment of Tiny Transformers on Low-Power MCUs
von: Jung, Victor J. B., et al.
Veröffentlicht: (2024)
von: Jung, Victor J. B., et al.
Veröffentlicht: (2024)
MATCHA: Efficient Deployment of Deep Neural Networks on Multi-Accelerator Heterogeneous Edge SoCs
von: Russo, Enrico, et al.
Veröffentlicht: (2026)
von: Russo, Enrico, et al.
Veröffentlicht: (2026)
Co-Design of CNN Accelerators for TinyML using Approximate Matrix Decomposition
von: Morales, José Juan Hernández, et al.
Veröffentlicht: (2026)
von: Morales, José Juan Hernández, et al.
Veröffentlicht: (2026)
BioTrain: Sub-MB, Sub-50mW On-Device Fine-Tuning for Edge-AI on Biosignals
von: Wang, Run, et al.
Veröffentlicht: (2026)
von: Wang, Run, et al.
Veröffentlicht: (2026)
Sequential Printed MLP Circuits for Super TinyML Multi-Sensory Applications
von: Saglam, Gurol, et al.
Veröffentlicht: (2024)
von: Saglam, Gurol, et al.
Veröffentlicht: (2024)
Circuits and Systems for Embodied AI: Exploring uJ Multi-Modal Perception for Nano-UAVs on the Kraken Shield
von: Potocnik, Viviane, et al.
Veröffentlicht: (2024)
von: Potocnik, Viviane, et al.
Veröffentlicht: (2024)
A Flexible Template for Edge Generative AI with High-Accuracy Accelerated Softmax & GELU
von: Belano, Andrea, et al.
Veröffentlicht: (2024)
von: Belano, Andrea, et al.
Veröffentlicht: (2024)
A Survey on Design Methodologies for Accelerating Deep Learning on Heterogeneous Architectures
von: Curzel, Serena, et al.
Veröffentlicht: (2023)
von: Curzel, Serena, et al.
Veröffentlicht: (2023)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Large Attention-Based Model Inference on Tile-Based Accelerators
von: Zhang, Chi, et al.
Veröffentlicht: (2026)
von: Zhang, Chi, et al.
Veröffentlicht: (2026)
Wet TinyML: Chemical Neural Network Using Gene Regulation and Cell Plasticity
von: Somathilaka, Samitha, et al.
Veröffentlicht: (2024)
von: Somathilaka, Samitha, et al.
Veröffentlicht: (2024)
A Dynamic Allocation Scheme for Adaptive Shared-Memory Mapping on Kilo-core RV Clusters for Attention-Based Model Deployment
von: Wang, Bowen, et al.
Veröffentlicht: (2025)
von: Wang, Bowen, et al.
Veröffentlicht: (2025)
RISC-V Based TinyML Accelerator for Depthwise Separable Convolutions in Edge AI
von: Yildirim, Muhammed, et al.
Veröffentlicht: (2025)
von: Yildirim, Muhammed, et al.
Veröffentlicht: (2025)
Distributed Inference with Minimal Off-Chip Traffic for Transformers on Low-Power MCUs
von: Bochem, Severin, et al.
Veröffentlicht: (2024)
von: Bochem, Severin, et al.
Veröffentlicht: (2024)
xTern: Energy-Efficient Ternary Neural Network Inference on RISC-V-Based Edge Systems
von: Rutishauser, Georg, et al.
Veröffentlicht: (2024)
von: Rutishauser, Georg, et al.
Veröffentlicht: (2024)
Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators
von: Shen, Aofeng, et al.
Veröffentlicht: (2025)
von: Shen, Aofeng, et al.
Veröffentlicht: (2025)
A Multicast-Capable AXI Crossbar for Many-core Machine Learning Accelerators
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
von: Colagrande, Luca, et al.
Veröffentlicht: (2025)
TOP: Towards Open & Predictable Heterogeneous SoCs
von: Valente, Luca, et al.
Veröffentlicht: (2024)
von: Valente, Luca, et al.
Veröffentlicht: (2024)
Energy Efficient Software Hardware CoDesign for Machine Learning: From TinyML to Large Language Models
von: Vahdatpour, Mohammad Saleh, et al.
Veröffentlicht: (2026)
von: Vahdatpour, Mohammad Saleh, et al.
Veröffentlicht: (2026)
Open-Source Heterogeneous SoCs for AI: The PULP Platform Experience
von: Conti, Francesco, et al.
Veröffentlicht: (2024)
von: Conti, Francesco, et al.
Veröffentlicht: (2024)
Siracusa: A 16 nm Heterogenous RISC-V SoC for Extended Reality with At-MRAM Neural Engine
von: Prasad, Arpan Suravi, et al.
Veröffentlicht: (2023)
von: Prasad, Arpan Suravi, et al.
Veröffentlicht: (2023)
How to keep pushing ML accelerator performance? Know your rooflines!
von: Verhelst, Marian, et al.
Veröffentlicht: (2025)
von: Verhelst, Marian, et al.
Veröffentlicht: (2025)
FERMI-ML: A Flexible and Resource-Efficient Memory-In-Situ SRAM Macro for TinyML acceleration
von: Lokhande, Mukul, et al.
Veröffentlicht: (2025)
von: Lokhande, Mukul, et al.
Veröffentlicht: (2025)
Optimizing Offload Performance in Heterogeneous MPSoCs
von: Colagrande, Luca, et al.
Veröffentlicht: (2024)
von: Colagrande, Luca, et al.
Veröffentlicht: (2024)
Hybrid Modular Redundancy: Exploring Modular Redundancy Approaches in RISC-V Multi-Core Computing Clusters for Reliable Processing in Space
von: Rogenmoser, Michael, et al.
Veröffentlicht: (2023)
von: Rogenmoser, Michael, et al.
Veröffentlicht: (2023)
A Hybrid Edge Classifier: Combining TinyML-Optimised CNN with RRAM-CMOS ACAM for Energy-Efficient Inference
von: Woodward, Kieran, et al.
Veröffentlicht: (2025)
von: Woodward, Kieran, et al.
Veröffentlicht: (2025)
TSB: Tiny Shared Block for Efficient DNN Deployment on NVCIM Accelerators
von: Qin, Yifan, et al.
Veröffentlicht: (2024)
von: Qin, Yifan, et al.
Veröffentlicht: (2024)
Evaluating IOMMU-Based Shared Virtual Addressing for RISC-V Embedded Heterogeneous SoCs
von: Koenig, Cyril, et al.
Veröffentlicht: (2025)
von: Koenig, Cyril, et al.
Veröffentlicht: (2025)
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
von: Colagrande, Luca, et al.
Veröffentlicht: (2026)
von: Colagrande, Luca, et al.
Veröffentlicht: (2026)
TinyFormer: Efficient Transformer Design and Deployment on Tiny Devices
von: Yang, Jianlei, et al.
Veröffentlicht: (2023)
von: Yang, Jianlei, et al.
Veröffentlicht: (2023)
SARIS: Accelerating Stencil Computations on Energy-Efficient RISC-V Compute Clusters with Indirect Stream Registers
von: Scheffler, Paul, et al.
Veröffentlicht: (2024)
von: Scheffler, Paul, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Deeploy: Enabling Energy-Efficient Deployment of Small Language Models On Heterogeneous Microcontrollers
von: Scherer, Moritz, et al.
Veröffentlicht: (2024) -
CHIMERA: A Flexible and Scalable 3.1 TOPS/W AI-MCU with Transformer Accelerator and 563 Gb/s Shared-L2 Memory Subsystem with QoS Guarantees
von: Leone, Lorenzo, et al.
Veröffentlicht: (2026) -
ITA: An Energy-Efficient Attention and Softmax Accelerator for Quantized Transformers
von: İslamoğlu, Gamze, et al.
Veröffentlicht: (2023) -
TrainDeeploy: Hardware-Accelerated Parameter-Efficient Fine-Tuning of Small Transformer Models at the Extreme Edge
von: Wang, Run, et al.
Veröffentlicht: (2026) -
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators
von: Zhang, Chi, et al.
Veröffentlicht: (2025)