A Flexible Template for Edge Generative AI with High-Accuracy Accelerated Softmax & GELU
Fuente:
arXiv
Saved in:
| Main Authors: | Belano, Andrea, Tortorella, Yvan, Garofalo, Angelo, Benini, Luca, Rossi, Davide, Conti, Francesco |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VEXP: A Low-Cost RISC-V ISA Extension for Accelerated Softmax Computation in Transformers
by: Wang, Run, et al.
Published: (2025)
by: Wang, Run, et al.
Published: (2025)
Hybrid Modular Redundancy: Exploring Modular Redundancy Approaches in RISC-V Multi-Core Computing Clusters for Reliable Processing in Space
by: Rogenmoser, Michael, et al.
Published: (2023)
by: Rogenmoser, Michael, et al.
Published: (2023)
Open-Source Heterogeneous SoCs for AI: The PULP Platform Experience
by: Conti, Francesco, et al.
Published: (2024)
by: Conti, Francesco, et al.
Published: (2024)
FractalSync: Lightweight Scalable Global Synchronization of Massive Bulk Synchronous Parallel AI Accelerators
by: Isachi, Victor, et al.
Published: (2025)
by: Isachi, Victor, et al.
Published: (2025)
CHIMERA: A Flexible and Scalable 3.1 TOPS/W AI-MCU with Transformer Accelerator and 563 Gb/s Shared-L2 Memory Subsystem with QoS Guarantees
by: Leone, Lorenzo, et al.
Published: (2026)
by: Leone, Lorenzo, et al.
Published: (2026)
HULK-V: a Heterogeneous Ultra-low-power Linux capable RISC-V SoC
by: Valente, Luca, et al.
Published: (2022)
by: Valente, Luca, et al.
Published: (2022)
O-POPE: High-Frequency Pipelined Outer Product based GEMM acceleration with minimal buffering overhead
by: Cammarata, Danilo, et al.
Published: (2026)
by: Cammarata, Danilo, et al.
Published: (2026)
MXDOTP: A RISC-V ISA Extension for Enabling Microscaling (MX) Floating-Point Dot Products
by: İslamoğlu, Gamze, et al.
Published: (2025)
by: İslamoğlu, Gamze, et al.
Published: (2025)
RedMulE-FT: A Reconfigurable Fault-Tolerant Matrix Multiplication Engine
by: Wiese, Philip, et al.
Published: (2025)
by: Wiese, Philip, et al.
Published: (2025)
ITA: An Energy-Efficient Attention and Softmax Accelerator for Quantized Transformers
by: İslamoğlu, Gamze, et al.
Published: (2023)
by: İslamoğlu, Gamze, et al.
Published: (2023)
Towards Reliable Systems: A Scalable Approach to AXI4 Transaction Monitoring
by: Liang, Chaoqun, et al.
Published: (2025)
by: Liang, Chaoqun, et al.
Published: (2025)
relOBI: A Reliable Low-latency Interconnect for Tightly-Coupled On-chip Communication
by: Rogenmoser, Michael, et al.
Published: (2025)
by: Rogenmoser, Michael, et al.
Published: (2025)
A Gigabit, DMA-enhanced Open-Source Ethernet Controller for Mixed-Criticality Systems
by: Liang, Chaoqun, et al.
Published: (2024)
by: Liang, Chaoqun, et al.
Published: (2024)
VMXDOTP: A RISC-V Vector ISA Extension for Efficient Microscaling (MX) Format Acceleration
by: Wipfli, Max, et al.
Published: (2026)
by: Wipfli, Max, et al.
Published: (2026)
Reusing Softmax Hardware Unit for GELU Computation in Transformers
by: Peltekis, Christodoulos, et al.
Published: (2024)
by: Peltekis, Christodoulos, et al.
Published: (2024)
Maestro: A 302 GFLOPS/W and 19.8GFLOPS RISC-V Vector-Tensor Architecture for Wearable Ultrasound Edge Computing
by: Sinigaglia, Mattia, et al.
Published: (2025)
by: Sinigaglia, Mattia, et al.
Published: (2025)
AXI-REALM: Safe, Modular and Lightweight Traffic Monitoring and Regulation for Heterogeneous Mixed-Criticality Systems
by: Benz, Thomas, et al.
Published: (2025)
by: Benz, Thomas, et al.
Published: (2025)
A Reliable, Time-Predictable Heterogeneous SoC for AI-Enhanced Mixed-Criticality Edge Applications
by: Garofalo, Angelo, et al.
Published: (2025)
by: Garofalo, Angelo, et al.
Published: (2025)
CVA6-VMRT: A Modular Approach Towards Time-Predictable Virtual Memory in a 64-bit Application Class RISC-V Processor
by: Reinwardt, Christopher, et al.
Published: (2025)
by: Reinwardt, Christopher, et al.
Published: (2025)
TOP: Towards Open & Predictable Heterogeneous SoCs
by: Valente, Luca, et al.
Published: (2024)
by: Valente, Luca, et al.
Published: (2024)
Who Checks the Checker? Enhancing Component-level Architectural SEU Fault Tolerance for End-to-End SoC Protection
by: Rogenmoser, Michael, et al.
Published: (2026)
by: Rogenmoser, Michael, et al.
Published: (2026)
Not All Faults Are Equal: Transient-Fault Sensitivity Characterization of an Open-Source RISC-V Vector Cluster
by: Cai, Maoyuan, et al.
Published: (2026)
by: Cai, Maoyuan, et al.
Published: (2026)
Quadrilatero: A RISC-V programmable matrix coprocessor for low-power edge applications
by: Cammarata, Danilo, et al.
Published: (2025)
by: Cammarata, Danilo, et al.
Published: (2025)
A Multicast-Capable AXI Crossbar for Many-core Machine Learning Accelerators
by: Colagrande, Luca, et al.
Published: (2025)
by: Colagrande, Luca, et al.
Published: (2025)
TrainDeeploy: Hardware-Accelerated Parameter-Efficient Fine-Tuning of Small Transformer Models at the Extreme Edge
by: Wang, Run, et al.
Published: (2026)
by: Wang, Run, et al.
Published: (2026)
vCLIC: Towards Fast Interrupt Handling in Virtualized RISC-V Mixed-criticality Systems
by: Zelioli, Enrico, et al.
Published: (2024)
by: Zelioli, Enrico, et al.
Published: (2024)
MATCHA: Efficient Deployment of Deep Neural Networks on Multi-Accelerator Heterogeneous Edge SoCs
by: Russo, Enrico, et al.
Published: (2026)
by: Russo, Enrico, et al.
Published: (2026)
SentryCore: A RISC-V Co-Processor System for Safe, Real-Time Control Applications
by: Rogenmoser, Michael, et al.
Published: (2024)
by: Rogenmoser, Michael, et al.
Published: (2024)
Fused-Tiled Layers: Minimizing Data Movement on RISC-V SoCs with Software-Managed Caches
by: Jung, Victor J. B., et al.
Published: (2025)
by: Jung, Victor J. B., et al.
Published: (2025)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
Circuits and Systems for Embodied AI: Exploring uJ Multi-Modal Perception for Nano-UAVs on the Kraken Shield
by: Potocnik, Viviane, et al.
Published: (2024)
by: Potocnik, Viviane, et al.
Published: (2024)
Distributed Inference with Minimal Off-Chip Traffic for Transformers on Low-Power MCUs
by: Bochem, Severin, et al.
Published: (2024)
by: Bochem, Severin, et al.
Published: (2024)
SpikeStream: Accelerating Spiking Neural Network Inference on RISC-V Clusters with Sparse Computation Extensions
by: Manoni, Simone, et al.
Published: (2025)
by: Manoni, Simone, et al.
Published: (2025)
Spatzformer: An Efficient Reconfigurable Dual-Core RISC-V V Cluster for Mixed Scalar-Vector Workloads
by: Perotti, Matteo, et al.
Published: (2024)
by: Perotti, Matteo, et al.
Published: (2024)
A Heterogeneous RISC-V based SoC for Secure Nano-UAV Navigation
by: Valente, Luca, et al.
Published: (2024)
by: Valente, Luca, et al.
Published: (2024)
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Large Attention-Based Model Inference on Tile-Based Accelerators
by: Zhang, Chi, et al.
Published: (2026)
by: Zhang, Chi, et al.
Published: (2026)
ControlPULP: A RISC-V On-Chip Parallel Power Controller for Many-Core HPC Processors with FPGA-Based Hardware-In-The-Loop Power and Thermal Emulation
by: Ottaviano, Alessandro, et al.
Published: (2023)
by: Ottaviano, Alessandro, et al.
Published: (2023)
ControlPULPlet: A Flexible Real-time Multi-core RISC-V Controller for 2.5D Systems-in-package
by: Ottaviano, Alessandro, et al.
Published: (2024)
by: Ottaviano, Alessandro, et al.
Published: (2024)
Late Breaking Results: Boosting Efficient Dual-Issue Execution on Lightweight RISC-V Cores
by: Colagrande, Luca, et al.
Published: (2026)
by: Colagrande, Luca, et al.
Published: (2026)
Dual-Issue Execution of Mixed Integer and Floating-Point Workloads on Energy-Efficient In-Order RISC-V Cores
by: Colagrande, Luca, et al.
Published: (2025)
by: Colagrande, Luca, et al.
Published: (2025)
Similar Items
-
VEXP: A Low-Cost RISC-V ISA Extension for Accelerated Softmax Computation in Transformers
by: Wang, Run, et al.
Published: (2025) -
Hybrid Modular Redundancy: Exploring Modular Redundancy Approaches in RISC-V Multi-Core Computing Clusters for Reliable Processing in Space
by: Rogenmoser, Michael, et al.
Published: (2023) -
Open-Source Heterogeneous SoCs for AI: The PULP Platform Experience
by: Conti, Francesco, et al.
Published: (2024) -
FractalSync: Lightweight Scalable Global Synchronization of Massive Bulk Synchronous Parallel AI Accelerators
by: Isachi, Victor, et al.
Published: (2025) -
CHIMERA: A Flexible and Scalable 3.1 TOPS/W AI-MCU with Transformer Accelerator and 563 Gb/s Shared-L2 Memory Subsystem with QoS Guarantees
by: Leone, Lorenzo, et al.
Published: (2026)