Taming the Exponential: A Fast Softmax Surrogate for Integer-Native Edge Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Danopoulos, Dimitrios, Lupi, Enrico, Kagan, Michael, Pierini, Maurizio |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AIE4ML: An End-to-End Framework for Compiling Neural Networks for the Next Generation of AMD AI Engines
von: Danopoulos, Dimitrios, et al.
Veröffentlicht: (2025)
von: Danopoulos, Dimitrios, et al.
Veröffentlicht: (2025)
TransAxx: Efficient Transformers with Approximate Computing
von: Danopoulos, Dimitrios, et al.
Veröffentlicht: (2024)
von: Danopoulos, Dimitrios, et al.
Veröffentlicht: (2024)
Design Rules for Extreme-Edge Scientific Computing on AI Engines
von: Ma, Zhenghua, et al.
Veröffentlicht: (2026)
von: Ma, Zhenghua, et al.
Veröffentlicht: (2026)
Hardware-Efficient Softmax and Layer Normalization with Guaranteed Normalization for Edge Devices
von: Choi, Dawon, et al.
Veröffentlicht: (2026)
von: Choi, Dawon, et al.
Veröffentlicht: (2026)
SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
von: Wang, Wenxun, et al.
Veröffentlicht: (2025)
von: Wang, Wenxun, et al.
Veröffentlicht: (2025)
A Data-Driven Approach to Dataflow-Aware Online Scheduling for Graph Neural Network Inference
von: Puigdemont, Pol, et al.
Veröffentlicht: (2024)
von: Puigdemont, Pol, et al.
Veröffentlicht: (2024)
Reusing Softmax Hardware Unit for GELU Computation in Transformers
von: Peltekis, Christodoulos, et al.
Veröffentlicht: (2024)
von: Peltekis, Christodoulos, et al.
Veröffentlicht: (2024)
ITA: An Energy-Efficient Attention and Softmax Accelerator for Quantized Transformers
von: İslamoğlu, Gamze, et al.
Veröffentlicht: (2023)
von: İslamoğlu, Gamze, et al.
Veröffentlicht: (2023)
VEXP: A Low-Cost RISC-V ISA Extension for Accelerated Softmax Computation in Transformers
von: Wang, Run, et al.
Veröffentlicht: (2025)
von: Wang, Run, et al.
Veröffentlicht: (2025)
KLLM: Fast LLM Inference with K-Means Quantization
von: Wu, Xueying, et al.
Veröffentlicht: (2025)
von: Wu, Xueying, et al.
Veröffentlicht: (2025)
PolyThrottle: Energy-efficient Neural Network Inference on Edge Devices
von: Yan, Minghao, et al.
Veröffentlicht: (2023)
von: Yan, Minghao, et al.
Veröffentlicht: (2023)
MaRVIn: A Cross-Layer Mixed-Precision RISC-V Framework for DNN Inference, from ISA Extension to Hardware Acceleration
von: Armeniakos, Giorgos, et al.
Veröffentlicht: (2025)
von: Armeniakos, Giorgos, et al.
Veröffentlicht: (2025)
SHIELD: A Segmented Hierarchical Memory Architecture for Energy-Efficient LLM Inference on Edge NPUs
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
von: Zhang, Jintao, et al.
Veröffentlicht: (2026)
SigmaQuant: Hardware-Aware Heterogeneous Quantization Method for Edge DNN Inference
von: Liu, Qunyou, et al.
Veröffentlicht: (2026)
von: Liu, Qunyou, et al.
Veröffentlicht: (2026)
COBRA: Algorithm-Architecture Co-optimized Binary Transformer Accelerator for Edge Inference
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
CHIME: Chiplet-based Heterogeneous Near-Memory Acceleration for Edge Multimodal LLM Inference
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
xTern: Energy-Efficient Ternary Neural Network Inference on RISC-V-Based Edge Systems
von: Rutishauser, Georg, et al.
Veröffentlicht: (2024)
von: Rutishauser, Georg, et al.
Veröffentlicht: (2024)
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
von: Chen, Yuzong, et al.
Veröffentlicht: (2025)
von: Chen, Yuzong, et al.
Veröffentlicht: (2025)
FLASH-D: FlashAttention with Hidden Softmax Division
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2025)
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2025)
A PPA-Driven 3D-IC Partitioning Selection Framework with Surrogate Models
von: Wang, Shang, et al.
Veröffentlicht: (2026)
von: Wang, Shang, et al.
Veröffentlicht: (2026)
From Physics to Surrogate Intelligence: A Unified Electro-Thermo-Optimization Framework for TSV Networks
von: Gharib, Mohamed, et al.
Veröffentlicht: (2026)
von: Gharib, Mohamed, et al.
Veröffentlicht: (2026)
Rescaling-Aware Training for Efficient Deployment of Deep Learning Models on Full-Integer Hardware
von: Mueller, Lion, et al.
Veröffentlicht: (2025)
von: Mueller, Lion, et al.
Veröffentlicht: (2025)
ConSmax: Hardware-Friendly Alternative Softmax with Learnable Parameters
von: Liu, Shiwei, et al.
Veröffentlicht: (2024)
von: Liu, Shiwei, et al.
Veröffentlicht: (2024)
Low-Cost FlashAttention with Fused Exponential and Multiplication Hardware Operators
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2025)
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2025)
LL-GNN: Low Latency Graph Neural Networks on FPGAs for High Energy Physics
von: Que, Zhiqiang, et al.
Veröffentlicht: (2022)
von: Que, Zhiqiang, et al.
Veröffentlicht: (2022)
Accelerating PoT Quantization on Edge Devices
von: Saha, Rappy, et al.
Veröffentlicht: (2024)
von: Saha, Rappy, et al.
Veröffentlicht: (2024)
Designing Efficient LLM Accelerators for Edge Devices
von: Haris, Jude, et al.
Veröffentlicht: (2024)
von: Haris, Jude, et al.
Veröffentlicht: (2024)
Runtime Tunable Tsetlin Machines for Edge Inference on eFPGAs
von: Rahman, Tousif, et al.
Veröffentlicht: (2025)
von: Rahman, Tousif, et al.
Veröffentlicht: (2025)
Exploiting temporal parallelism for LSTM Autoencoder acceleration on FPGA
von: Leftheriotis, Aimilios, et al.
Veröffentlicht: (2026)
von: Leftheriotis, Aimilios, et al.
Veröffentlicht: (2026)
A Precision-Scalable RISC-V DNN Processor with On-Device Learning Capability at the Extreme Edge
von: Huang, Longwei, et al.
Veröffentlicht: (2023)
von: Huang, Longwei, et al.
Veröffentlicht: (2023)
FiCABU: A Fisher-Based, Context-Adaptive Machine Unlearning Processor for Edge AI
von: Cho, Eun-Su, et al.
Veröffentlicht: (2025)
von: Cho, Eun-Su, et al.
Veröffentlicht: (2025)
Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective
von: Li, Jinhao, et al.
Veröffentlicht: (2024)
von: Li, Jinhao, et al.
Veröffentlicht: (2024)
Dynamic Tsetlin Machine Accelerators for On-Chip Training at the Edge using FPGAs
von: Mao, Gang, et al.
Veröffentlicht: (2025)
von: Mao, Gang, et al.
Veröffentlicht: (2025)
MAx-DNN: Multi-Level Arithmetic Approximation for Energy-Efficient DNN Hardware Accelerators
von: Leon, Vasileios, et al.
Veröffentlicht: (2025)
von: Leon, Vasileios, et al.
Veröffentlicht: (2025)
Leveraging Highly Approximated Multipliers in DNN Inference
von: Zervakis, Georgios, et al.
Veröffentlicht: (2024)
von: Zervakis, Georgios, et al.
Veröffentlicht: (2024)
Leveraging Stochastic Depth Training for Adaptive Inference
von: Korol, Guilherme, et al.
Veröffentlicht: (2025)
von: Korol, Guilherme, et al.
Veröffentlicht: (2025)
Atleus: Accelerating Transformers on the Edge Enabled by 3D Heterogeneous Manycore Architectures
von: Dhingra, Pratyush, et al.
Veröffentlicht: (2025)
von: Dhingra, Pratyush, et al.
Veröffentlicht: (2025)
Hardware Software Optimizations for Fast Model Recovery on Reconfigurable Architectures
von: Xu, Bin, et al.
Veröffentlicht: (2025)
von: Xu, Bin, et al.
Veröffentlicht: (2025)
TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
von: Qiao, Ye, et al.
Veröffentlicht: (2025)
Decentor-V: Lightweight ML Training on Low-Power RISC-V Edge Devices
von: Ribeiro, Marcelo, et al.
Veröffentlicht: (2025)
von: Ribeiro, Marcelo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AIE4ML: An End-to-End Framework for Compiling Neural Networks for the Next Generation of AMD AI Engines
von: Danopoulos, Dimitrios, et al.
Veröffentlicht: (2025) -
TransAxx: Efficient Transformers with Approximate Computing
von: Danopoulos, Dimitrios, et al.
Veröffentlicht: (2024) -
Design Rules for Extreme-Edge Scientific Computing on AI Engines
von: Ma, Zhenghua, et al.
Veröffentlicht: (2026) -
Hardware-Efficient Softmax and Layer Normalization with Guaranteed Normalization for Edge Devices
von: Choi, Dawon, et al.
Veröffentlicht: (2026) -
SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
von: Wang, Wenxun, et al.
Veröffentlicht: (2025)