SoftmAP: Software-Hardware Co-design for Integer-Only Softmax on Associative Processors
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rakka, Mariam, Li, Jinhao, Dai, Guohao, Eltawil, Ahmed, Fouda, Mohammed E., Kurdahi, Fadi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EdgeCIM: A Hardware-Software Co-Design for CIM-Based Acceleration of Small Language Models
von: Bazzi, Jinane, et al.
Veröffentlicht: (2026)
von: Bazzi, Jinane, et al.
Veröffentlicht: (2026)
BF-IMNA: A Bit Fluid In-Memory Neural Architecture for Neural Network Acceleration
von: Rakka, Mariam, et al.
Veröffentlicht: (2024)
von: Rakka, Mariam, et al.
Veröffentlicht: (2024)
Towards Efficient IMC Accelerator Design Through Joint Hardware-Workload Co-optimization
von: Krestinskaya, Olga, et al.
Veröffentlicht: (2024)
von: Krestinskaya, Olga, et al.
Veröffentlicht: (2024)
Joint Hardware-Workload Co-Optimization for In-Memory Computing Accelerators
von: Krestinskaya, Olga, et al.
Veröffentlicht: (2026)
von: Krestinskaya, Olga, et al.
Veröffentlicht: (2026)
SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
von: Wang, Wenxun, et al.
Veröffentlicht: (2025)
von: Wang, Wenxun, et al.
Veröffentlicht: (2025)
MERE: Hardware-Software Co-Design for Masking Cache Miss Latency in Embedded Processors
von: You, Dean, et al.
Veröffentlicht: (2025)
von: You, Dean, et al.
Veröffentlicht: (2025)
Hardware-Software Co-design for 3D-DRAM-based LLM Serving Accelerator
von: Li, Cong, et al.
Veröffentlicht: (2026)
von: Li, Cong, et al.
Veröffentlicht: (2026)
Inclusive-PIM: Hardware-Software Co-design for Broad Acceleration on Commercial PIM Architectures
von: Alsop, Johnathan, et al.
Veröffentlicht: (2023)
von: Alsop, Johnathan, et al.
Veröffentlicht: (2023)
ADDT -- A Digital Twin Framework for Proactive Safety Validation in Autonomous Driving Systems
von: Yu, Bo, et al.
Veröffentlicht: (2025)
von: Yu, Bo, et al.
Veröffentlicht: (2025)
HSCO-Bench: An Agent-Driven End-to-End Hardware-Software Co-design Benchmark for Systems-on-Chip
von: Tsai, Pei-Huan, et al.
Veröffentlicht: (2026)
von: Tsai, Pei-Huan, et al.
Veröffentlicht: (2026)
CIMNAS: A Joint Framework for Compute-In-Memory-Aware Neural Architecture Search
von: Krestinskaya, Olga, et al.
Veröffentlicht: (2025)
von: Krestinskaya, Olga, et al.
Veröffentlicht: (2025)
QiMeng: Fully Automated Hardware and Software Design for Processor Chip
von: Zhang, Rui, et al.
Veröffentlicht: (2025)
von: Zhang, Rui, et al.
Veröffentlicht: (2025)
Finesse: An Agile Design Framework for Pairing-based Cryptography via Software/Hardware Co-Design
von: Pan, Tianwei, et al.
Veröffentlicht: (2025)
von: Pan, Tianwei, et al.
Veröffentlicht: (2025)
GenDRAM:Hardware-Software Co-Design of General Platform in DRAM
von: Lu, Tsung-Han, et al.
Veröffentlicht: (2026)
von: Lu, Tsung-Han, et al.
Veröffentlicht: (2026)
Direct Integer Division in RNS and its Hardware Solutions
von: Olsen, Eric B.
Veröffentlicht: (2026)
von: Olsen, Eric B.
Veröffentlicht: (2026)
TurboFuzz: FPGA Accelerated Hardware Fuzzing for Processor Agile Verification
von: Zhong, Yang, et al.
Veröffentlicht: (2025)
von: Zhong, Yang, et al.
Veröffentlicht: (2025)
Reusing Softmax Hardware Unit for GELU Computation in Transformers
von: Peltekis, Christodoulos, et al.
Veröffentlicht: (2024)
von: Peltekis, Christodoulos, et al.
Veröffentlicht: (2024)
Implementation of Compute Intensive Algorithms on Software Configurable Processor
von: Ganesha, et al.
Veröffentlicht: (2025)
von: Ganesha, et al.
Veröffentlicht: (2025)
METRO: A Software-Hardware Co-Design of Interconnections for Spatial DNN Accelerators
von: Wang, Zhao, et al.
Veröffentlicht: (2021)
von: Wang, Zhao, et al.
Veröffentlicht: (2021)
Hardware-Software Co-Design for Accelerating Transformer Inference Leveraging Compute-in-Memory
von: Kim, Dong Eun, et al.
Veröffentlicht: (2025)
von: Kim, Dong Eun, et al.
Veröffentlicht: (2025)
NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering
von: Zhou, Zhe, et al.
Veröffentlicht: (2024)
von: Zhou, Zhe, et al.
Veröffentlicht: (2024)
Image processing Application Development on Software Configurable Processor Array
von: Prabhu, Ganesh, et al.
Veröffentlicht: (2025)
von: Prabhu, Ganesh, et al.
Veröffentlicht: (2025)
MCMComm: Hardware-Software Co-Optimization for End-to-End Communication in Multi-Chip-Modules
von: Raj, Ritik, et al.
Veröffentlicht: (2025)
von: Raj, Ritik, et al.
Veröffentlicht: (2025)
PIM-FW: Hardware-Software Co-Design of All-pairs Shortest Paths in DRAM
von: Lu, Tsung-Han, et al.
Veröffentlicht: (2025)
von: Lu, Tsung-Han, et al.
Veröffentlicht: (2025)
Taming the Exponential: A Fast Softmax Surrogate for Integer-Native Edge Inference
von: Danopoulos, Dimitrios, et al.
Veröffentlicht: (2026)
von: Danopoulos, Dimitrios, et al.
Veröffentlicht: (2026)
Sparsity-Aware Streaming SNN Accelerator with Output-Channel Dataflow for Automatic Modulation Classification
von: Yang, Kuilian, et al.
Veröffentlicht: (2026)
von: Yang, Kuilian, et al.
Veröffentlicht: (2026)
Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective
von: Li, Jinhao, et al.
Veröffentlicht: (2024)
von: Li, Jinhao, et al.
Veröffentlicht: (2024)
Hardware-Software Co-Design for Event-Driven SNN Deployment on Low-Cost Neuromorphic FPGAs
von: Lee, Jiwoon, et al.
Veröffentlicht: (2026)
von: Lee, Jiwoon, et al.
Veröffentlicht: (2026)
Aquas: Enhancing Domain Specialization through Holistic Hardware-Software Co-Optimization based on MLIR
von: Zou, Yuyang, et al.
Veröffentlicht: (2025)
von: Zou, Yuyang, et al.
Veröffentlicht: (2025)
Hardware-Efficient Softmax and Layer Normalization with Guaranteed Normalization for Edge Devices
von: Choi, Dawon, et al.
Veröffentlicht: (2026)
von: Choi, Dawon, et al.
Veröffentlicht: (2026)
Lyra: A Hardware-Accelerated RISC-V Verification Framework with Generative Model-Based Processor Fuzzing
von: Huo, Juncheng, et al.
Veröffentlicht: (2025)
von: Huo, Juncheng, et al.
Veröffentlicht: (2025)
SkyByte: Architecting an Efficient Memory-Semantic CXL-based SSD with OS and Hardware Co-design
von: Zhang, Haoyang, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyang, et al.
Veröffentlicht: (2025)
MARCA: Mamba Accelerator with ReConfigurable Architecture
von: Li, Jinhao, et al.
Veröffentlicht: (2024)
von: Li, Jinhao, et al.
Veröffentlicht: (2024)
SeDA: Secure and Efficient DNN Accelerators with Hardware/Software Synergy
von: Xuan, Wei, et al.
Veröffentlicht: (2025)
von: Xuan, Wei, et al.
Veröffentlicht: (2025)
Microarchitectural Co-Optimization for Sustained Throughput of RISC-V Multi-Lane Chaining Vector Processors
von: Wang, Weiying, et al.
Veröffentlicht: (2026)
von: Wang, Weiying, et al.
Veröffentlicht: (2026)
Hardware-Software Co-Design of Scalable, Energy-Efficient Analog Recurrent Computations
von: Fyon, Arthur, et al.
Veröffentlicht: (2026)
von: Fyon, Arthur, et al.
Veröffentlicht: (2026)
Sparsity-Aware Hardware-Software Co-Design of Spiking Neural Networks: An Overview
von: Aliyev, Ilkin, et al.
Veröffentlicht: (2024)
von: Aliyev, Ilkin, et al.
Veröffentlicht: (2024)
ConSmax: Hardware-Friendly Alternative Softmax with Learnable Parameters
von: Liu, Shiwei, et al.
Veröffentlicht: (2024)
von: Liu, Shiwei, et al.
Veröffentlicht: (2024)
C2HLSC: Can LLMs Bridge the Software-to-Hardware Design Gap?
von: Collini, Luca, et al.
Veröffentlicht: (2024)
von: Collini, Luca, et al.
Veröffentlicht: (2024)
STAR: An Efficient Softmax Engine for Attention Model with RRAM Crossbar
von: Zhai, Yifeng, et al.
Veröffentlicht: (2024)
von: Zhai, Yifeng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
EdgeCIM: A Hardware-Software Co-Design for CIM-Based Acceleration of Small Language Models
von: Bazzi, Jinane, et al.
Veröffentlicht: (2026) -
BF-IMNA: A Bit Fluid In-Memory Neural Architecture for Neural Network Acceleration
von: Rakka, Mariam, et al.
Veröffentlicht: (2024) -
Towards Efficient IMC Accelerator Design Through Joint Hardware-Workload Co-optimization
von: Krestinskaya, Olga, et al.
Veröffentlicht: (2024) -
Joint Hardware-Workload Co-Optimization for In-Memory Computing Accelerators
von: Krestinskaya, Olga, et al.
Veröffentlicht: (2026) -
SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
von: Wang, Wenxun, et al.
Veröffentlicht: (2025)