Flex-PE: Flexible and SIMD Multi-Precision Processing Element for AI Workloads
Fuente:
arXiv
Saved in:
| Main Authors: | Lokhande, Mukul, Raut, Gopal, Vishvakarma, Santosh Kumar |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TREA: Low-precision Time-Multiplexed, Resource-Efficient Edge Accelerator for Object Detection and Classification
by: Sharma, Vijay Pratap, et al.
Published: (2026)
by: Sharma, Vijay Pratap, et al.
Published: (2026)
Res-DPU: Resource-shared Digital Processing-in-memory Unit for Edge-AI Workloads
by: Lokhande, Mukul, et al.
Published: (2025)
by: Lokhande, Mukul, et al.
Published: (2025)
POLARON: Precision-aware On-device Learning and Adaptive Runtime-cONfigurable AI acceleration
by: Lokhande, Mukul, et al.
Published: (2025)
by: Lokhande, Mukul, et al.
Published: (2025)
L-SPINE: A Low-Precision SIMD Spiking Neural Compute Engine for Resource-efficient Edge Inference
by: Kumar, Sonu, et al.
Published: (2026)
by: Kumar, Sonu, et al.
Published: (2026)
FERMI-ML: A Flexible and Resource-Efficient Memory-In-Situ SRAM Macro for TinyML acceleration
by: Lokhande, Mukul, et al.
Published: (2025)
by: Lokhande, Mukul, et al.
Published: (2025)
Retrospective: A CORDIC Based Configurable Activation Function for NN Applications
by: Kokane, Omkar, et al.
Published: (2025)
by: Kokane, Omkar, et al.
Published: (2025)
XR-NPE: High-Throughput Mixed-precision SIMD Neural Processing Engine for Extended Reality Perception Workloads
by: Chaudhari, Tejas, et al.
Published: (2025)
by: Chaudhari, Tejas, et al.
Published: (2025)
SPADE: A SIMD Posit-enabled compute engine for Accelerating DNN Efficiency
by: Kumar, Sonu, et al.
Published: (2026)
by: Kumar, Sonu, et al.
Published: (2026)
CARMEN: CORDIC-Accelerated Resource-Efficient Multi-Precision Inference Engine for Deep Learning
by: Kumar, Sonu, et al.
Published: (2026)
by: Kumar, Sonu, et al.
Published: (2026)
Leveraging SIMD for Accelerating Large-number Arithmetic
by: Das, Subhrajit, et al.
Published: (2026)
by: Das, Subhrajit, et al.
Published: (2026)
Open Challenges for a Production-ready Cloud Environment on top of RISC-V hardware
by: Call, Aaron, et al.
Published: (2025)
by: Call, Aaron, et al.
Published: (2025)
From GPUs to RRAMs: Distributed In-Memory Primal-Dual Hybrid Gradient Method for Solving Large-Scale Linear Optimization Problem
by: Vo, Huynh Q. N., et al.
Published: (2025)
by: Vo, Huynh Q. N., et al.
Published: (2025)
CLAASIC: a Cortex-Inspired Hardware Accelerator
by: Puente, Valentin, et al.
Published: (2016)
by: Puente, Valentin, et al.
Published: (2016)
DFabric: Scaling Out Data Parallel Applications with CXL-Ethernet Hybrid Interconnects
by: Zhang, Xu, et al.
Published: (2024)
by: Zhang, Xu, et al.
Published: (2024)
HYDRA: Hybrid Data Multiplexing and Run-time Layer Configurable DNN Accelerator
by: Kumar, Sonu, et al.
Published: (2024)
by: Kumar, Sonu, et al.
Published: (2024)
Transforming the Hybrid Cloud for Emerging AI Workloads
by: Chen, Deming, et al.
Published: (2024)
by: Chen, Deming, et al.
Published: (2024)
FlexStep: Enabling Flexible Error Detection in Multi/Many-core Real-time Systems
by: Wang, Tinglue, et al.
Published: (2025)
by: Wang, Tinglue, et al.
Published: (2025)
FlexVector: A SpMM Vector Processor with Flexible VRF for GCNs on Varying-Sparsity Graphs
by: Li, Bohan, et al.
Published: (2026)
by: Li, Bohan, et al.
Published: (2026)
EULER-ADAS: Energy-Efficient & SIMD-Unified Logarithmic-Posit Engine for Precision-Reconfigurable Approximate ADAS Acceleration
by: Lokhande, Mukul, et al.
Published: (2026)
by: Lokhande, Mukul, et al.
Published: (2026)
Proteus: Enabling High-Performance Processing-Using-DRAM with Dynamic Bit-Precision, Adaptive Data Representation, and Flexible Arithmetic
by: Oliveira, Geraldo F., et al.
Published: (2025)
by: Oliveira, Geraldo F., et al.
Published: (2025)
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
by: Afzal, Ayesha, et al.
Published: (2026)
by: Afzal, Ayesha, et al.
Published: (2026)
Data-aware Dynamic Execution of Irregular Workloads on Heterogeneous Systems
by: Bai, Zhenyu, et al.
Published: (2025)
by: Bai, Zhenyu, et al.
Published: (2025)
E-ReCON: An Energy- and Resource-Efficient Precision-Configurable Sparse nvCIM Macro for Conventional and Spiking Neural Edge Inference
by: Tenwar, Ankit Kumar, et al.
Published: (2026)
by: Tenwar, Ankit Kumar, et al.
Published: (2026)
Managed-Retention Memory: A New Class of Memory for the AI Era
by: Legtchenko, Sergey, et al.
Published: (2025)
by: Legtchenko, Sergey, et al.
Published: (2025)
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference
by: Ortega, Cristobal, et al.
Published: (2024)
by: Ortega, Cristobal, et al.
Published: (2024)
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
by: Adnan, Muhammad, et al.
Published: (2024)
by: Adnan, Muhammad, et al.
Published: (2024)
HARP: A Taxonomy for Heterogeneous and Hierarchical Processors for Mixed-reuse Workloads
by: Garg, Raveesh, et al.
Published: (2025)
by: Garg, Raveesh, et al.
Published: (2025)
PUDTune: Multi-Level Charging for High-Precision Calibration in Processing-Using-DRAM
by: Kubo, Tatsuya, et al.
Published: (2025)
by: Kubo, Tatsuya, et al.
Published: (2025)
PID-Comm: A Fast and Flexible Collective Communication Framework for Commodity Processing-in-DIMM Devices
by: Noh, Si Ung, et al.
Published: (2024)
by: Noh, Si Ung, et al.
Published: (2024)
Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis
by: Chu, Xiaoyu, et al.
Published: (2024)
by: Chu, Xiaoyu, et al.
Published: (2024)
TreeVQA: A Tree-Structured Execution Framework for Shot Reduction in Variational Quantum Algorithms
by: Hou, Yuewen, et al.
Published: (2025)
by: Hou, Yuewen, et al.
Published: (2025)
Architecting Distributed Quantum Computers: Design Insights from Resource Estimation
by: Filippov, Dmitry, et al.
Published: (2025)
by: Filippov, Dmitry, et al.
Published: (2025)
ForgetMeNot: Understanding and Modeling the Impact of Forever Chemicals Toward Sustainable Large-Scale Computing
by: Roy, Rohan Basu, et al.
Published: (2025)
by: Roy, Rohan Basu, et al.
Published: (2025)
Carbon Connect: An Ecosystem for Sustainable Computing
by: Lee, Benjamin C., et al.
Published: (2024)
by: Lee, Benjamin C., et al.
Published: (2024)
Reference Architecture of a Quantum-Centric Supercomputer
by: Seelam, Seetharami, et al.
Published: (2026)
by: Seelam, Seetharami, et al.
Published: (2026)
Implementation and Evaluation of GBDI Memory Compression Algorithm Using C/C++ on a Broader Range of Workloads
by: Aina, Adeyemi
Published: (2025)
by: Aina, Adeyemi
Published: (2025)
Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators
by: Shen, Aofeng, et al.
Published: (2025)
by: Shen, Aofeng, et al.
Published: (2025)
Efficient Optimization Accelerator Framework for Multistate Ising Problems
by: Garg, Chirag, et al.
Published: (2025)
by: Garg, Chirag, et al.
Published: (2025)
CORVET: A CORDIC-Powered, Resource-Frugal Mixed-Precision Vector Processing Engine for High-Throughput AIoT applications
by: Kumar, Sonu, et al.
Published: (2026)
by: Kumar, Sonu, et al.
Published: (2026)
Profiling Concurrent Vision Inference Workloads on NVIDIA Jetson -- Extended
by: Chakraborty, Abhinaba, et al.
Published: (2025)
by: Chakraborty, Abhinaba, et al.
Published: (2025)
Similar Items
-
TREA: Low-precision Time-Multiplexed, Resource-Efficient Edge Accelerator for Object Detection and Classification
by: Sharma, Vijay Pratap, et al.
Published: (2026) -
Res-DPU: Resource-shared Digital Processing-in-memory Unit for Edge-AI Workloads
by: Lokhande, Mukul, et al.
Published: (2025) -
POLARON: Precision-aware On-device Learning and Adaptive Runtime-cONfigurable AI acceleration
by: Lokhande, Mukul, et al.
Published: (2025) -
L-SPINE: A Low-Precision SIMD Spiking Neural Compute Engine for Resource-efficient Edge Inference
by: Kumar, Sonu, et al.
Published: (2026) -
FERMI-ML: A Flexible and Resource-Efficient Memory-In-Situ SRAM Macro for TinyML acceleration
by: Lokhande, Mukul, et al.
Published: (2025)