DHFP-PE: Dual-Precision Hybrid Floating Point Processing Element for AI Acceleration
Fuente:
arXiv
Saved in:
| Main Authors: | Kumar, Shubham, Sharma, Vijay Pratap, Neema, Vaibhav, Vishvakarma, Santosh Kumar |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bhasha-Rupantarika: Algorithm-Hardware Co-design approach for Multilingual Neural Machine Translation
by: Lokhande, Mukul, et al.
Published: (2025)
by: Lokhande, Mukul, et al.
Published: (2025)
CARMEN: CORDIC-Accelerated Resource-Efficient Multi-Precision Inference Engine for Deep Learning
by: Kumar, Sonu, et al.
Published: (2026)
by: Kumar, Sonu, et al.
Published: (2026)
Flex-PE: Flexible and SIMD Multi-Precision Processing Element for AI Workloads
by: Lokhande, Mukul, et al.
Published: (2024)
by: Lokhande, Mukul, et al.
Published: (2024)
TREA: Low-precision Time-Multiplexed, Resource-Efficient Edge Accelerator for Object Detection and Classification
by: Sharma, Vijay Pratap, et al.
Published: (2026)
by: Sharma, Vijay Pratap, et al.
Published: (2026)
POLARON: Precision-aware On-device Learning and Adaptive Runtime-cONfigurable AI acceleration
by: Lokhande, Mukul, et al.
Published: (2025)
by: Lokhande, Mukul, et al.
Published: (2025)
ASAP-FE: Energy-Efficient Feature Extraction Enabling Multi-Channel Keyword Spotting on Edge Processors
by: Choi, Jongin, et al.
Published: (2025)
by: Choi, Jongin, et al.
Published: (2025)
A 71.2-$μ$W Speech Recognition Accelerator with Recurrent Spiking Neural Network
by: Yang, Chih-Chyau, et al.
Published: (2025)
by: Yang, Chih-Chyau, et al.
Published: (2025)
A 14uJ/Decision Keyword Spotting Accelerator with In-SRAM-Computing and On Chip Learning for Customization
by: Chiang, Yu-Hsiang, et al.
Published: (2022)
by: Chiang, Yu-Hsiang, et al.
Published: (2022)
Bio-RV: Low-Power Resource-Efficient RISC-V Processor for Biomedical Applications
by: Sharma, Vijay Pratap, et al.
Published: (2026)
by: Sharma, Vijay Pratap, et al.
Published: (2026)
Real-Time Piano Note Frequency Detection Using FPGA and FFT Core
by: Anik, Shafayet M., et al.
Published: (2025)
by: Anik, Shafayet M., et al.
Published: (2025)
Res-DPU: Resource-shared Digital Processing-in-memory Unit for Edge-AI Workloads
by: Lokhande, Mukul, et al.
Published: (2025)
by: Lokhande, Mukul, et al.
Published: (2025)
Acoustic Local Positioning With Encoded Emission Beacons
by: Urena, Jesus, et al.
Published: (2024)
by: Urena, Jesus, et al.
Published: (2024)
A Low-Power Streaming Speech Enhancement Accelerator For Edge Devices
by: Wu, Ci-Hao, et al.
Published: (2025)
by: Wu, Ci-Hao, et al.
Published: (2025)
SPADE: A SIMD Posit-enabled compute engine for Accelerating DNN Efficiency
by: Kumar, Sonu, et al.
Published: (2026)
by: Kumar, Sonu, et al.
Published: (2026)
HYDRA: Hybrid Data Multiplexing and Run-time Layer Configurable DNN Accelerator
by: Kumar, Sonu, et al.
Published: (2024)
by: Kumar, Sonu, et al.
Published: (2024)
L-SPINE: A Low-Precision SIMD Spiking Neural Compute Engine for Resource-efficient Edge Inference
by: Kumar, Sonu, et al.
Published: (2026)
by: Kumar, Sonu, et al.
Published: (2026)
E-ReCON: An Energy- and Resource-Efficient Precision-Configurable Sparse nvCIM Macro for Conventional and Spiking Neural Edge Inference
by: Tenwar, Ankit Kumar, et al.
Published: (2026)
by: Tenwar, Ankit Kumar, et al.
Published: (2026)
Folding Attention: Memory and Power Optimization for On-Device Transformer-based Streaming Speech Recognition
by: Li, Yang, et al.
Published: (2023)
by: Li, Yang, et al.
Published: (2023)
FERMI-ML: A Flexible and Resource-Efficient Memory-In-Situ SRAM Macro for TinyML acceleration
by: Lokhande, Mukul, et al.
Published: (2025)
by: Lokhande, Mukul, et al.
Published: (2025)
CORVET: A CORDIC-Powered, Resource-Frugal Mixed-Precision Vector Processing Engine for High-Throughput AIoT applications
by: Kumar, Sonu, et al.
Published: (2026)
by: Kumar, Sonu, et al.
Published: (2026)
EULER-ADAS: Energy-Efficient & SIMD-Unified Logarithmic-Posit Engine for Precision-Reconfigurable Approximate ADAS Acceleration
by: Lokhande, Mukul, et al.
Published: (2026)
by: Lokhande, Mukul, et al.
Published: (2026)
QForce-RL: Quantized FPGA-Optimized Reinforcement Learning Compute Engine
by: Jha, Anushka, et al.
Published: (2025)
by: Jha, Anushka, et al.
Published: (2025)
Prototype: A Keyword Spotting-Based Intelligent Audio SoC for IoT
by: Liang, Huihong, et al.
Published: (2025)
by: Liang, Huihong, et al.
Published: (2025)
XR-NPE: High-Throughput Mixed-precision SIMD Neural Processing Engine for Extended Reality Perception Workloads
by: Chaudhari, Tejas, et al.
Published: (2025)
by: Chaudhari, Tejas, et al.
Published: (2025)
Ultra-low power on-chip learning of speech commands with phase-change memories
by: Miriyala, Venkata Pavan Kumar, et al.
Published: (2020)
by: Miriyala, Venkata Pavan Kumar, et al.
Published: (2020)
DeltaKWS: A 65nm 36nJ/Decision Bio-inspired Temporal-Sparsity-Aware Digital Keyword Spotting IC with 0.6V Near-Threshold SRAM
by: Chen, Qinyu, et al.
Published: (2024)
by: Chen, Qinyu, et al.
Published: (2024)
TsetlinKWS: A 65nm 16.58uW, 0.63mm2 State-Driven Convolutional Tsetlin Machine-Based Accelerator For Keyword Spotting
by: Lin, Baizhou, et al.
Published: (2025)
by: Lin, Baizhou, et al.
Published: (2025)
An Energy-Efficient RFET-Based Stochastic Computing Neural Network Accelerator
by: Lu, Sheng, et al.
Published: (2025)
by: Lu, Sheng, et al.
Published: (2025)
CORDIC Is All You Need
by: Kokane, Omkar, et al.
Published: (2025)
by: Kokane, Omkar, et al.
Published: (2025)
FSL-HDnn: A 40 nm Few-shot On-Device Learning Accelerator with Integrated Feature Extraction and Hyperdimensional Computing
by: Xu, Weihong, et al.
Published: (2025)
by: Xu, Weihong, et al.
Published: (2025)
Modeling the Energy Consumption of the HEVC Software Encoding Process using Processor events
by: Ramasubbu, Geetha, et al.
Published: (2024)
by: Ramasubbu, Geetha, et al.
Published: (2024)
Technology-Circuit-Algorithm Tri-Design for Processing-in-Pixel-in-Memory (P2M)
by: Kaiser, Md Abdullah-Al, et al.
Published: (2023)
by: Kaiser, Md Abdullah-Al, et al.
Published: (2023)
Voltage-Controlled Magnetic Tunnel Junction based ADC-less Global Shutter Processing-in-Pixel for Extreme-Edge Intelligence
by: Kaiser, Md Abdullah-Al, et al.
Published: (2024)
by: Kaiser, Md Abdullah-Al, et al.
Published: (2024)
An Efficient Architecture and High-Throughput Implementation of CCSDS-123.0-B-2 Hybrid Entropy Coder Targeting Space-Grade SRAM FPGA Technology
by: Chatziantoniou, Panagiotis, et al.
Published: (2022)
by: Chatziantoniou, Panagiotis, et al.
Published: (2022)
Enhancing Real-World Active Speaker Detection with Multi-Modal Extraction Pre-Training
by: Tao, Ruijie, et al.
Published: (2024)
by: Tao, Ruijie, et al.
Published: (2024)
Retrospective: A CORDIC Based Configurable Activation Function for NN Applications
by: Kokane, Omkar, et al.
Published: (2025)
by: Kokane, Omkar, et al.
Published: (2025)
A 129FPS Full HD Real-Time Accelerator for 3D Gaussian Splatting
by: Chang, Fang-Chi, et al.
Published: (2026)
by: Chang, Fang-Chi, et al.
Published: (2026)
Sub-Millisecond Event-Based Eye Tracking on a Resource-Constrained Microcontroller
by: Giordano, Marco, et al.
Published: (2025)
by: Giordano, Marco, et al.
Published: (2025)
Pack my weights and run! Minimizing overheads for in-memory computing accelerators
by: Houshmand, Pouya, et al.
Published: (2024)
by: Houshmand, Pouya, et al.
Published: (2024)
FPCA: Field-Programmable Pixel Convolutional Array for Extreme-Edge Intelligence
by: Yin, Zihan, et al.
Published: (2024)
by: Yin, Zihan, et al.
Published: (2024)
Similar Items
-
Bhasha-Rupantarika: Algorithm-Hardware Co-design approach for Multilingual Neural Machine Translation
by: Lokhande, Mukul, et al.
Published: (2025) -
CARMEN: CORDIC-Accelerated Resource-Efficient Multi-Precision Inference Engine for Deep Learning
by: Kumar, Sonu, et al.
Published: (2026) -
Flex-PE: Flexible and SIMD Multi-Precision Processing Element for AI Workloads
by: Lokhande, Mukul, et al.
Published: (2024) -
TREA: Low-precision Time-Multiplexed, Resource-Efficient Edge Accelerator for Object Detection and Classification
by: Sharma, Vijay Pratap, et al.
Published: (2026) -
POLARON: Precision-aware On-device Learning and Adaptive Runtime-cONfigurable AI acceleration
by: Lokhande, Mukul, et al.
Published: (2025)