PolyThrottle: Energy-efficient Neural Network Inference on Edge Devices
Fuente:
arXiv
Guardado en:
| Autores principales: | Yan, Minghao, Wang, Hongyi, Venkataraman, Shivaram |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
xTern: Energy-Efficient Ternary Neural Network Inference on RISC-V-Based Edge Systems
por: Rutishauser, Georg, et al.
Publicado: (2024)
por: Rutishauser, Georg, et al.
Publicado: (2024)
Accelerating PoT Quantization on Edge Devices
por: Saha, Rappy, et al.
Publicado: (2024)
por: Saha, Rappy, et al.
Publicado: (2024)
SHIELD: A Segmented Hierarchical Memory Architecture for Energy-Efficient LLM Inference on Edge NPUs
por: Zhang, Jintao, et al.
Publicado: (2026)
por: Zhang, Jintao, et al.
Publicado: (2026)
Designing Efficient LLM Accelerators for Edge Devices
por: Haris, Jude, et al.
Publicado: (2024)
por: Haris, Jude, et al.
Publicado: (2024)
Efficient and Mathematically Robust Operations for Certified Neural Networks Inference
por: Geyer, Fabien, et al.
Publicado: (2024)
por: Geyer, Fabien, et al.
Publicado: (2024)
PolyLUT: Ultra-low Latency Polynomial Inference with Hardware-Aware Structured Pruning
por: Andronic, Marta, et al.
Publicado: (2025)
por: Andronic, Marta, et al.
Publicado: (2025)
Low-Energy On-Device Personalization for MCUs
por: Huang, Yushan, et al.
Publicado: (2024)
por: Huang, Yushan, et al.
Publicado: (2024)
Hardware-Efficient Softmax and Layer Normalization with Guaranteed Normalization for Edge Devices
por: Choi, Dawon, et al.
Publicado: (2026)
por: Choi, Dawon, et al.
Publicado: (2026)
A Precision-Scalable RISC-V DNN Processor with On-Device Learning Capability at the Extreme Edge
por: Huang, Longwei, et al.
Publicado: (2023)
por: Huang, Longwei, et al.
Publicado: (2023)
COBRA: Algorithm-Architecture Co-optimized Binary Transformer Accelerator for Edge Inference
por: Qiao, Ye, et al.
Publicado: (2025)
por: Qiao, Ye, et al.
Publicado: (2025)
On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration
por: Xiang, Maoyang, et al.
Publicado: (2025)
por: Xiang, Maoyang, et al.
Publicado: (2025)
PolyLUT: Learning Piecewise Polynomials for Ultra-Low Latency FPGA LUT-based Inference
por: Andronic, Marta, et al.
Publicado: (2023)
por: Andronic, Marta, et al.
Publicado: (2023)
SPARQ: Spiking Early-Exit Neural Networks for Energy-Efficient Edge AI
por: Patne, Parth, et al.
Publicado: (2026)
por: Patne, Parth, et al.
Publicado: (2026)
MiCo: End-to-End Mixed Precision Neural Network Co-Exploration Framework for Edge AI
por: Jiang, Zijun, et al.
Publicado: (2025)
por: Jiang, Zijun, et al.
Publicado: (2025)
Decentor-V: Lightweight ML Training on Low-Power RISC-V Edge Devices
por: Ribeiro, Marcelo, et al.
Publicado: (2025)
por: Ribeiro, Marcelo, et al.
Publicado: (2025)
A Data-Driven Approach to Dataflow-Aware Online Scheduling for Graph Neural Network Inference
por: Puigdemont, Pol, et al.
Publicado: (2024)
por: Puigdemont, Pol, et al.
Publicado: (2024)
Architectural Implications of Neural Network Inference for High Data-Rate, Low-Latency Scientific Applications
por: Weng, Olivia, et al.
Publicado: (2024)
por: Weng, Olivia, et al.
Publicado: (2024)
SigmaQuant: Hardware-Aware Heterogeneous Quantization Method for Edge DNN Inference
por: Liu, Qunyou, et al.
Publicado: (2026)
por: Liu, Qunyou, et al.
Publicado: (2026)
Taming the Exponential: A Fast Softmax Surrogate for Integer-Native Edge Inference
por: Danopoulos, Dimitrios, et al.
Publicado: (2026)
por: Danopoulos, Dimitrios, et al.
Publicado: (2026)
TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs
por: Qiao, Ye, et al.
Publicado: (2025)
por: Qiao, Ye, et al.
Publicado: (2025)
Clo-HDnn: A 4.66 TFLOPS/W and 3.78 TOPS/W Continual On-Device Learning Accelerator with Energy-efficient Hyperdimensional Computing via Progressive Search
por: Song, Chang Eun, et al.
Publicado: (2025)
por: Song, Chang Eun, et al.
Publicado: (2025)
Evaluating Four FPGA-accelerated Space Use Cases based on Neural Network Algorithms for On-board Inference
por: Antunes, Pedro, et al.
Publicado: (2026)
por: Antunes, Pedro, et al.
Publicado: (2026)
CHIME: Chiplet-based Heterogeneous Near-Memory Acceleration for Edge Multimodal LLM Inference
por: Chen, Yanru, et al.
Publicado: (2025)
por: Chen, Yanru, et al.
Publicado: (2025)
From LLM to Silicon: RL-Driven ASIC Architecture Exploration for On-Device AI Inference
por: Ganti, Ravindra, et al.
Publicado: (2026)
por: Ganti, Ravindra, et al.
Publicado: (2026)
TreeLUT: An Efficient Alternative to Deep Neural Networks for Inference Acceleration Using Gradient Boosted Decision Trees
por: Khataei, Alireza, et al.
Publicado: (2025)
por: Khataei, Alireza, et al.
Publicado: (2025)
P3-LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical Formats
por: Chen, Yuzong, et al.
Publicado: (2025)
por: Chen, Yuzong, et al.
Publicado: (2025)
PointODE: Lightweight Point Cloud Learning with Neural Ordinary Differential Equations on Edge
por: Sugiura, Keisuke, et al.
Publicado: (2025)
por: Sugiura, Keisuke, et al.
Publicado: (2025)
An FPGA-Based SoC Architecture with a RISC-V Controller for Energy-Efficient Temporal-Coding Spiking Neural Networks
por: Sekonji, Mohammad Javad, et al.
Publicado: (2026)
por: Sekonji, Mohammad Javad, et al.
Publicado: (2026)
Graph Neural Networks Based Analog Circuit Link Prediction
por: Pan, Guanyuan, et al.
Publicado: (2025)
por: Pan, Guanyuan, et al.
Publicado: (2025)
ROSA: Robust and Energy-Efficient Microring-Based Optical Neural Networks via Optical Shift-and-Add and Layer-Wise Hybrid Mapping
por: Zhang, Huifan, et al.
Publicado: (2026)
por: Zhang, Huifan, et al.
Publicado: (2026)
PolyLUT-Add: FPGA-based LUT Inference with Wide Inputs
por: Lou, Binglei, et al.
Publicado: (2024)
por: Lou, Binglei, et al.
Publicado: (2024)
The prediction of the quality of results in Logic Synthesis using Transformer and Graph Neural Networks
por: Yang, Chenghao, et al.
Publicado: (2022)
por: Yang, Chenghao, et al.
Publicado: (2022)
Algorithmic Strategies for Sustainable Reuse of Neural Network Accelerators with Permanent Faults
por: Alama, Youssef A. Ait, et al.
Publicado: (2024)
por: Alama, Youssef A. Ait, et al.
Publicado: (2024)
NeuralMatrix: Compute the Entire Neural Networks with Linear Matrix Operations for Efficient Inference
por: Sun, Ruiqi, et al.
Publicado: (2023)
por: Sun, Ruiqi, et al.
Publicado: (2023)
A Hybrid Edge Classifier: Combining TinyML-Optimised CNN with RRAM-CMOS ACAM for Energy-Efficient Inference
por: Woodward, Kieran, et al.
Publicado: (2025)
por: Woodward, Kieran, et al.
Publicado: (2025)
Enhancing Reliability of Neural Networks at the Edge: Inverted Normalization with Stochastic Affine Transformations
por: Ahmed, Soyed Tuhin, et al.
Publicado: (2024)
por: Ahmed, Soyed Tuhin, et al.
Publicado: (2024)
Approximate Multiplier Induced Error Propagation in Deep Neural Networks
por: Alahakoon, A. M. H. H., et al.
Publicado: (2025)
por: Alahakoon, A. M. H. H., et al.
Publicado: (2025)
Sustainable Transformer Neural Network Acceleration with Stochastic Photonic Computing
por: Afifi, S., et al.
Publicado: (2026)
por: Afifi, S., et al.
Publicado: (2026)
Histogram-Equalized Quantization for logic-gated Residual Neural Networks
por: Nguyen, Van Thien, et al.
Publicado: (2025)
por: Nguyen, Van Thien, et al.
Publicado: (2025)
Accelerating Sparse Graph Neural Networks with Tensor Core Optimization
por: Wu, Ka Wai
Publicado: (2024)
por: Wu, Ka Wai
Publicado: (2024)
Ejemplares similares
-
xTern: Energy-Efficient Ternary Neural Network Inference on RISC-V-Based Edge Systems
por: Rutishauser, Georg, et al.
Publicado: (2024) -
Accelerating PoT Quantization on Edge Devices
por: Saha, Rappy, et al.
Publicado: (2024) -
SHIELD: A Segmented Hierarchical Memory Architecture for Energy-Efficient LLM Inference on Edge NPUs
por: Zhang, Jintao, et al.
Publicado: (2026) -
Designing Efficient LLM Accelerators for Edge Devices
por: Haris, Jude, et al.
Publicado: (2024) -
Efficient and Mathematically Robust Operations for Certified Neural Networks Inference
por: Geyer, Fabien, et al.
Publicado: (2024)