Energy-Aware Deep Learning on Resource-Constrained Hardware
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Millar, Josh, Haddadi, Hamed, Madhavapeddy, Anil |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Benchmarking Ultra-Low-Power $μ$NPUs
von: Millar, Josh, et al.
Veröffentlicht: (2025)
von: Millar, Josh, et al.
Veröffentlicht: (2025)
Low-Energy On-Device Personalization for MCUs
von: Huang, Yushan, et al.
Veröffentlicht: (2024)
von: Huang, Yushan, et al.
Veröffentlicht: (2024)
Rescaling-Aware Training for Efficient Deployment of Deep Learning Models on Full-Integer Hardware
von: Mueller, Lion, et al.
Veröffentlicht: (2025)
von: Mueller, Lion, et al.
Veröffentlicht: (2025)
Exploring Quantization and Mapping Synergy in Hardware-Aware Deep Neural Network Accelerators
von: Klhufek, Jan, et al.
Veröffentlicht: (2024)
von: Klhufek, Jan, et al.
Veröffentlicht: (2024)
HASS: Hardware-Aware Sparsity Search for Dataflow DNN Accelerator
von: Yu, Zhewen, et al.
Veröffentlicht: (2024)
von: Yu, Zhewen, et al.
Veröffentlicht: (2024)
CATransformers: Carbon Aware Transformers Through Joint Model-Hardware Optimization
von: Wang, Irene, et al.
Veröffentlicht: (2025)
von: Wang, Irene, et al.
Veröffentlicht: (2025)
Hardware-Aware Neural Dropout Search for Reliable Uncertainty Prediction on FPGA
von: Zhang, Zehuan, et al.
Veröffentlicht: (2024)
von: Zhang, Zehuan, et al.
Veröffentlicht: (2024)
Energy Efficient Software Hardware CoDesign for Machine Learning: From TinyML to Large Language Models
von: Vahdatpour, Mohammad Saleh, et al.
Veröffentlicht: (2026)
von: Vahdatpour, Mohammad Saleh, et al.
Veröffentlicht: (2026)
Low Power Vision Transformer Accelerator with Hardware-Aware Pruning and Optimized Dataflow
von: Hsiung, Ching-Lin, et al.
Veröffentlicht: (2025)
von: Hsiung, Ching-Lin, et al.
Veröffentlicht: (2025)
SigmaQuant: Hardware-Aware Heterogeneous Quantization Method for Edge DNN Inference
von: Liu, Qunyou, et al.
Veröffentlicht: (2026)
von: Liu, Qunyou, et al.
Veröffentlicht: (2026)
Hardware-Software Co-Design of Scalable, Energy-Efficient Analog Recurrent Computations
von: Fyon, Arthur, et al.
Veröffentlicht: (2026)
von: Fyon, Arthur, et al.
Veröffentlicht: (2026)
AP-DRL: A Synergistic Algorithm-Hardware Framework for Automatic Task Partitioning of Deep Reinforcement Learning on Versal ACAP
von: Li, Enlai, et al.
Veröffentlicht: (2026)
von: Li, Enlai, et al.
Veröffentlicht: (2026)
PolyLUT: Ultra-low Latency Polynomial Inference with Hardware-Aware Structured Pruning
von: Andronic, Marta, et al.
Veröffentlicht: (2025)
von: Andronic, Marta, et al.
Veröffentlicht: (2025)
MAx-DNN: Multi-Level Arithmetic Approximation for Energy-Efficient DNN Hardware Accelerators
von: Leon, Vasileios, et al.
Veröffentlicht: (2025)
von: Leon, Vasileios, et al.
Veröffentlicht: (2025)
A Joint Learning Approach to Hardware Caching and Prefetching
von: Yuan, Samuel, et al.
Veröffentlicht: (2025)
von: Yuan, Samuel, et al.
Veröffentlicht: (2025)
Hardware-Aware Fine-Tuning of Spiking Q-Networks on the SpiNNaker2 Neuromorphic Platform
von: Arfa, Sirine, et al.
Veröffentlicht: (2025)
von: Arfa, Sirine, et al.
Veröffentlicht: (2025)
A Hardware-Aware, Per-Layer Methodology for Post-Training Quantization of Large Language Models
von: Killian, Earl
Veröffentlicht: (2026)
von: Killian, Earl
Veröffentlicht: (2026)
Hardware-Aware Data and Instruction Mapping for AI Tasks: Balancing Parallelism, I/O and Memory Tradeoffs
von: Chowdhury, Md Rownak Hossain, et al.
Veröffentlicht: (2025)
von: Chowdhury, Md Rownak Hossain, et al.
Veröffentlicht: (2025)
A Survey on Deep Learning Hardware Accelerators for Heterogeneous HPC Platforms
von: Silvano, Cristina, et al.
Veröffentlicht: (2023)
von: Silvano, Cristina, et al.
Veröffentlicht: (2023)
Torch2Chip: An End-to-end Customizable Deep Neural Network Compression and Deployment Toolkit for Prototype Hardware Accelerator Design
von: Meng, Jian, et al.
Veröffentlicht: (2024)
von: Meng, Jian, et al.
Veröffentlicht: (2024)
TCL: Enabling Fast and Efficient Cross-Hardware Tensor Program Optimization via Continual Learning
von: Shen, Chaoyao, et al.
Veröffentlicht: (2026)
von: Shen, Chaoyao, et al.
Veröffentlicht: (2026)
EvolveGen: Algorithmic Level Hardware Model Checking Benchmark Generation through Reinforcement Learning
von: Hu, Guangyu, et al.
Veröffentlicht: (2026)
von: Hu, Guangyu, et al.
Veröffentlicht: (2026)
Reusing Softmax Hardware Unit for GELU Computation in Transformers
von: Peltekis, Christodoulos, et al.
Veröffentlicht: (2024)
von: Peltekis, Christodoulos, et al.
Veröffentlicht: (2024)
Hardware Software Optimizations for Fast Model Recovery on Reconfigurable Architectures
von: Xu, Bin, et al.
Veröffentlicht: (2025)
von: Xu, Bin, et al.
Veröffentlicht: (2025)
PQA: Exploring the Potential of Product Quantization in DNN Hardware Acceleration
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2023)
von: AbouElhamayed, Ahmed F., et al.
Veröffentlicht: (2023)
ATHEENA: A Toolflow for Hardware Early-Exit Network Automation
von: Biggs, Benjamin, et al.
Veröffentlicht: (2023)
von: Biggs, Benjamin, et al.
Veröffentlicht: (2023)
On-Chip Hardware-Aware Quantization for Mixed Precision Neural Networks
von: Huang, Wei, et al.
Veröffentlicht: (2023)
von: Huang, Wei, et al.
Veröffentlicht: (2023)
QiMeng: Fully Automated Hardware and Software Design for Processor Chip
von: Zhang, Rui, et al.
Veröffentlicht: (2025)
von: Zhang, Rui, et al.
Veröffentlicht: (2025)
Hardware-Friendly Delayed-Feedback Reservoir for Multivariate Time-Series Classification
von: Ikeda, Sosei, et al.
Veröffentlicht: (2025)
von: Ikeda, Sosei, et al.
Veröffentlicht: (2025)
Low-Cost FlashAttention with Fused Exponential and Multiplication Hardware Operators
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2025)
von: Alexandridis, Kosmas, et al.
Veröffentlicht: (2025)
Hardware implementation of timely reliable Bayesian decision-making using memristors
von: Song, Lekai, et al.
Veröffentlicht: (2024)
von: Song, Lekai, et al.
Veröffentlicht: (2024)
Algorithm and Hardware Co-Design for Efficient Complex-Valued Uncertainty Estimation
von: Zhang, Zehuan, et al.
Veröffentlicht: (2026)
von: Zhang, Zehuan, et al.
Veröffentlicht: (2026)
Hardware-Efficient Softmax and Layer Normalization with Guaranteed Normalization for Edge Devices
von: Choi, Dawon, et al.
Veröffentlicht: (2026)
von: Choi, Dawon, et al.
Veröffentlicht: (2026)
Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective
von: Li, Jinhao, et al.
Veröffentlicht: (2024)
von: Li, Jinhao, et al.
Veröffentlicht: (2024)
VeriBug: An Attention-based Framework for Bug-Localization in Hardware Designs
von: Stracquadanio, Giuseppe, et al.
Veröffentlicht: (2024)
von: Stracquadanio, Giuseppe, et al.
Veröffentlicht: (2024)
When Forgetting Builds Reliability: LLM Unlearning for Reliable Hardware Code Generation
von: Liang, Yiwen, et al.
Veröffentlicht: (2025)
von: Liang, Yiwen, et al.
Veröffentlicht: (2025)
On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration
von: Xiang, Maoyang, et al.
Veröffentlicht: (2025)
von: Xiang, Maoyang, et al.
Veröffentlicht: (2025)
SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
von: Wang, Wenxun, et al.
Veröffentlicht: (2025)
von: Wang, Wenxun, et al.
Veröffentlicht: (2025)
LLM4DV: Using Large Language Models for Hardware Test Stimuli Generation
von: Zhang, Zixi, et al.
Veröffentlicht: (2023)
von: Zhang, Zixi, et al.
Veröffentlicht: (2023)
Exploring the Limitations of Kolmogorov-Arnold Networks in Classification: Insights to Software Training and Hardware Implementation
von: Tran, Van Duy, et al.
Veröffentlicht: (2024)
von: Tran, Van Duy, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Benchmarking Ultra-Low-Power $μ$NPUs
von: Millar, Josh, et al.
Veröffentlicht: (2025) -
Low-Energy On-Device Personalization for MCUs
von: Huang, Yushan, et al.
Veröffentlicht: (2024) -
Rescaling-Aware Training for Efficient Deployment of Deep Learning Models on Full-Integer Hardware
von: Mueller, Lion, et al.
Veröffentlicht: (2025) -
Exploring Quantization and Mapping Synergy in Hardware-Aware Deep Neural Network Accelerators
von: Klhufek, Jan, et al.
Veröffentlicht: (2024) -
HASS: Hardware-Aware Sparsity Search for Dataflow DNN Accelerator
von: Yu, Zhewen, et al.
Veröffentlicht: (2024)