Hardware Software Optimizations for Fast Model Recovery on Reconfigurable Architectures
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Xu, Bin, Banerjee, Ayan, Gupta, Sandeep |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Enabling Physical AI at the Edge: Hardware-Accelerated Recovery of System Dynamics
par: Xu, Bin, et autres
Publié: (2025)
par: Xu, Bin, et autres
Publié: (2025)
TCL: Enabling Fast and Efficient Cross-Hardware Tensor Program Optimization via Continual Learning
par: Shen, Chaoyao, et autres
Publié: (2026)
par: Shen, Chaoyao, et autres
Publié: (2026)
QiMeng: Fully Automated Hardware and Software Design for Processor Chip
par: Zhang, Rui, et autres
Publié: (2025)
par: Zhang, Rui, et autres
Publié: (2025)
CATransformers: Carbon Aware Transformers Through Joint Model-Hardware Optimization
par: Wang, Irene, et autres
Publié: (2025)
par: Wang, Irene, et autres
Publié: (2025)
Hardware-Software Co-Design of Scalable, Energy-Efficient Analog Recurrent Computations
par: Fyon, Arthur, et autres
Publié: (2026)
par: Fyon, Arthur, et autres
Publié: (2026)
Energy Efficient Software Hardware CoDesign for Machine Learning: From TinyML to Large Language Models
par: Vahdatpour, Mohammad Saleh, et autres
Publié: (2026)
par: Vahdatpour, Mohammad Saleh, et autres
Publié: (2026)
FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs
par: Xie, Xilong, et autres
Publié: (2025)
par: Xie, Xilong, et autres
Publié: (2025)
SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference
par: Wang, Wenxun, et autres
Publié: (2025)
par: Wang, Wenxun, et autres
Publié: (2025)
Exploring the Limitations of Kolmogorov-Arnold Networks in Classification: Insights to Software Training and Hardware Implementation
par: Tran, Van Duy, et autres
Publié: (2024)
par: Tran, Van Duy, et autres
Publié: (2024)
AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIM
par: Zhang, Yuanpeng, et autres
Publié: (2025)
par: Zhang, Yuanpeng, et autres
Publié: (2025)
Reconfigurable Edge Hardware for Intelligent IDS: Systematic Approach
par: Foudhaili, Wadid, et autres
Publié: (2024)
par: Foudhaili, Wadid, et autres
Publié: (2024)
Low Power Vision Transformer Accelerator with Hardware-Aware Pruning and Optimized Dataflow
par: Hsiung, Ching-Lin, et autres
Publié: (2025)
par: Hsiung, Ching-Lin, et autres
Publié: (2025)
Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective
par: Li, Jinhao, et autres
Publié: (2024)
par: Li, Jinhao, et autres
Publié: (2024)
Rescaling-Aware Training for Efficient Deployment of Deep Learning Models on Full-Integer Hardware
par: Mueller, Lion, et autres
Publié: (2025)
par: Mueller, Lion, et autres
Publié: (2025)
On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration
par: Xiang, Maoyang, et autres
Publié: (2025)
par: Xiang, Maoyang, et autres
Publié: (2025)
LLM4DV: Using Large Language Models for Hardware Test Stimuli Generation
par: Zhang, Zixi, et autres
Publié: (2023)
par: Zhang, Zixi, et autres
Publié: (2023)
A Hardware-Aware, Per-Layer Methodology for Post-Training Quantization of Large Language Models
par: Killian, Earl
Publié: (2026)
par: Killian, Earl
Publié: (2026)
EvolveGen: Algorithmic Level Hardware Model Checking Benchmark Generation through Reinforcement Learning
par: Hu, Guangyu, et autres
Publié: (2026)
par: Hu, Guangyu, et autres
Publié: (2026)
A Joint Learning Approach to Hardware Caching and Prefetching
par: Yuan, Samuel, et autres
Publié: (2025)
par: Yuan, Samuel, et autres
Publié: (2025)
Energy-Aware Deep Learning on Resource-Constrained Hardware
par: Millar, Josh, et autres
Publié: (2025)
par: Millar, Josh, et autres
Publié: (2025)
Reusing Softmax Hardware Unit for GELU Computation in Transformers
par: Peltekis, Christodoulos, et autres
Publié: (2024)
par: Peltekis, Christodoulos, et autres
Publié: (2024)
TrainDeeploy: Hardware-Accelerated Parameter-Efficient Fine-Tuning of Small Transformer Models at the Extreme Edge
par: Wang, Run, et autres
Publié: (2026)
par: Wang, Run, et autres
Publié: (2026)
FORTALESA: Fault-Tolerant Reconfigurable Systolic Array for DNN Inference
par: Cherezova, Natalia, et autres
Publié: (2025)
par: Cherezova, Natalia, et autres
Publié: (2025)
PQA: Exploring the Potential of Product Quantization in DNN Hardware Acceleration
par: AbouElhamayed, Ahmed F., et autres
Publié: (2023)
par: AbouElhamayed, Ahmed F., et autres
Publié: (2023)
ATHEENA: A Toolflow for Hardware Early-Exit Network Automation
par: Biggs, Benjamin, et autres
Publié: (2023)
par: Biggs, Benjamin, et autres
Publié: (2023)
HASS: Hardware-Aware Sparsity Search for Dataflow DNN Accelerator
par: Yu, Zhewen, et autres
Publié: (2024)
par: Yu, Zhewen, et autres
Publié: (2024)
Learning Generalizable Program and Architecture Representations for Performance Modeling
par: Li, Lingda, et autres
Publié: (2023)
par: Li, Lingda, et autres
Publié: (2023)
Hardware/Software Co-Design of RISC-V Extensions for Accelerating Sparse DNNs on FPGAs
par: Sabih, Muhammad, et autres
Publié: (2025)
par: Sabih, Muhammad, et autres
Publié: (2025)
An FPGA-Based Reconfigurable Accelerator for Convolution-Transformer Hybrid EfficientViT
par: Shao, Haikuo, et autres
Publié: (2024)
par: Shao, Haikuo, et autres
Publié: (2024)
Hardware-Friendly Delayed-Feedback Reservoir for Multivariate Time-Series Classification
par: Ikeda, Sosei, et autres
Publié: (2025)
par: Ikeda, Sosei, et autres
Publié: (2025)
Low-Cost FlashAttention with Fused Exponential and Multiplication Hardware Operators
par: Alexandridis, Kosmas, et autres
Publié: (2025)
par: Alexandridis, Kosmas, et autres
Publié: (2025)
Hardware implementation of timely reliable Bayesian decision-making using memristors
par: Song, Lekai, et autres
Publié: (2024)
par: Song, Lekai, et autres
Publié: (2024)
Hardware-Aware Neural Dropout Search for Reliable Uncertainty Prediction on FPGA
par: Zhang, Zehuan, et autres
Publié: (2024)
par: Zhang, Zehuan, et autres
Publié: (2024)
Algorithm and Hardware Co-Design for Efficient Complex-Valued Uncertainty Estimation
par: Zhang, Zehuan, et autres
Publié: (2026)
par: Zhang, Zehuan, et autres
Publié: (2026)
Hardware-Efficient Softmax and Layer Normalization with Guaranteed Normalization for Edge Devices
par: Choi, Dawon, et autres
Publié: (2026)
par: Choi, Dawon, et autres
Publié: (2026)
VeriBug: An Attention-based Framework for Bug-Localization in Hardware Designs
par: Stracquadanio, Giuseppe, et autres
Publié: (2024)
par: Stracquadanio, Giuseppe, et autres
Publié: (2024)
From LLM to Silicon: RL-Driven ASIC Architecture Exploration for On-Device AI Inference
par: Ganti, Ravindra, et autres
Publié: (2026)
par: Ganti, Ravindra, et autres
Publié: (2026)
PowerGenie: Analytically-Guided Evolutionary Discovery of Superior Reconfigurable Power Converters
par: Gao, Jian, et autres
Publié: (2026)
par: Gao, Jian, et autres
Publié: (2026)
When Forgetting Builds Reliability: LLM Unlearning for Reliable Hardware Code Generation
par: Liang, Yiwen, et autres
Publié: (2025)
par: Liang, Yiwen, et autres
Publié: (2025)
SigmaQuant: Hardware-Aware Heterogeneous Quantization Method for Edge DNN Inference
par: Liu, Qunyou, et autres
Publié: (2026)
par: Liu, Qunyou, et autres
Publié: (2026)
Documents similaires
-
Enabling Physical AI at the Edge: Hardware-Accelerated Recovery of System Dynamics
par: Xu, Bin, et autres
Publié: (2025) -
TCL: Enabling Fast and Efficient Cross-Hardware Tensor Program Optimization via Continual Learning
par: Shen, Chaoyao, et autres
Publié: (2026) -
QiMeng: Fully Automated Hardware and Software Design for Processor Chip
par: Zhang, Rui, et autres
Publié: (2025) -
CATransformers: Carbon Aware Transformers Through Joint Model-Hardware Optimization
par: Wang, Irene, et autres
Publié: (2025) -
Hardware-Software Co-Design of Scalable, Energy-Efficient Analog Recurrent Computations
par: Fyon, Arthur, et autres
Publié: (2026)